Screen interaction method and device, display equipment and storage medium
By analyzing key hand points and screen pose information in three-dimensional space, remote interactive control is achieved, solving the problem of limitations in interaction methods in existing technologies, improving the flexibility and accuracy of interaction, and reducing equipment costs.
Patent Information
- Application Number
- CN202410542527.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-04-30
- Publication Date
- 2025-10-31
AI Technical Summary
In existing technologies, mouse and touch screen operation methods limit the performance of intelligent interaction when users interact with the display screen, resulting in less flexible and accurate interaction.
By acquiring key points of the controller's hands, using 3D spatial analysis to determine virtual rays, and combining this with the pose information of the target screen, remote interactive control can be achieved, avoiding the need for additional equipment configuration.
It improves the flexibility and accuracy of interaction, reduces device configuration costs, and enhances intelligent interaction performance.
Smart Images

Figure CN120872218A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of intelligent interaction technology, specifically to a screen interaction method, device, display equipment, and storage medium. Background Technology
[0002] With the development of intelligent interaction technology, the application of interaction scenarios has become increasingly diversified. When users interact with the display screen, they generally control the interaction by changing the screen cursor or screen focus with a mouse, or by controlling the interaction through touch screen. However, these interaction methods limit users to operating the mouse only from where it is placed, or to standing in front of the screen and interacting through touch screen, which restricts intelligent interaction and reduces its performance. Summary of the Invention
[0003] This application provides a screen interaction method, apparatus, display device, and storage medium, aiming to solve the problem of low intelligent interaction performance in the prior art.
[0004] In a first aspect, this application provides a screen interaction method, including:
[0005] The key hand points of the controller who controls the target screen are obtained, and the key hand points are determined by user localization and skeleton tracking of the target image containing the controller.
[0006] Based on the three-dimensional position information of each of the aforementioned key hand points, at least two target key points are determined, and a virtual ray formed by the three-dimensional position information of the at least two target key points points points from the controller's hand to the target screen.
[0007] Based on the three-dimensional position information of the target key points and the pose information of the target screen, the target interaction position of the target screen is determined, and the target screen is interactively controlled.
[0008] This solution eliminates the need for additional interactive devices such as mice, reducing equipment configuration costs. Furthermore, it uses hand pointing analysis in three-dimensional space to determine the target interaction position, ensuring the accuracy of the interaction and thus improving interaction performance.
[0009] In one embodiment of this application, before acquiring the key hand points of the controller controlling the target screen, the method includes:
[0010] Acquire the target image for controlling the target screen;
[0011] Human detection is performed on the target image using a human detection model to identify the main controller in the target image and the hand area of the main controller;
[0012] The key points of the hand region are detected by the hand key point detection model to determine the two-dimensional position information of the hand joints of the controller, as well as the confidence level of the two-dimensional position information;
[0013] Based on the two-dimensional position information and the confidence level corresponding to each two-dimensional position information, key points of the hand are determined.
[0014] This solution uses a model to detect the hand region and key points, improving the efficiency of key point detection. At the same time, it determines the key points of the hand based on the confidence level corresponding to the two-dimensional position information, avoiding the limitations of image key point detection and ensuring the rationality and accuracy of the key points of the hand.
[0015] In one embodiment of this application, determining at least two target key points based on the three-dimensional position information of each of the hand key points includes:
[0016] The three-dimensional position information of the key hand points is obtained by predicting the three-dimensional position information of the key hand points using a hand joint model;
[0017] Based on the three-dimensional position information of each hand key point and the relative positional relationship between each hand key point, the hand key points of the same hand limb are determined;
[0018] Based on the hand key points of each hand limb and the pose information of the target screen, the target hand key points of the hand limb pointing to the target screen are determined.
[0019] This solution uses a hand joint model to predict the three-dimensional position information of the two-dimensional hand key points. The process of obtaining the three-dimensional position information of the hand key points can improve the calculation accuracy of the three-dimensional position information. At the same time, it determines the direction of each hand limb to ensure the accuracy of determining the target hand key points of the hand limbs pointing to the target screen.
[0020] In one embodiment of this application, determining the target hand key points of the hand limb pointing to the target screen based on the hand key points of each hand limb and the pose information of the target screen includes:
[0021] The bending angle of each segment is calculated based on the three-dimensional position information of every two adjacent key points of the same hand limb.
[0022] The bending angle of each segment in each hand limb is counted to obtain the bending degree of the hand limb;
[0023] Acquire the target hand limb with a curvature smaller than a preset curvature;
[0024] Based on the three-dimensional position information of the target hand limb and the pose information of the target screen, the key points of the target hand limb pointing to the target screen are determined.
[0025] In this scheme, during the calculation of virtual rays for the hand, if all hand limbs are in a bent state, no ray is generated for that hand. The purpose is to avoid the judgment and calculation of unnecessary hand limbs, reduce the amount of data calculation, and improve the accuracy of pointing judgment.
[0026] In one embodiment of this application, determining the target interaction position of the target screen based on the three-dimensional position information of the target key points and the pose information of the target screen includes:
[0027] Based on the three-dimensional position information of the target key points, the main control key points and auxiliary key points are determined. The main control key point is the fingertip key point, and the auxiliary key points are at least one of the wrist key point, interphalangeal key point, and palmofinite key point.
[0028] Based on the pose information of the target screen and the virtual ray formed by the main control key point and the auxiliary key point, the target interaction position of the target screen is determined.
[0029] In this scheme, the three-dimensional position information of the target key points is used to determine the main control key points and auxiliary key points, thereby improving the accuracy of the pointing calculation and avoiding the limitations caused by directly using the target's hand key points to calculate the target's interactive position.
[0030] In one embodiment of this application, the interactive control of the target screen includes:
[0031] If a cursor control exists on the target screen, then control the cursor control to move to the target interactive position;
[0032] In response to a trigger operation at the target interaction location, interact with the target screen.
[0033] This solution generates cursor movement instructions based on the target interaction position to control the cursor control to move to the target interaction position, thereby realizing gesture-controlled cursor movement interaction, improving the flexibility of interaction control, and reducing the cost of interaction devices.
[0034] In one embodiment of this application, the determination of the pose information includes the following steps:
[0035] Acquire a reference image and a calibration image, wherein the target image and the reference image include a reference object, the reference image is acquired by a target imaging device that acquires the target image, and the calibration image is acquired by an auxiliary imaging device that captures the target screen, and the target imaging device and the auxiliary imaging device have a common viewing area;
[0036] The first pose information of the reference object is determined based on the image coordinates of the reference object in the reference image;
[0037] Based on the image coordinates of the reference object in the calibration image, the second pose information of the reference object is determined.
[0038] The pose information of the target screen is determined based on the relative pose information between the first pose information and the second pose information.
[0039] Secondly, this application also provides a screen interaction device, the device comprising:
[0040] The acquisition module is used to acquire the key hand points of the master controller who controls the target screen. The key hand points are determined by user positioning and skeleton tracking of the target image containing the master controller.
[0041] The determination module is used to determine at least two target key points based on the three-dimensional position information of each of the hand key points, and the virtual ray formed by the three-dimensional position information of the at least two target key points points points from the hand of the controller to the target screen;
[0042] The interactive control module is used to determine the target interactive position of the target screen based on the three-dimensional position information of the target key points and the pose information of the target screen, and to perform interactive control on the target screen.
[0043] Thirdly, this application also provides a display device, the display device comprising:
[0044] One or more processors;
[0045] Memory; and
[0046] One or more applications, wherein the one or more applications are stored in the memory and configured to be executed by the processor to implement the steps in any of the screen interaction methods described above.
[0047] Fourthly, this application also provides a computer-readable storage medium having a computer program stored thereon, the computer program being loaded by a processor to perform the steps in any of the screen interaction methods described herein.
[0048] This application provides a screen interaction method, apparatus, display device, and storage medium. It acquires key hand points of a user controlling a target screen, determined through user localization and skeletal tracking of a target image containing the user. Based on the three-dimensional position information of each key hand point, at least two target key points are determined. A virtual ray formed by the three-dimensional position information of these two target key points points points from the user's hand towards the target screen. Based on the three-dimensional position information of the target key points and the pose information of the target screen, a target interaction position on the target screen is determined, enabling interactive control of the target screen. This solution analyzes the pose of the key hand points and the target screen in three-dimensional space to determine the target interaction position of the user, thus enabling interactive control. It eliminates the need for additional interactive devices such as a mouse, reducing equipment configuration costs. Furthermore, the spatial pointing analysis based on three-dimensional space ensures the accuracy of the interaction and improves its performance. Attached Figure Description
[0049] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0050] Figure 1 This is a schematic diagram of a scenario for the screen interaction method provided in an embodiment of this application;
[0051] Figure 2 This is a schematic flowchart of an embodiment of the screen interaction method provided in this application.
[0052] Figure 3 A schematic flowchart of one implementation scheme for determining key hand points in the screen interaction method provided in this application embodiment;
[0053] Figure 4 A schematic diagram of one implementation scheme for determining target key points in the screen interaction method provided for the implementation scheme of this application;
[0054] Figure 5 A schematic diagram of constraint information between three-dimensional hands included in the hand joint model provided for the implementation scheme of this application;
[0055] Figure 6 A schematic diagram of one implementation scheme for determining the target interaction position in the screen interaction method provided in this application;
[0056] Figure 7This is a schematic flowchart of one implementation scheme for pose information calibration in the target screen interaction method provided in this application embodiment;
[0057] Figure 8 Flowchart of another embodiment of the screen interaction method provided in this application;
[0058] Figure 9 This is a schematic diagram of an embodiment of the screen interaction device provided in this application.
[0059] Figure 10 This is a schematic diagram of an embodiment of the display device provided in this application. Detailed Implementation
[0060] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0061] In the description of this invention, it should be understood that the terms "center," "longitudinal," "lateral," "length," "width," "thickness," "upper," "lower," "front," "rear," "left," "right," "vertical," "horizontal," "top," "bottom," "inner," and "outer," etc., indicating orientation or positional relationships based on the orientation or positional relationships shown in the accompanying drawings, are only for the convenience of describing the invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation, and therefore should not be construed as a limitation of the invention. Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Thus, features defined with "first" and "second" may explicitly or implicitly include one or more of the stated features. In the description of this invention, "a plurality of" means two or more, unless otherwise explicitly specified.
[0062] In this embodiment, "and / or" describes the relationship between associated objects, indicating that three relationships can exist. For example, A and / or B can represent three cases: A alone, A and B simultaneously, and B alone. Additionally, the character " / ", unless otherwise specified, generally indicates that the preceding and following associated objects have an "or" relationship.
[0063] In this application, the term "exemplary" is used to mean "serving as an example, illustration, or description." Any embodiment described as "exemplary" in this application is not necessarily to be construed as being more preferred or advantageous than other embodiments. The following description is provided to enable any person skilled in the art to make and use the invention. Details are set forth in the following description for purposes of explanation. It should be understood that those skilled in the art will recognize that the invention can be made without using these specific details. In other instances, well-known structures and processes will not be described in detail to avoid obscuring the description of the invention with unnecessary detail. Therefore, the invention is not intended to be limited to the embodiments shown, but is consistent with the broadest scope of the principles and features disclosed in this application.
[0064] Gesture recognition is a technology that enables interaction by analyzing human hand gestures. It converts human gestures into computer-understandable instructions, thus enabling interaction with computers or other devices. Gesture recognition technology is applied in multiple fields, including human-computer interaction, virtual reality, augmented reality, smart homes, and healthcare, offering advantages such as naturalness, convenience, and intuitiveness in interaction design.
[0065] In existing solutions, large screens primarily use touchscreens and remote controls for interaction. Touchscreens offer precise positioning but require close-range operation, and the large screen size makes the user experience less than ideal. Remote controls support interaction from medium to long distances but generally do not support actual cursor positioning. Remote controls require additional hardware, increasing sales costs and incurring additional maintenance costs (charging, preventing loss, etc.), especially for public facilities.
[0066] Therefore, embodiments of this application provide a screen interaction method, apparatus, device, and computer-readable storage medium (hereinafter referred to as storage medium). By acquiring a target image of a target screen for remote control, analyzing the key points of the controller's hand, determining a virtual ray between the hand and the target screen based on the three-dimensional position information of the key points, and further confirming the target interaction position based on the pose information of the virtual ray and the target screen, remote interactive control is achieved, improving the flexibility of remote interactive control, avoiding the increased cost of using remote control devices, and enhancing intelligent interactive performance. These will be described in detail below.
[0067] The screen interaction method in this embodiment of the invention is applied to a screen interaction device, which is set in a display device. The display device is provided with one or more processors, a memory, and one or more applications, wherein one or more applications are stored in the memory and configured to be executed by the processor to implement the screen interaction method. The display device can be a terminal, such as a mobile phone or a tablet computer, or it can be a server or a service cluster composed of multiple servers.
[0068] like Figure 1 As shown, Figure 1 This is a schematic diagram of a screen interaction method according to an embodiment of the present application. The screen interaction scenario in this embodiment includes a display device 100 (the display device 100 integrates a screen interaction device), and a computer-readable storage medium corresponding to the screen interaction is run in the display device 100 to perform the screen interaction steps.
[0069] Understandable, Figure 1 The display device in the scenario of the screen interaction method shown, or the device included in the display device, does not constitute a limitation on the embodiments of the present invention. That is, the number or type of device included in the scenario of the screen interaction method, or the number or type of device included in each device, does not affect the overall implementation of the technical solution in the embodiments of the present invention, and can all be considered as equivalent substitutions or derivatives of the technical solutions claimed in the embodiments of the present invention.
[0070] In this embodiment of the invention, the display device 100 is mainly used for: acquiring key hand points of a controller who controls a target screen, wherein the key hand points are determined by user positioning and skeletal tracking of a target image containing the controller; determining at least two target key points based on the three-dimensional position information of each key hand point, wherein a virtual ray formed by the three-dimensional position information of the at least two target key points points points from the controller's hand to the target screen; determining the target interaction position of the target screen based on the three-dimensional position information of the target key points and the pose information of the target screen, and performing interactive control on the target screen.
[0071] In this embodiment of the invention, the display device 100 can be an independent display device, or a network of display devices or a cluster of display devices. For example, the display device 100 described in this embodiment of the invention includes, but is not limited to, a computer, a network host, a single network display device, a set of multiple network display devices, or a cloud display device composed of multiple display devices. The cloud display device is composed of a large number of computers or network display devices based on cloud computing.
[0072] Those skilled in the art will understand that Figure 1The application environment shown is merely one application scenario of the solution in this application and does not constitute a limitation on the application scenario of the solution in this application. Other application environments may include those that are more specific to this application. Figure 1 The number of display devices shown, or the network connectivity of display devices, for example... Figure 1 Only one display device is shown in the diagram. It is understood that the scenario of this screen interaction method may also include one or more other display devices, which are not specifically limited here. The display device 100 may also include a memory for storing data, such as storing image information acquired by shooting.
[0073] Furthermore, in the screen interaction method scenario of this application, the display device 100 sets the display device 200 as the target screen, or the display device 100 does not have a communication connection with an external display device. The display device 200 is used to output the result of the screen interaction method executed in the display device. The display device 100 can access the background database 300 (the background database can be in the local storage of the display device, or it can be set in the cloud), and the background database 300 stores information related to screen interaction.
[0074] It should be noted that, Figure 1 The schematic diagram of the screen interaction method shown is merely an example. The scenarios of the screen interaction method described in the embodiments of the present invention are intended to more clearly illustrate the technical solutions of the embodiments of the present invention and do not constitute a limitation on the technical solutions provided in the embodiments of the present invention.
[0075] Based on the scenarios described above for screen interaction methods, embodiments of screen interaction methods are proposed.
[0076] like Figure 2 The diagram shown is a flowchart of an embodiment of the screen interaction method in this application. The screen interaction method includes steps S201-S203:
[0077] S201. Obtain the key points of the hands of the master controller who controls the target screen.
[0078] Among them, the key points of the hand are determined by user localization and skeletal tracking of the target image containing the main controller.
[0079] The target image is an image of the environment corresponding to the target screen, and the target image includes at least one controller who controls the target screen.
[0080] It is understood that the target image can be captured by a target shooting device set on the target screen itself, or by a shooting device that transmits data with the target screen. Specifically, this application does not specifically limit the source of the target image.
[0081] In this scenario, the controller is the user who interacts with the target screen through gestures, and the target screen adjusts its display based on the controller's gesture changes.
[0082] Among them, the key points of the hand are the key points used to characterize the position of each hand limb of the controller. For example, they include key points corresponding to each hand limb such as the thumb, index finger, middle finger, ring finger and little finger. The key points of a hand limb can include: fingertip key points, wrist key points, interphalangeal key points, palm and finger key points, etc.
[0083] In some embodiments of this application, hand key points may include identification information, which is used to identify which hand key point is represented by the corresponding hand key point, such as representing a fingertip key point, a wrist key point, etc.
[0084] In some other embodiments of this application, there is a hand constraint relationship between hand key points. The hand constraint relationship can be a mapping relationship between hand key points and a preset hand key point model or a set of hand key point positions. For example, the hand constraint relationship is that the index finger and the middle finger are set to be adjacent.
[0085] User location, which is the identification of the user's location in the target image, can be understood as being achieved through methods such as user facial recognition and human body recognition.
[0086] The skeletal tracking representation locates and tracks the human skeleton of each user to identify skeletal changes and determine the controller and key points of the controller's hands based on these changes. It is understood that skeletal tracking can be implemented through a model. For example, user images are input into a preset recognition model for skeletal prediction. The type of preset recognition model is not limited; for example, it could be a Long Short-Term Memory Network (LSTM) or a Spatiotemporal Convolutional Neural Network (3D-CNN) to determine the key points of the human skeleton corresponding to the target image, and thus determine the key points of the controller's hands. It is understood that the preset recognition model can be trained to output key points of the hands, i.e., it can perform skeletal tracking only on the controller's hands to obtain the key points of the hands.
[0087] In one embodiment of this application, the screen interaction method is applied to a display device, which includes a target screen, a corresponding shooting device, and a processor. For example, the display device can be a smart TV, a conference large screen device, etc. In the application scenario, the user can trigger screen remote control and send it to the processor through the screen interaction controls displayed on the target screen or the screen interaction controls displayed on the mobile device connected to the target screen. The processor responds to the screen remote control operation, controls the shooting device to start, and performs target image acquisition. After obtaining the target image, the processor inputs the target image into a preset recognition model for user positioning and skeleton tracking, and then outputs the key points of the controller's hand.
[0088] S202. Based on the three-dimensional position information of each of the hand key points, at least two target key points are determined, and a virtual ray formed by the three-dimensional position information of the at least two target key points points points from the hand of the controller to the target screen.
[0089] Among them, the three-dimensional position information, that is, the position coordinates of the hand key points in three-dimensional space, can be understood as the three-dimensional space being, but not limited to, the world coordinate system corresponding to the camera that acquires the target image, or a preset three-dimensional coordinate system that has a positional mapping relationship with the hand key points. It can be understood that, in the implementation scheme of this application, the pose information in the three-dimensional space is pre-existing in the target screen, and the pose information of the target screen in the three-dimensional space can be obtained through camera calibration.
[0090] The virtual ray is represented as pointing from the controller's hand to the target screen; it is understood that the virtual ray does not actually exist.
[0091] Specifically, in the implementation scheme of this application, after obtaining the key points of the hand, the key points of the hand are mapped to three-dimensional space to obtain the three-dimensional position information corresponding to the key points of the hand in three-dimensional space. Furthermore, based on the relative positional relationship between the three-dimensional position information, the target key points corresponding to the hand limbs are determined, thereby determining the spatial orientation of the hand limbs.
[0092] It is understood that the hand includes multiple hand limbs, that is, the controller's hand may include different directions corresponding to different hand limbs. In the implementation scheme of this application, the direction corresponding to each hand limb is determined, and combined with the posture information of the target screen in three-dimensional space, the target hand key point corresponding to the target direction pointing to the target screen is determined.
[0093] The hand limbs include, but are not limited to: the thumb, index finger, middle finger, ring finger, little finger, palm, and wrist to arm; furthermore, each hand limb corresponds to at least two key points of the hand. For example, the target key points of the index finger include: the key point of the index fingertip, the key point between the index finger bones, and the key point of the index finger palm.
[0094] Specifically, in the embodiments of this application, the direction of the hand limb is determined based on the relative positional relationship between the three-dimensional positional information of the key points of the same hand limb; for example, the direction of the index finger is determined based on the relative positional relationship between the three-dimensional positional information of the fingertip key point and the three-dimensional positional information of the interphalangeal key point among the key points of the hand representing the index finger.
[0095] Furthermore, after determining the direction corresponding to each hand limb, the key hand points of the target hand limb pointing to the target screen are determined based on the pose information of the target screen in three-dimensional space. The pose information of the target screen and the three-dimensional position information of the key hand points are located in the same three-dimensional space.
[0096] Furthermore, after determining the target hand limb, target hand key point information is determined based on the hand key points of the target hand limb. For example, all hand key points of the target hand limb can be set as target key points. Alternatively, in some other embodiments of this application, in order to reduce the amount of data and improve data processing efficiency, at least two hand key points that can characterize the direction of the target hand limb can be extracted as target key points. This application does not limit the specific details.
[0097] Understandably, the specific implementation of determining at least two target key points may include: determining the curvature of the target hand limb based on the relative positional relationship between the key points representing the target hand limb; determining the target key points based on the magnitude relationship between the curvature of the target hand limb and a preset curvature. For example, if the curvature of the hand limb is greater than the preset curvature, it indicates that the target hand limb is not straight, and the fingertip key point and the interphalangeal key point adjacent to the fingertip key point are extracted as target key points. If the curvature of the hand limb is less than or equal to the preset curvature, it indicates that the target hand limb is straight, and at least two hand key points are extracted as target key points.
[0098] It is understandable that the three-dimensional information corresponding to at least two target key points can represent a directional straight line, and thus the virtual ray formed by the three-dimensional information corresponding to at least two target key points points points from the controller's hand to the target screen.
[0099] S203. Based on the three-dimensional position information of the target key points and the pose information of the target screen, determine the target interaction position of the target screen and perform interactive control on the target screen.
[0100] Among them, the pose information of the target screen, that is, the three-dimensional position information of the target screen at the key points of the hand, corresponds to the position information in three-dimensional space, including position and orientation.
[0101] Specifically, in the implementation scheme of this application, after obtaining the target key points, the relative distance between the target key points and the target screen is determined based on the pose information of the target key points and the target screen. Further, based on the relative distance, the three-dimensional intersection point between the virtual rays corresponding to at least two target key points and the target screen is determined. Further, based on the preset mapping relationship between the screen coordinates corresponding to the target screen and the spatial position in three-dimensional space, the three-dimensional intersection point is converted into the screen coordinate point of the target screen as the target interaction position. Further, interactive control is performed based on the screen interaction position.
[0102] Specifically, this application does not limit the specific implementation of interacting with the target screen, and may include cursor position control, screen focus switching control, etc.
[0103] This application's implementation scheme analyzes the target image to obtain key hand points of the controller corresponding to the target screen. Furthermore, based on the three-dimensional position information of these key hand points, it determines the target key points corresponding to the virtual ray representing the hand pointing towards the target screen from a three-dimensional spatial perspective. This ensures accurate determination of the target key points, avoiding limitations of two-dimensional space. Further, combined with the pose information of the target screen, it determines the target interaction position of the target screen, thereby enabling interactive control of the target screen. This ensures the flexibility of the controller's position movement during interaction, eliminates the need for additional interactive devices such as a mouse, reduces equipment configuration costs, and ensures accuracy by analyzing the hand's pointing direction in three-dimensional space to determine the target interaction position, thus improving interaction performance.
[0104] For example, in one embodiment of this application, interactive control of the target screen specifically includes the following steps:
[0105] (1) If a cursor control exists in the target screen, control the cursor control to move to the target interactive position;
[0106] (2) In response to a trigger operation at the target interaction location, interact with the target screen.
[0107] Specifically, in the implementation scheme of this application, if a cursor control exists on the target screen, it means that the target screen achieves interaction through the cursor control. After determining the target interaction position, the target interaction position is determined as the target position of the cursor. That is, a cursor movement command is generated according to the target interaction position to control the cursor control to move to the target interaction position. In response to the trigger operation of the target interaction position, the cursor interacts with the target screen to realize the interaction of gesture control of cursor movement, improve the flexibility of interaction control, and reduce the cost of interaction devices.
[0108] In some other embodiments of this application, if the interaction in the current display interface of the target screen is a focus switching control, then the target interaction position is taken as the target focus, and the target focus is sent to the focus switching control module of the processor. The focus switching control module performs focus rendering on the display control corresponding to the target interaction position so that the display control corresponding to the target interaction position is highlighted, thereby realizing focus switching.
[0109] Furthermore, based on the above implementation scheme, this application also provides a specific implementation scheme for determining key hand points, see [link to relevant documentation]. Figure 3 , Figure 3 This is a flowchart illustrating one implementation of the hand key point determination method in the screen interaction method provided in this application, specifically including steps S301-S304:
[0110] S301. Obtain the target image for controlling the target screen.
[0111] Specifically, in this embodiment of the application, a target image is acquired by a camera installed on the target screen, and the target image includes at least one user.
[0112] S302. Perform human detection on the target image using a human detection model to determine the main controller in the target image and the hand area of the main controller.
[0113] Specifically, in the implementation scheme of this application, human detection is performed on the target image using a human detection model. This means that the human detection model uses a pre-trained deep learning model to detect the location regions of all human bodies and the location regions of each person's hands from the target image of the video stream. Here, the location region is defined as a rectangular area on the target image, generally represented using the top-left corner + width and height, top-left corner + top-right corner, or center point + width and height. This application does not require a specific representation.
[0114] S303. The key points of the hand region are detected by the hand key point detection model to determine the two-dimensional position information of the hand joints of the controller and the confidence level of the two-dimensional position information.
[0115] Among them, the confidence level of the two-dimensional position information is used to evaluate the accuracy of the hand joints. It is understandable that due to the occlusion of the hand in the image and the incompleteness of the image information, the accuracy of the two-dimensional position information of the hand joints may be biased. Therefore, the corresponding confidence level is output as auxiliary information to determine the key points of the hand.
[0116] Specifically, in the implementation scheme of this application, the hand keypoint model uses a pre-trained deep learning model to extract hand keypoints from the hand location region in the original video stream based on the results of the human detection model. Here, hand keypoints refer to the positional information of the corresponding joints of the hand, typically consisting of 21 or 26 points, with each keypoint corresponding to a fixed joint. The positional information of the keypoints can be 2D coordinates or may include relative / estimated depth information (i.e., 2.5D coordinates).
[0117] It should be noted that this application does not restrict the deep learning model. It can be the classic OpenPose network (OpenPose is a deep learning-based multi-person pose estimation system designed to detect and track key points of the human body from images or videos. Its core is a deep neural network capable of end-to-end pose estimation), the lightweight AlphaPose (AlphaPose is a deep learning-based multi-person pose estimation system designed to detect and track key points of the human body from images or videos, similar to OpenPose), the high-precision HRNet (HRNet: High-Resolution Network, a deep learning network for human pose estimation that offers better performance while preserving higher resolution information compared to traditional pose estimation methods), and other similar networks. The network must contain at least the 2D coordinate information of all or some of the key points, and may also provide additional information such as confidence level, relative / estimated depth, relative distance between finger joints, and predicted values of finger joint rotation angles.
[0118] That is, it can be understood that, in the implementation scheme of this application, the two-dimensional position information of the hand joints of the controller determined by the hand key point detection model can be 2D coordinate information or 2.5D coordinate information.
[0119] Understandably, this application does not impose restrictions on the deep learning model. It can be a single model simultaneously detecting the person and corresponding hand region (one-stage), or multiple models detecting the person and hand region separately (multi-stage). In multi-stage models, hand region detection can use the entire image input or be performed individually for each person based on the detected human body region. There are also no restrictions on the network structure of the deep learning model. It can be a lightweight YOLO ("YOLO" is a popular object detection algorithm, short for "You Only Look Once." It is a real-time object detection algorithm that can quickly and accurately detect multiple objects in an image or video, providing bounding boxes and class labels for each object.) series of networks, or EfficientDet (EfficientDet is an object detection model based on the EfficientNet architecture, designed to achieve higher efficiency and accuracy in object detection tasks. It employs a series of innovative technologies, including network structure design, Feature Pyramid Network (FPN), and Bidirectional Feature Network (BiFPN), to achieve excellent performance in object detection tasks), or Transformer-based ViTDet or DETR, etc.
[0120] ViTDet is an object detection model based on the Vision Transformer (ViT), applying the Transformer architecture to object detection tasks. Unlike traditional Convolutional Neural Networks (CNNs), ViTDet uses a global self-attention mechanism to capture global contextual information in images. ViTDet uses several variations to adapt to object detection tasks; for example, it adds positional encoding information to the input image to preserve the spatial structure of the image. By introducing object detection-related designs into the Transformer architecture, ViTDet achieves performance comparable to or even better than traditional object detection models on some datasets. DETR is an end-to-end object detection model that does not require traditional anchor boxes, candidate boxes, etc., but directly outputs the objects in the image and their locations through the Transformer architecture. DETR uses an encoder-decoder architecture, where the encoder extracts features from the input image, and the decoder predicts the object's category and location. DETR uses a self-attention mechanism to globally model the features of the input image, thus effectively detecting objects at different scales and resolutions. DETR demonstrates high performance on some commonly used object detection datasets and has end-to-end advantages, making the training and inference processes simpler and more efficient.
[0121] S304. Based on the two-dimensional position information and the confidence level corresponding to each two-dimensional position information, determine the key points of the hand.
[0122] Furthermore, in the embodiments of this application, the joint information can be further corrected by the model based on the confidence level of the two-dimensional position information and the hand constraint relationship between the two-dimensional position information, thereby determining the corrected key points of the hand.
[0123] In this embodiment, the hand region and key points are detected through a model, which improves the efficiency of key point detection. At the same time, the key points of the hand are determined based on the confidence level corresponding to the two-dimensional position information, avoiding the limitations of image key point detection and ensuring the rationality and accuracy of the key points of the hand.
[0124] Furthermore, based on any of the above implementation schemes, this application also provides an implementation scheme for determining key target points, see [link to relevant documentation]. Figure 4 , Figure 4 A flowchart illustrating one implementation of the screen interaction method for determining target key points provided in this application, specifically including steps S401-S403:
[0125] S401. The three-dimensional position information of the key hand points is predicted by using the hand joint model to obtain the three-dimensional position information of the key hand points.
[0126] The hand joint model is constructed based on kinematic structure, including constraint information between the three-dimensional hand parts. The type of joint model is not limited; it can be a parametric template model (e.g., MANO, Model for Articulated Hands with Object Interaction, a model for hand pose estimation. It is a deep learning-based method designed to accurately estimate the three-dimensional pose and gestures of the hand), or an optimized set of constraints for finger length learned based on statistical information. For an example, see [link to example]. Figure 5 , Figure 5 This diagram illustrates the constraint information between the three-dimensional hand joints in the hand joint model provided for the implementation of this application. The constraint information includes 21 key hand points, forming a tree structure with the root node at the bottom and child nodes at the top. The wrist key point is the root node. First-level child nodes are the five palmodigital key points connected to the wrist key point. Second-level child nodes are the points corresponding to the lower-middle joints of the fingers closest to the fingertip key point, connected to each palmodigital key point. Fourth-level nodes are the fingertip key points of each finger (not shown in the diagram, representing the fingertip tip). Third-level child nodes are the points corresponding to the upper-middle joints of the fingers located between the palmodigital key points and the fingertip key points, connected to each palmodigital key point. Bending the parent node affects the 3D position of the child nodes. Due to the limitations of human anatomy, the direction and range of bending for each node are restricted, thus reducing the degrees of freedom.
[0127] In this scheme, each point on the fingers of the hand joint model has 1-2 degrees of freedom in bending along one or two axes, and the wrist itself contains 6 degrees of freedom (position in three dimensions and rotation in three dimensions), thus forming constraint information between the three-dimensional hand parts. If there are no constraints on the 3D points, 21*3 variables need to be solved. In this implementation scheme, after kinematic model constraints, with the finger joint length fixed, only 23 variables need to be solved, significantly reducing the difficulty of solving and the number of degrees of freedom. Therefore, the process of predicting the three-dimensional information of the key hand points by using the hand joint model to obtain the three-dimensional information of the key hand points can reduce the amount of data processing while improving the calculation accuracy of the three-dimensional position information.
[0128] In this embodiment, a hand keypoint detection model is used to detect keypoints in the hand region, outputting two-dimensional position information and confidence levels for each keypoint. In other embodiments, the hand keypoint detection model also outputs relative / estimated depth, relative distance between finger joints, and predicted values of finger joint rotation angles. Further, based on the preliminary results output by the network (two-dimensional position information, and the confidence levels of the two-dimensional position information, or including relative / estimated depth, relative distance between finger joints, and predicted values of finger joint rotation angles), a hand joint model is used to further solve for a set of hand 3D keypoint positions that are most consistent with (with the smallest error) the actual output values from all viewpoints (this process is the step mentioned in the above embodiment of determining hand keypoints based on the two-dimensional position information and the confidence levels corresponding to each two-dimensional position information). This position is the final determined three-dimensional position information of the hand keypoints, which also satisfies some constraints on hand length (based on statistical information from offline data).
[0129] S402. Based on the three-dimensional position information of each hand key point and the relative positional relationship between each hand key point, determine the hand key points of the same hand limb.
[0130] Specifically, after determining the three-dimensional position information of each hand key point, the hand key points representing the same hand limb are determined according to the relative position constraints between the three-dimensional information. Then, the hand key points representing the same hand limb are formed as a group of hand key points for the purpose of calculating the direction of the hand limb.
[0131] For example, the prediction using the hand joint model can directly predict the three-dimensional position information of the hand key points and divide the hand key points based on the constraint information carried inside, thereby determining the hand key points of the same hand limb.
[0132] S403. Based on the hand key points of each hand limb and the pose information of the target screen, determine the target hand key points of the hand limb pointing to the target screen.
[0133] Specifically, after determining the key hand points of the same hand limb, the direction of the hand limb is determined based on the relative positional relationship between the three-dimensional positional information corresponding to the key hand point group of the same hand limb. Furthermore, the target hand key points of the hand limb pointing to the target screen are determined based on the pose information of the target screen in three-dimensional space.
[0134] It is understandable that if there are at least two hand limbs pointing to the target screen, a preset hand limb is selected from the hand limbs as the target hand limb, and the key points of the target hand are determined based on the key points of the target hand limb.
[0135] For example, for each group of key points on the hand, the direction of the corresponding hand limb is determined by calculating the angular offset vector between the key points on the hand in the corresponding coordinate space. If the index finger and middle finger of the hand limb both point to the screen, but the preset rule sets the index finger to be used to determine the virtual ray, then the index finger is determined as the target hand limb.
[0136] This embodiment uses a hand joint model to predict the three-dimensional information of key hand points, thereby reducing the amount of data processing and improving the accuracy of the three-dimensional position information calculation. At the same time, it determines the direction of each hand limb to ensure the accuracy of determining the target hand key point of the hand limb pointing to the target screen.
[0137] Furthermore, based on any of the above embodiments, this application also provides a specific embodiment for determining target key points, wherein determining the target hand key points of the hand limb pointing to the target screen based on the hand key points of each hand limb and the pose information of the target screen specifically includes the following steps:
[0138] (1) Based on the three-dimensional position information of each two adjacent target hand key points on the same hand limb, the bending angle of each limb segment is calculated;
[0139] (2) Calculate the bending angle of each segment in each hand limb to obtain the bending degree of the hand limb;
[0140] (3) Obtain the target hand limb with a curvature smaller than the preset curvature;
[0141] (4) Based on the three-dimensional position information of the target hand limb and the pose information of the target screen, determine the target hand key points pointing to the target hand limb of the target screen.
[0142] The curvature includes the curvature of each hand limb, for example, the curvature of the thumb, index finger, middle finger, ring finger, and little finger.
[0143] In one embodiment of this application, the curvature can be determined by calculating the relative offset between the three-dimensional position information corresponding to the key points of the hand using a preset offset calculation formula. For example, the relative offset can be the angular offset, position offset, etc. between the key points of the hand.
[0144] In other embodiments of this application, the curvature can also be calculated by a preset model. For example, the three-dimensional position information corresponding to the key points of the hand is input into the preset model, and the curvature of the hand limb is output through model analysis. It is understood that the preset model can be obtained through training.
[0145] It is understood that in some embodiments of this application, if the scenario is limited to a target hand limb pointing to the target interaction position of the target screen, it is not necessary to calculate the curvature of all hand limbs. Only the curvature corresponding to a specific target hand limb needs to be calculated. If the curvature is greater than the preset curvature, it means that the gesture is invalid, and the step of determining the target hand key point pointing to the target screen based on the three-dimensional position information of the target hand limb and the pose information of the target screen is not executed.
[0146] Similarly, in a scenario where the target hand limb is defined as the part used to determine the target interaction position pointing to the target screen, the hand key points in any of the above implementation schemes may only include the hand key points corresponding to the target hand limb. For example, if the index finger is defined as the target hand limb, then only the hand key points corresponding to the index finger need to be obtained for processing. For the specific implementation process, please refer to any of the implementation schemes, which will not be described in detail here.
[0147] Specifically, this application excludes bent hand limbs through curvature calculation. That is, during the virtual ray calculation of the hand, if all hand limbs are in a bent state, then no ray is generated for that hand limb. The purpose is to avoid the judgment calculation of unnecessary hand limbs, reduce the amount of data calculation, and improve the accuracy of pointing judgment.
[0148] Furthermore, based on any of the above implementation schemes, this application also provides a specific implementation scheme for determining the target interaction location, see [link to relevant documentation]. Figure 6 , Figure 6 A flowchart illustrating one implementation of the screen interaction method for determining the target interaction position provided in this application is shown, specifically including steps S601-S602:
[0149] S601. Based on the three-dimensional position information of the target key points, determine the main control key points and auxiliary key points.
[0150] Specifically, in the implementation scheme of this application, the primary control key point is the fingertip key point, and the auxiliary key points are at least one of the wrist key point, interphalangeal key point, and palmofinar key point.
[0151] Specifically, auxiliary key points can be determined based on the number of hand key points included in the target key points. For example, if the target key points include any one of wrist key points, interphalangeal key points, or palmofinite key points, then any one of the included wrist key points, interphalangeal key points, or palmofinite key points is set as an auxiliary key point. If the target key points include at least two of the included wrist key points, interphalangeal key points, or palmofinite key points, then auxiliary key points are determined based on the pointing offset between at least two of the included wrist key points, interphalangeal key points, or palmofinite key points and the main control key point. For example, if... The target key points include wrist key points and interphalangeal key points. A first direction between the wrist key point and the main control key point, and a second direction between the interphalangeal key point and the main control key point are calculated. If the relative distance between the first and second directions is less than a preset relative distance, any point between the wrist key point and the interphalangeal key point is set as an auxiliary key point. If the relative distance between the first and second directions is greater than or equal to the preset relative distance, the midpoint between the wrist key point and the interphalangeal key point is set as an auxiliary key point. This application does not specifically limit the specific implementation method.
[0152] S602. Determine the target interaction position of the target screen based on the pose information of the target screen and the virtual ray formed by the main control key point and the auxiliary key point.
[0153] Specifically, in some embodiments of this application, after determining the main control key points and auxiliary key points, the relative distance between the main control key points or auxiliary key points and the pose information of the target screen is calculated. Then, based on the relative distance, the main control key points, and the auxiliary key points, the three-dimensional intersection point of the virtual ray formed by the main control key points and the auxiliary key points on the target screen is determined. Finally, the target interaction position of the target screen is determined based on the three-dimensional intersection point. It is understood that this process can be implemented by a triangle principle algorithm or by a model.
[0154] In this implementation plan, the main control key points and auxiliary key points are determined based on the three-dimensional position information of the target key points, thereby improving the accuracy of the pointing calculation and avoiding the limitations caused by directly using the target's hand key points to calculate the target's interactive position.
[0155] Furthermore, this application also provides one embodiment for determining the pose of a target screen. Specifically, the determination of the pose information includes the following steps:
[0156] (1) Acquire a reference image and a calibration image, wherein the target image and the reference image include a reference object, the reference image is acquired by a target shooting device that acquires the target image, and the calibration image is acquired by an auxiliary shooting device that captures the target screen, and the target shooting device and the auxiliary shooting device have a common viewing area;
[0157] (2) Determine the first pose information of the reference object based on the image coordinates of the reference object in the reference image;
[0158] (3) Determine the second pose information of the reference object based on the image coordinates of the reference object in the calibration image.
[0159] (4) Determine the pose information of the target screen based on the relative pose information between the first pose information and the second pose information.
[0160] See Figure 7 , Figure 7 This is a schematic flowchart of one implementation scheme for pose information calibration in the target screen interaction method provided in this application. Specifically, after calibration begins, the target shooting device C0 and the auxiliary shooting device C1 jointly shoot the external reference object t and calculate the pose of the reference object within the target shooting device. Position within the auxiliary shooting device Further calculations were performed on the pose of the auxiliary imaging device within the target imaging device. Based on the calculated pose of the target screen marker within the auxiliary shooting device Then calculate the pose of the target imaging device C0 relative to the target screen marker. The final calibration is complete.
[0161] Specifically, in the embodiments of this application, it is necessary to rely on the 3D pose information of the target screen in the coordinate system of the shooting device corresponding to the target image acquisition device, that is, the accurate positional relationship between the target screen and the target shooting device. In some embodiments of this application, this parameter can be completed through a single calibration at the factory. In this solution, since the target shooting device C0 and the target screen s are parallel to each other and do not share a view, an additional auxiliary shooting device C1 is required for calibration.
[0162] The calibration method requires ensuring that the auxiliary imaging device C1 is positioned to see the target screen and shares a common field of view with the target imaging device C0. This common field of view means that the viewing angles of the two imaging devices overlap within a certain range, allowing them to see the same object (reference point). There are no strict requirements regarding the placement of the imaging devices; they can be placed opposite each other (each seeing the other) or sideways (not appearing in the line of sight of the other).
[0163] In addition to the auxiliary shooting devices, a reference object t needs to be placed in the common viewing area of the shooting devices to ensure that all shooting devices can see the reference object, and to calculate the 3D pose of the reference object (position and rotation relative to the shooting device) in each shooting device. Here, 'i' represents any one of the shooting devices. The reference object can be a tag board with a QR code, such as AprilTag or RuneTag, or an object with a unique pattern and shape. It can be used as a reference object as long as its 3D pose can be determined from a single image taken by all the shooting devices. By combining the poses of the reference object within the two shooting devices, the relative pose relationship between the two shooting devices can be obtained, i.e., the pose of the auxiliary shooting device within the target shooting device.
[0164] After obtaining the relative poses of the two imaging devices, target screen markers (calibration patterns, such as Apriltag) can be displayed on the target screen s to calculate the pose of the target screen within the auxiliary imaging device. Finally, the position of the target screen marker in the target imaging device can be obtained. This pose is the pose of the target screen s in the C0 coordinate system of the target shooting device. Once the calibration is complete, the pose information of the target screen is obtained.
[0165] This application does not limit the number of target imaging devices and auxiliary imaging devices. For multiple target imaging devices, the calibration method is the same as for a single target imaging device, and it can improve the accuracy of gesture calculation. For multiple auxiliary imaging devices, the calibration accuracy can be improved, while also reducing manual measurement of the calibration object's dimensions, thus increasing calibration efficiency and automation.
[0166] Furthermore, based on the above implementation scheme, this application also provides an implementation scheme for a screen interaction method, as detailed below. Figure 8 , Figure 8The flowchart of another implementation scheme of the screen interaction method provided in this application is as follows: In the screen interaction method, a human detection model is used to detect all people (users) in the scene corresponding to the target image, obtaining the position areas of all people in the scene and the position areas of each person's hands. If no master controller is currently selected, the master controller detection process is entered; if a master controller already exists, the step of obtaining the key hand points of the master controller who controls the target screen is executed, and the master controller tracking process is entered, which only requires key hand point detection and gesture recognition of the master controller. That is, if a master controller is detected, gesture and screen cursor calculation are performed, and master controller information is recorded (based on the three-dimensional position information of each of the key hand points, at least two target key points are determined, and the virtual ray formed by the three-dimensional position information of at least two target key points points points from the master controller's hand to the target screen). If a gesture action is triggered, the target interaction position of the target screen is determined based on the three-dimensional position information of the target key points and the pose information of the target screen, and the target screen is interactively controlled. The gesture triggering refers to the determination of the relationship between the curvature of the hand limb and the preset curvature in the above implementation scheme.
[0167] Specifically, the master controller detection process includes: treating each user as a candidate master controller, using a hand keypoint detection model to detect hand keypoints in the hand area of each candidate master controller, and sequentially calling a hand keypoint deep learning module to recognize hand movements based on the 3D position information of each hand keypoint for both hands of each candidate master controller, thereby calculating hand gestures. If the calculated gesture matches a predefined master controller gesture, then that candidate is selected as the master controller. In subsequent algorithm processes, only the interaction actions of this master controller will be processed, ignoring the interaction actions of other people (if any) in the video frame.
[0168] Currently, the commercial display industry generally does not incorporate gesture control technology. Supporting gesture control via built-in or external cameras will provide a significant competitive advantage without incurring excessive additional costs. The screen interaction method provided in this application acquires key hand points of the user controlling the target screen. These key hand points are determined through user localization and skeletal tracking of a target image containing the user. Based on the three-dimensional position information of each key hand point, at least two target key points are determined. A virtual ray formed by the three-dimensional position information of these at least two target key points points points from the user's hand towards the target screen. Based on the three-dimensional position information of the target key points and the pose information of the target screen, the target interaction position of the target screen is determined, enabling interactive control of the target screen. The algorithm of this solution is based solely on a monocular RGB camera. Compared to other products on the market that require infrared, depth, and other sensors, this sensor is simpler and lower in cost, enhancing product competitiveness and achieving a leading domestic technological level. Furthermore, this solution proposes a bare-hand gesture ray interaction method, allowing users to simulate a laser pointer and interact with a large screen using their bare hands. This reduces the need for additional hardware components and enables pointing capabilities in multi-person discussion scenarios, improving the efficiency of communication. Further, this solution proposes a method for calibrating the relative 3D pose (position and rotation) between the large screen camera and the screen, solving the calibration problem under conditions without shared vision and providing crucial support for accurately calculating the intersection point of the user's hand ray and the screen. Moreover, this solution uses 3D keypoint information to calculate the gesture ray direction, combined with the camera and screen positional relationship calibrated at the device's factory, to provide accurate finger pointing direction, offering higher accuracy compared to general 2D methods. This interaction method is consistent with traditional pointing methods, aligning with user intuition, reducing the user's learning cost, and improving intelligent interaction performance.
[0169] To better implement the screen interaction method in the embodiments of this application, a screen interaction device is also provided in the embodiments of this application, such as... Figure 9 As shown, the screen interaction device includes modules 901-903:
[0170] The acquisition module 901 is used to acquire the key hand points of the master controller controlling the target screen. The key hand points are determined by user positioning and skeleton tracking of the target image containing the master controller.
[0171] The determining module 902 is used to determine at least two target key points based on the three-dimensional position information of each of the hand key points, and the virtual ray formed by the three-dimensional position information of the at least two target key points points points from the hand of the master controller to the target screen.
[0172] The interactive control module 903 is used to determine the target interactive position of the target screen based on the three-dimensional position information of the target key points and the pose information of the target screen, and to perform interactive control on the target screen.
[0173] In one embodiment of this application, before the acquisition module 901 acquires the key points of the hand of the controller controlling the target screen, it further includes a method for:
[0174] Acquire the target image for controlling the target screen;
[0175] Human detection is performed on the target image using a human detection model to identify the main controller in the target image and the hand area of the main controller;
[0176] The key points of the hand region are detected by the hand key point detection model to determine the two-dimensional position information of the hand joints of the controller, as well as the confidence level of the two-dimensional position information;
[0177] Based on the two-dimensional position information and the confidence level corresponding to each two-dimensional position information, key points of the hand are determined.
[0178] This solution uses a model to detect the hand region and key points, improving the efficiency of key point detection. At the same time, it determines the key points of the hand based on the confidence level corresponding to the two-dimensional position information, avoiding the limitations of image key point detection and ensuring the rationality and accuracy of the key points of the hand.
[0179] In one embodiment of this application, the determining module is configured to determine at least two target key points based on the three-dimensional position information of each of the hand key points, including:
[0180] The three-dimensional position information of the key hand points is obtained by predicting the three-dimensional position information of the key hand points using a hand joint model;
[0181] Based on the three-dimensional position information of each hand key point and the relative positional relationship between each hand key point, the hand key points of the same hand limb are determined;
[0182] Based on the hand key points of each hand limb and the pose information of the target screen, the target hand key points of the hand limb pointing to the target screen are determined.
[0183] This solution uses a hand joint model to predict the three-dimensional position information of the two-dimensional hand key points. The process of obtaining the three-dimensional position information of the hand key points can improve the calculation accuracy of the three-dimensional position information. At the same time, it determines the direction of each hand limb to ensure the accuracy of determining the target hand key points of the hand limbs pointing to the target screen.
[0184] In one embodiment of this application, the determining module is configured to determine the target hand key points of the hand limb pointing to the target screen based on the hand key points of each hand limb and the pose information of the target screen, including:
[0185] The bending angle of each segment is calculated based on the three-dimensional position information of every two adjacent key points of the same hand limb.
[0186] The bending angle of each segment in each hand limb is counted to obtain the bending degree of the hand limb;
[0187] Acquire the target hand limb with a curvature smaller than a preset curvature;
[0188] Based on the three-dimensional position information of the target hand limb and the pose information of the target screen, the key points of the target hand limb pointing to the target screen are determined.
[0189] In this scheme, during the calculation of virtual rays for the hand, if all hand limbs are in a bent state, no ray is generated for that hand. The purpose is to avoid the judgment and calculation of unnecessary hand limbs, reduce the amount of data calculation, and improve the accuracy of pointing judgment.
[0190] In one embodiment of this application, the interaction control module 903 is configured to determine the target interaction position of the target screen based on the three-dimensional position information of the target key points and the pose information of the target screen, including:
[0191] Based on the three-dimensional position information of the target key points, the main control key points and auxiliary key points are determined. The main control key point is the fingertip key point, and the auxiliary key points are at least one of the wrist key point, interphalangeal key point, and palmofinite key point.
[0192] Based on the pose information of the target screen and the virtual ray formed by the main control key point and the auxiliary key point, the target interaction position of the target screen is determined.
[0193] In this scheme, the three-dimensional position information of the target key points is used to determine the main control key points and auxiliary key points, thereby improving the accuracy of the pointing calculation and avoiding the limitations caused by directly using the target's hand key points to calculate the target's interactive position.
[0194] In one embodiment of this application, the interactive control module 903 is used to interactively control the target screen, including:
[0195] If a cursor control exists on the target screen, then control the cursor control to move to the target interactive position;
[0196] In response to a trigger operation at the target interaction location, interact with the target screen.
[0197] This solution generates cursor movement instructions based on the target interaction position to control the cursor control to move to the target interaction position, thereby realizing gesture-controlled cursor movement interaction, improving the flexibility of interaction control, and reducing the cost of interaction devices.
[0198] In one embodiment of this application, the interactive control module is further used to determine pose information, specifically including the following steps:
[0199] Acquire a reference image and a calibration image, wherein the target image and the reference image include a reference object, the reference image is acquired by a target imaging device that acquires the target image, and the calibration image is acquired by an auxiliary imaging device that captures the target screen, and the target imaging device and the auxiliary imaging device have a common viewing area;
[0200] The first pose information of the reference object is determined based on the image coordinates of the reference object in the reference image;
[0201] Based on the image coordinates of the reference object in the calibration image, the second pose information of the reference object is determined.
[0202] The pose information of the target screen is determined based on the relative pose information between the first pose information and the second pose information.
[0203] The screen interaction device provided in this application includes an acquisition module for acquiring key hand points of the user controlling the target screen. These key hand points are determined through user localization and skeletal tracking of a target image containing the user. A determination module is used to determine at least two target key points based on the three-dimensional position information of each key hand point. A virtual ray formed by the three-dimensional position information of these at least two target key points points points from the user's hand towards the target screen. An interaction control module is used to determine the target interaction position of the target screen based on the three-dimensional position information of the target key points and the pose information of the target screen, and then perform interactive control on the target screen. This solution analyzes the pose of the key hand points and the target screen from a three-dimensional spatial perspective to determine the target interaction position of the user on the target screen, thereby performing interactive control. It eliminates the need for additional interactive devices such as a mouse, reducing equipment configuration costs. Furthermore, the spatial pointing analysis based on three-dimensional space ensures the accuracy of the interaction and improves interaction performance.
[0204] Furthermore, it is understood that in some other embodiments of this application, a display device is also provided, which integrates any of the screen interaction devices provided in the embodiments of the present invention, the display device comprising:
[0205] One or more processors;
[0206] Memory; and
[0207] One or more applications, wherein the one or more applications are stored in the memory and configured to be executed by the processor of the steps in the screen interaction method described in any of the above embodiments of the screen interaction method.
[0208] As can be seen from the above embodiments, in some embodiments of this application, the processors and memory in the display device are integrated on the circuit board body included in the display device, and the circuit board body is disposed in the display device.
[0209] It is understood that in some other embodiments of this application, the processor and memory in the display device may not be integrated into the display device, that is, the processor and memory are respectively disposed in the display device as components of the display device.
[0210] like Figure 10 As shown, Figure 10 This is a schematic diagram of an embodiment of the display device provided in this application.
[0211] Specifically, a display device may include components such as a processor 1001 with one or more processing cores, a memory 1002 with one or more computer-readable storage media, a power supply 1003, and an input unit 1004. Those skilled in the art will understand that... Figure 10 The display device structure shown does not constitute a limitation on the display device and may include more or fewer components than shown, or combine certain components, or have different component arrangements. Wherein:
[0212] The processor 1001 is the image processing center, connecting various parts of the display device via various interfaces and lines. It executes software programs and / or modules stored in the memory 1002, and calls data stored in the memory 1002, to perform various functions and process data of the display device, thereby providing overall monitoring of the display device. It is understood that the processor 1001 communicates with the controller via signal transmission. Optionally, the processor 1001 may include one or more processing cores; preferably, the processor 1001 may integrate an application processor and a modem processor, wherein the application processor mainly handles the operating system, user interface, and applications, while the modem processor mainly handles wireless communication. It is understood that the aforementioned modem processor may also not be integrated into the processor 1001.
[0213] The memory 1002 can be used to store software programs and modules. The processor 1001 executes various functional applications and data processing by running the software programs and modules stored in the memory 1002. The memory 1002 may mainly include a program storage area and a data storage area. The program storage area may store the operating system, application programs required for at least one function (such as sound playback function, image playback function, etc.), etc.; the data storage area may store data created according to the use of the display device, etc. In addition, the memory 1002 may include high-speed random access memory, and may also include non-volatile memory, such as at least one disk storage device, flash memory device, or other volatile solid-state storage device. Accordingly, the memory 1002 may also include a memory controller to provide the processor 1001 with access to the memory 1002.
[0214] In some embodiments of this application, the screen interaction device can be implemented as a computer program, and the computer program can be implemented as follows: Figure 10 The device operates on the display device shown. The display device's memory can store the various program modules that make up the screen interaction device, for example, Figure 9 The diagram shows a response acquisition module 901, a determination module 902, and an interaction control module 903. The computer program comprised of these modules causes the processor to execute the steps of the screen interaction methods in the various embodiments of this application described in this specification.
[0215] For example, Figure 10 The display device shown can be used as follows Figure 9The response acquisition module 901 in the screen interaction device shown executes step S201. The display device can execute step S202 through the determination module 902. The display device can execute step S203 through the interaction control module 903. The display device includes a processor, memory, and network interface connected via a system bus. The processor of the display device provides computing and control capabilities. The memory of the display device includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The network interface of the display device is used to communicate with external display devices via a network connection. When the computer program is executed by the processor, it implements a screen interaction method.
[0216] The display device also includes a power supply 1003 that supplies power to the various components. Preferably, the power supply 1003 can be logically connected to the processor 1001 through a power management system, thereby enabling functions such as charging, discharging, and power consumption management through the power management system. The power supply 1003 may also include one or more DC or AC power supplies, recharging systems, power fault detection circuits, power converters or inverters, power status indicators, and other arbitrary components.
[0217] The display device may also include an input unit 1004, which can be used to receive input digital or character information, and generate keyboard, mouse, joystick, optical or trackball signal inputs related to user settings and function control.
[0218] Although not shown, the display device may also include display units, etc., which will not be described in detail here. Specifically, in this embodiment, the processor 1001 in the display device loads the executable files corresponding to the processes of one or more application programs into the memory 1002 according to the following instructions, and the processor 1001 runs the application programs stored in the memory 1002 to realize various functions, as follows:
[0219] The key hand points of the controller who controls the target screen are obtained, and the key hand points are determined by user localization and skeleton tracking of the target image containing the controller.
[0220] Based on the three-dimensional position information of each of the aforementioned key hand points, at least two target key points are determined, and a virtual ray formed by the three-dimensional position information of the at least two target key points points points from the controller's hand to the target screen.
[0221] Based on the three-dimensional position information of the target key points and the pose information of the target screen, the target interaction position of the target screen is determined, and the target screen is interactively controlled.
[0222] Those skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be performed by instructions, or by instructions controlling related hardware. These instructions can be stored in a computer-readable storage medium and loaded and executed by a processor.
[0223] Therefore, embodiments of the present invention provide a computer-readable storage medium, which may include: read-only memory (ROM), random access memory (RAM), magnetic disk, or optical disk, etc. A computer program is stored thereon, which is loaded by a processor to execute the steps in any of the screen interaction methods provided in the embodiments of the present invention. For example, the computer program loaded by the processor can execute the following steps:
[0224] The key hand points of the controller who controls the target screen are obtained, and the key hand points are determined by user localization and skeleton tracking of the target image containing the controller.
[0225] Based on the three-dimensional position information of each of the aforementioned key hand points, at least two target key points are determined, and a virtual ray formed by the three-dimensional position information of the at least two target key points points points from the controller's hand to the target screen.
[0226] Based on the three-dimensional position information of the target key points and the pose information of the target screen, the target interaction position of the target screen is determined, and the target screen is interactively controlled.
[0227] In the above embodiments, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the detailed descriptions of other embodiments above, which will not be repeated here.
[0228] In practice, each of the above units or structures can be implemented as an independent entity or can be arbitrarily combined to be implemented as the same or several entities. For the specific implementation of each of the above units or structures, please refer to the previous method embodiments, which will not be repeated here.
[0229] For details on the implementation of each of the above operations, please refer to the previous examples, which will not be repeated here.
[0230] The foregoing has provided a detailed description of a screen interaction method, apparatus, display device, and storage medium provided in the embodiments of this application. Specific examples have been used to illustrate the principles and implementation methods of the present invention. The descriptions of the above embodiments are only for the purpose of helping to understand the method and core ideas of the present invention. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of the present invention. Therefore, the content of this specification should not be construed as a limitation of the present invention.
Claims
1. A screen interaction method, characterized in that, include: The key hand points of the controller who controls the target screen are obtained, and the key hand points are determined by user localization and skeleton tracking of the target image containing the controller. Based on the three-dimensional position information of each of the aforementioned key hand points, at least two target key points are determined, and a virtual ray formed by the three-dimensional position information of the at least two target key points points points from the controller's hand to the target screen. Based on the three-dimensional position information of the target key points and the pose information of the target screen, the target interaction position of the target screen is determined, and the target screen is interactively controlled.
2. The screen interaction method according to claim 1, characterized in that, Before acquiring the key hand points of the controller controlling the target screen, the method includes: Acquire the target image for controlling the target screen; Human detection is performed on the target image using a human detection model to identify the main controller in the target image and the hand area of the main controller; The key points of the hand region are detected by the hand key point detection model to determine the two-dimensional position information of the hand joints of the controller, as well as the confidence level of the two-dimensional position information; Based on the two-dimensional position information and the confidence level corresponding to each two-dimensional position information, key points of the hand are determined.
3. The screen interaction method according to claim 1, characterized in that, The step of determining at least two target key points based on the three-dimensional position information of each of the hand key points includes: The three-dimensional position information of the key hand points is obtained by predicting the three-dimensional position information of the key hand points using a hand joint model; Based on the three-dimensional position information of each hand key point and the relative positional relationship between each hand key point, the hand key points of the same hand limb are determined; Based on the hand key points of each hand limb and the pose information of the target screen, the target hand key points of the hand limb pointing to the target screen are determined.
4. The screen interaction method according to claim 3, characterized in that, The step of determining the target hand key points of the hand limb pointing to the target screen based on the hand key points of each hand limb and the pose information of the target screen includes: The bending angle of each segment is calculated based on the three-dimensional position information of every two adjacent key points of the same hand limb. The bending angle of each segment in each hand limb is counted to obtain the bending degree of the hand limb; Acquire the target hand limb with a curvature smaller than a preset curvature; Based on the three-dimensional position information of the target hand limb and the pose information of the target screen, the target hand key points pointing to the target hand limb on the target screen are determined.
5. The screen interaction method according to claim 1, characterized in that, Determining the target interaction position of the target screen based on the three-dimensional position information of the target key points and the pose information of the target screen includes: Based on the three-dimensional position information of the target key points, the main control key points and auxiliary key points are determined. The main control key point is the fingertip key point, and the auxiliary key points are at least one of the wrist key point, interphalangeal key point, and palmofinite key point. Based on the pose information of the target screen and the virtual ray formed by the main control key point and the auxiliary key point, the target interaction position of the target screen is determined.
6. The screen interaction method according to claim 1, characterized in that, The interactive control of the target screen includes: If a cursor control exists on the target screen, then control the cursor control to move to the target interactive position; In response to a trigger operation at the target interaction location, interact with the target screen.
7. The screen interaction method according to any one of claims 1-6, characterized in that, Determining the pose information includes the following steps: Acquire a reference image and a calibration image, wherein the target image and the reference image include a reference object, the reference image is acquired by a target imaging device that acquires the target image, and the calibration image is acquired by an auxiliary imaging device that captures the target screen, and the target imaging device and the auxiliary imaging device have a common viewing area; The first pose information of the reference object is determined based on the image coordinates of the reference object in the reference image; Based on the image coordinates of the reference object in the calibration image, the second pose information of the reference object is determined. The pose information of the target screen is determined based on the relative pose information between the first pose information and the second pose information.
8. A screen interaction device, characterized in that, The device includes: The acquisition module is used to acquire the key hand points of the master controller who controls the target screen. The key hand points are determined by user positioning and skeleton tracking of the target image containing the master controller. The determination module is used to determine at least two target key points based on the three-dimensional position information of each of the hand key points, and the virtual ray formed by the three-dimensional position information of the at least two target key points points points from the hand of the controller to the target screen; The interactive control module is used to determine the target interactive position of the target screen based on the three-dimensional position information of the target key points and the pose information of the target screen, and to perform interactive control on the target screen.
9. A display device, characterized in that, The display device includes: One or more processors; Memory; and One or more applications, wherein the one or more applications are stored in the memory and configured to be executed by the processor to implement the steps of the screen interaction method according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, It stores a computer program, which is loaded by a processor to perform the steps of the screen interaction method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Interaction control method and device, electronic equipment and storage medium
CN111949111A
Method for interacting with virtual reality equipment and virtual reality equipment
CN112198962A
Non-contact screen control method and device, electronic equipment and readable storage medium
CN114360049A
Interactive control method and apparatus, electronic device and storage medium
US20220066545A1