A two-hand position recognition method, system, device and medium
By using a depth camera to acquire 3D information and calculate the weighted center point, the problem of limited gesture recognition range in existing technologies is solved, achieving high-precision and efficient gesture recognition at different distances, thus improving the convenience and accuracy of user operation.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- SHENZHEN GUANGJIAN TECH CO LTD
- Filing Date
- 2023-06-29
- Publication Date
- 2026-07-28
AI Technical Summary
Existing technologies in gesture recognition are limited to a fixed range and cannot effectively cover gesture applications at greater distances, resulting in poor usability, accuracy, recognition efficiency, and human-computer interaction effects.
Using a depth camera to acquire 3D information of the left hand, right hand, and eyes, the line connecting the eye position and the hand anchor point is calculated and intersected with the display at position a. The weighting center point c is determined by different weight values. Different weight values are given based on the distance of the eyes and hands from the display, thereby realizing the switching of control modes at different distances.
It enables effective gesture recognition at both close and long distances, improving the operating range and convenience, enhancing the accuracy of position control and the consistency of response, and reducing the false recognition rate.
Smart Images

Figure CN116844229B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of gesture recognition technology, and more specifically, to a method, system, device, and medium for recognizing the position of both hands. Background Technology
[0002] Gesture recognition aims to identify human gestures using mathematical algorithms. Gestures can originate from any body movement or state. Users can use simple gestures to control or interact with devices without touching them. Gesture recognition can be viewed as a way for computers to understand human language, thereby enabling human-computer interaction and functional performance.
[0003] There has been considerable research on the application of gestures. For example, a certain existing technology provides a gesture recognition device, method, and system. The gesture recognition device includes: at least one sensor disposed corresponding to the position of a finger, the sensor being used to identify the movement information of the moving finger relative to other fingers; an input mode determination unit, used to determine the input mode required by the recognition device based on the gesture movement information and angle change information detected by at least one sensor, wherein the input mode includes at least one of a simulated keyboard input method, simulated function keys, and simulated mouse; and an input content generation unit, which, when the input mode determination unit determines the required input mode, generates one of the following based on the gesture movement information and angle change information detected by at least one sensor: the corresponding input position on the keyboard, the switching of function keys, or the direction of mouse movement; thus, virtual human-computer interaction can be realized, and fast and accurate simulated input content can be obtained.
[0004] A certain prior art relates to a gesture recognition method, a gesture recognition module, and a gesture recognition system. The gesture recognition method includes: performing a binarization process on an image to obtain a binarized image, wherein the binarized image includes a plurality of foreground pixels and a plurality of background pixels; determining whether the plurality of foreground pixels in the binarized image surround at least one first background pixel; and when the plurality of foreground pixels surround the first background pixel, determining that a gesture conforms to a preset gesture.
[0005] However, existing technologies are limited to applications within a fixed range, restricting gesture recognition to a very narrow area, such as within 20cm. They fail to effectively cover gesture applications at greater distances, such as over 1 meter. Existing technologies suffer from issues regarding the usable range, accuracy, recognition efficiency, usability, and human-computer interaction of gestures.
[0006] The above background information is provided only to aid in understanding the inventive concept and technical solution of this invention. It does not necessarily belong to the prior art of this patent application. In the absence of clear evidence that the above information was disclosed on the filing date of this patent application, the above background information should not be used to evaluate the novelty and inventiveness of this application. Summary of the Invention
[0007] Therefore, this invention uses a depth camera to obtain three-dimensional information of the left hand, right hand and eyes, calculates the line connecting the eyeball position and the hand anchor point, intersects the display at position a, and determines the final weighted center point c through different weight values. It has the advantages of large operability, high accuracy, low cost, simple operation and flexible settings.
[0008] In a first aspect, the present invention provides a method for recognizing the position of both hands, characterized by comprising the following steps:
[0009] Step S1: Use a depth camera to obtain 3D information of the left hand, right hand, and eyes;
[0010] Step S2: Determine the eye position, the center point of the face, and the hand anchor point; wherein, the hand anchor point is determined by the left hand gesture, the right hand gesture, and the relative positions of the left and right hands;
[0011] Step S3: Calculate the line connecting the eyeball position and the hand anchor point, and intersect the display at position a; calculate the line connecting the face center point and the hand anchor point, and intersect the display at position b;
[0012] Step S4: Based on the hand anchor point and the distance between the eye and the display, assign weight values α and β to position a and position b respectively, and calculate the weighted center point c of position a and position b; where α + β = 1.
[0013] Optionally, the hand position recognition method is characterized in that, in step S2, the positions of the left and right pupils are determined as the eyeball positions.
[0014] Optionally, the hand position recognition method is characterized in that, in step S3, a line is drawn from the position of the left pupil to the hand anchor point, intersecting the display at position a1, and a weight value μ is assigned; a line is drawn from the position of the right pupil to the hand anchor point, intersecting the display at position a1, and a weight value ν is assigned; and the weighted center point a of positions a1 and a2 is calculated; wherein, μ + ν = 1, and the values of μ and ν are related to the distance difference between the left pupil and the right pupil and the display.
[0015] Optionally, the method for recognizing the position of both hands is characterized in that determining the hand anchor points includes:
[0016] Step S21: Determine the valid hand based on the correspondence and positional relationship between the left hand, right hand and face;
[0017] Step S22: Identify the key points of the effective hand and determine the hand state based on the distribution of the key points;
[0018] Step S23: If the hand state is the first state, then take some key points of the effective hand as the left hand anchor point or the right hand anchor point; if the hand state is the second state, then take the center point of at least some key points of the effective hand as the left hand anchor point or the right hand anchor point; take the midpoint of the left hand anchor point or the right hand anchor point as the hand anchor point.
[0019] Optionally, the method for recognizing the position of both hands is characterized in that step S21 includes:
[0020] Step S211: Determine the first face and the first hand based on whether the hand is in front of the face;
[0021] Step S212: Remove faces and hands whose facial normal vectors form an angle smaller than the first angle with the display, to obtain a second face and a second hand;
[0022] Step S213: Compare the second hand with the preset hand posture and determine the hand with the highest similarity as the valid hand.
[0023] Optionally, the method for recognizing the position of both hands is characterized in that step S23 includes:
[0024] Step S231: If the hand state is the first state and the number of protruding fingers is 1, then take the key point of the fingertip of the protruding finger as the left hand anchor point or the right hand anchor point.
[0025] Step S232: If the hand state is the first state and the number of protruding fingers is greater than 1, then take the center point of the key point of the fingertip of the protruding finger as the left hand anchor point or the right hand anchor point.
[0026] Step S233: If the hand state is the second state, then the center point of at least some key points of the effective hand is determined as the left hand anchor point or the right hand anchor point;
[0027] Step S234: Determine the midpoint of the left-hand anchor point or the right-hand anchor point as the hand anchor point.
[0028] Optionally, the two-hand position recognition method is characterized in that the weighted center point c has a correction parameter ε relative to the actual position of the mouse on the display; wherein the correction parameter ε is related to the position of the weighted center point c on the display.
[0029] In a second aspect, the present invention provides a two-hand position recognition system for implementing the two-hand position recognition method described in any of the above claims, characterized in that it includes:
[0030] The acquisition module is used to obtain 3D information of the left hand, right hand, and eyes using a depth camera;
[0031] A determination module is used to determine the position of the eyeballs, the center point of the face, and the hand anchor points; wherein the hand anchor points are determined by the left hand gesture, the right hand gesture, and the relative positions of the left and right hands;
[0032] The connection module is used to calculate the line connecting the eyeball position and the hand anchor point, and intersect the display at position a; calculate the line connecting the face center point and the hand anchor point, and intersect the display at position b;
[0033] The calculation module is used to assign weight values α and β to position a and position b respectively based on the hand anchor point and the distance of the eye from the display, and to calculate the weighted center point c of position a and position b; where α + β = 1.
[0034] Thirdly, the present invention provides a two-hand position recognition device, characterized in that it comprises:
[0035] processor;
[0036] A memory in which executable instructions of the processor are stored;
[0037] The processor is configured to perform the steps of any of the above-described hand position recognition methods by executing the executable instructions.
[0038] Fourthly, the present invention provides a computer-readable storage medium for storing a program, characterized in that, when the program is executed, it implements the steps of the hand position recognition method described in any one of the preceding claims.
[0039] Compared with the prior art, the present invention has the following beneficial effects:
[0040] This invention utilizes a depth camera to obtain three-dimensional information of the left hand, right hand, and eyes. Processing can be completed using the contour information of the hands and eyes. The accuracy requirements for the information of the hands and eyes are low, and effective recognition can be achieved at both close and long distances, greatly increasing the range and convenience of user operation.
[0041] This invention calculates hand anchor points using the left and right hands, allowing users to achieve more precise control over the interface by adjusting the positions of the left and right hands, greatly improving the accuracy of position control.
[0042] This invention calculates the intersection point between the line connecting the eye and hand and the display screen, eliminating the need for complex calculations. It is characterized by its simplicity and accuracy. Furthermore, it can utilize the intersection feature to select between the hand and the face, thereby achieving the recognition of valid faces and reducing the probability of false recognition.
[0043] This invention assigns different weight values to the distance between the hand and eyes and the display, enabling smooth transitions between control modes of the human body and the display at different distances. This allows for effective response to human body control at different distances, improving the consistency of the response. Attached Figure Description
[0044] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are merely embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort. Other features, objects, and advantages of the present invention will become more apparent by reading the following detailed description of non-limiting embodiments with reference to the accompanying drawings:
[0045] Figure 1 This is a flowchart illustrating the steps of a two-hand position recognition method according to an embodiment of the present invention;
[0046] Figure 2 This is a schematic diagram of a two-hand position recognition embodiment of the present invention;
[0047] Figure 3 This is a schematic diagram of a weighted center point in an embodiment of the present invention;
[0048] Figure 4 This is a flowchart illustrating one step in determining the hand anchor point according to an embodiment of the present invention;
[0049] Figure 5 This is a flowchart illustrating the steps for determining a valid hand in an embodiment of the present invention;
[0050] Figure 6 This is a flowchart illustrating the steps for classifying and determining the hand anchor points in an embodiment of the present invention.
[0051] Figure 7 This is a schematic diagram of a calibration system according to an embodiment of the present invention;
[0052] Figure 8 This is a schematic diagram of a non-contact calibration system according to an embodiment of the present invention;
[0053] Figure 9 This is a schematic diagram of another non-contact calibration system in an embodiment of the present invention;
[0054] Figure 10 This is a schematic diagram of the structure of a two-hand position recognition system according to an embodiment of the present invention;
[0055] Figure 11 This is a schematic diagram of the structure of a two-hand position recognition device according to an embodiment of the present invention; and
[0056] Figure 12 This is a schematic diagram of the structure of a computer-readable storage medium in an embodiment of the present invention. Detailed Implementation
[0057] The present invention will now be described in detail with reference to specific embodiments. These embodiments will help those skilled in the art to further understand the present invention, but do not limit the invention in any way. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of the present invention. These all fall within the scope of protection of the present invention.
[0058] The terms “first,” “second,” “third,” “fourth,” etc. (if present) in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a particular order or sequence. It should be understood that such data can be interchanged where appropriate so that embodiments of the invention described herein can be implemented, for example, in orders other than those illustrated or described herein. Furthermore, the terms “comprising” and “having,” and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0059] The present invention provides a method for recognizing the position of both hands, which aims to solve the problems existing in the prior art.
[0060] The technical solutions of the present invention and how they solve the above-mentioned technical problems will be described in detail below with specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments. The embodiments of the present invention will now be described with reference to the accompanying drawings.
[0061] This invention utilizes a depth camera to obtain three-dimensional information of the left hand, right hand, and eyes, calculates the line connecting the eyeball position and the hand anchor point, intersects the line with the display at position a, and determines the final weighted center point c through different weight values. It has the advantages of wide operability, high precision, low cost, simple operation, and flexible settings.
[0062] Figure 1 This is a flowchart illustrating the steps of a two-hand position recognition method according to an embodiment of the present invention. Figure 1 As shown, the steps of a two-hand position recognition method in an embodiment of the present invention include:
[0063] Step S1: Use a depth camera to obtain 3D information of the left hand, right hand and eyes.
[0064] In this step, the depth camera used can be of any type, including but not limited to structured light cameras, TOF cameras, and stereo cameras. The hands and eyes are identified from the acquired depth image, thus obtaining their 3D information. For example... Figure 2 As shown, both the depth camera and the display are facing the human body. The human body places its left and right hands in front of its eyes to control the display.
[0065] Step S2: Determine the position of the eyeballs, the center point of the face, and the anchor point of the hands.
[0066] In this step, the hand anchor point is determined by the left hand gesture, the right hand gesture, and the relative positions of the left and right hands. The eye position can be distinguished based on eye color. The face center point is the point closest to the center of the face; it can be the center of all points on the face surface or the center point on a plane perpendicular to the face's orientation. The face center point can even be a key facial point other than the eyes. The face center point characterizes facial features and determines the distance to the display or hand. The hand anchor point is a point that represents the hand while also conforming to human visual characteristics. The left hand anchor point is determined by the left hand gesture, the right hand anchor point by the right hand gesture, and then the hand anchor point is determined by combining the left and right hand anchor points. Various methods can be used to determine the left or right hand anchor point, such as using gestures, or using the hand center point as the left or right hand anchor point. The hand anchor point can be obtained by taking an intermediate value from the left and right hand anchor points, such as the midpoint or other proportional points.
[0067] In some embodiments, the positions of the left and right pupils are determined separately as the eyeball positions. In this case, the positions of the left and right eyeballs are determined separately. This embodiment is suitable for situations where the distance between the eyes and the depth camera is relatively short, and the pupil position can be directly determined using eyeball information, resulting in higher accuracy. For example, when a person is positioned in front of a computer, the eyes are relatively close to the screen.
[0068] Step S3: Calculate the line connecting the eyeball position and the hand anchor point, and intersect the display at position a; calculate the line connecting the face center point and the hand anchor point, and intersect the display at position b.
[0069] In this step, 3D reconstruction is performed to obtain the intersection point of the line connecting the hand anchor point, the eye position, and the center point of the face with the display. In this embodiment, the positions of the depth camera and the display are fixed and known. The depth camera and the display can be integrated or assembled. When the depth camera and the display are integrated, the depth camera can be located at the edge of the display or inside the display, such as through a display opening. When the depth camera and the display are assembled, they can be fixed separately using fixing devices, and then calibrated to obtain a more accurate positional relationship between the display and the depth camera. The eye position can refer to the position of the left and right eyes separately, or the average position of the left and right eyes. When the eye position is the position of the left and right eyes, lines are drawn connecting the left and right eyeballs to the hand anchor point, the intersection point with the display is calculated, and then the intersection point is manipulated. When the eye position is the average position of the left and right eyes, first take the average position of the left and right eyeballs, then connect it with the hand anchor point, and calculate the intersection point with the monitor.
[0070] In some embodiments, a line is drawn from the position of the left pupil to the hand anchor point, intersecting the display at position a1, and assigned a weight value μ; a line is drawn from the position of the right pupil to the hand anchor point, intersecting the display at position a1, and assigned a weight value ν; and a weighted center point a is calculated for positions a1 and a2; wherein, μ + ν = 1, and the values of μ and ν are related to the distance difference between the left pupil and the right pupil and the display.
[0071] In some embodiments, the weighted center point c has a correction parameter ε relative to the actual position of the mouse on the display; wherein, the correction parameter ε is related to the position of the weighted center point c on the display.
[0072] Step S4: Based on the hand anchor point and the distance between the eye and the display, assign weight values α and β to position a and position b respectively, and calculate the weighted center point c of position a and position b.
[0073] In this step, α + β = 1. Based on the hand anchor point and the distance between the eye and the display, different weight values are assigned to the corresponding positions 'a' and 'b' for the hand and eye, respectively, so that the weight of the eyeball increases when near and the weight of the face increases when far. When the distance between the hand and eye and the display increases, the weight value α decreases and the weight value β increases; conversely, when the distance between the hand and eye and the display increases, the weight value α increases and the weight value β decreases. Figure 3 As shown, the weighted center point c is located on the line connecting position a and position b, and is biased towards the point with the larger weight value.
[0074] Figure 4 This is a flowchart illustrating one step in determining the hand anchor point according to an embodiment of the present invention. Figure 4 As shown, one step in determining the hand anchor point in an embodiment of the present invention includes:
[0075] Step S21: Determine the valid hand based on the correspondence and positional relationship between the left hand, right hand and face.
[0076] In this step, when multiple people are present in the target space, multiple hands and figures are likely to appear. This is more likely to occur at a distance, such as when watching television. This step eliminates irrelevant hands by checking whether the hands are in front of the face and by considering the correspondence between the left and right hands, thus obtaining the valid hands. A valid hand can be one or more.
[0077] Step S22: Identify the key points of the effective hand and determine the hand state based on the distribution of the key points.
[0078] In this step, keypoint recognition algorithms are used to identify key points of the valid hand. For example, 21 commonly used hand keypoints are used. There are various methods to determine the hand state based on these keypoints. For instance, simple gesture recognition can be achieved by calculating the angles between the detected hand keypoints. For example, the angle between the thumb vectors 0-2 and 3-4 can be calculated; angles greater than a certain value are defined as bent, and angles less than a certain value are defined as straight.
[0079] Step S23: If the hand state is the first state, then take some key points of the effective hand as the left hand anchor point or the right hand anchor point; if the hand state is the second state, then take the center point of at least some key points of the effective hand as the left hand anchor point or the right hand anchor point; take the midpoint of the left hand anchor point or the right hand anchor point as the hand anchor point.
[0080] In this step, the first state consists of pre-set gestures, with some key points being more prominent, such as extending one finger or two fingers. The second state consists of pre-set gestures, with key points being more evenly distributed, or without any pre-set gestures, such as spreading all five fingers or making a fist.
[0081] This embodiment identifies the effective hand and regions the hand state of the effective hand, determining some key points or center points as left-hand or right-hand anchor points, thereby determining the hand anchor points. This is more in line with human perception, lowers the threshold for gesture control, and allows users to operate more precisely, conveniently, and efficiently.
[0082] Figure 5 This is a flowchart illustrating the steps for determining a valid hand in an embodiment of the present invention. Figure 5 As shown, one step in determining a valid hand in an embodiment of the present invention includes:
[0083] Step S211: Determine the first face and the first hand based on whether the hand is in front of the face.
[0084] In this step, the first face and first hand are determined based on their positional relationship with the face. The space in front of the face is defined as "front," not just directly in front. "Front" includes the space in front of the face, typically with a cross-sectional area of at least 120 cm². 2 The depth range from the face to the display screen is acceptable. The number of hands in the first step is twice the number of hands in the first step. This step also requires recognizing the matching relationship between the face and the hands, and determining the left and right hands corresponding to the face.
[0085] Step S212: Remove faces and hands whose facial normal vectors form an angle smaller than the first angle with the display, to obtain a second face and a second hand.
[0086] In this step, faces that are not facing the monitor are removed, thereby interrupting the operation when a person is looking at others or other areas, effectively preventing accidental operation.
[0087] Step S213: Compare the second hand with the preset hand posture and determine the hand with the highest similarity as the valid hand.
[0088] In this step, the final valid hand is selected from multiple second hands. The number of valid hands can be one pair, two pairs, or even more. For example, when a person is facing a monitor, one or two people can perform gesture operations, such as in two-player games and other interactive content, increasing the fun. The number of valid hands can be preset.
[0089] This embodiment further filters the identified hands and can obtain the corresponding valid hands according to preset conditions, so that it can be effectively identified even when there are multiple people.
[0090] Figure 6 This is a flowchart illustrating the steps for classifying and determining the hand anchor points in an embodiment of the present invention. Figure 6 As shown, in one embodiment of the present invention, a step for classifying and determining the hand anchor point includes:
[0091] Step S231: If the hand state is the first state and the number of protruding fingers is 1, then take the key point of the fingertip of the protruding finger as the left hand anchor point or the right hand anchor point.
[0092] In this step, if only one finger is protruding, the key point at the tip of that finger is taken as the anchor point for the left or right hand. For example, if only one index finger is protruding, the intersection point on the monitor is determined by the line connecting the eyeball and face to the tip of the index finger.
[0093] Step S232: If the hand state is the first state and the number of protruding fingers is greater than 1, then take the center point of the key point of the fingertip of the protruding finger as the left hand anchor point or the right hand anchor point.
[0094] In this step, if there are two or more protruding fingers, the center point of the fingertip of the protruding finger is taken as the hand anchor point. When calculating the center point, if there are two fingers, the center of the key points of the two fingertips is taken; if there are more than two fingers, the centroid of the triangle or polygon formed by the key points of the fingertips of multiple fingers is taken.
[0095] Step S233: If the hand state is the second state, then the center point of at least some key points of the effective hand is determined as the left hand anchor point or the right hand anchor point.
[0096] In this step, the second state represents gestures with relatively evenly distributed keypoints or undefined gestures. It's necessary to process all keypoints or some of them to determine the center point of each keypoint as the left or right hand anchor point. The gestures processed in this step typically involve performing specific operations, such as grabbing or right-clicking, requiring calculation of the overall orientation.
[0097] Step S234: Determine the midpoint of the left-hand anchor point or the right-hand anchor point as the hand anchor point.
[0098] In this step, the hand anchor point is not located on either the left or right hand, but rather in a position between the left and right hands. By adjusting the positions of the left and right hands, users can more precisely control the intersection of the hand anchor point and the display, thereby achieving more precise control over gesture accuracy.
[0099] This embodiment categorizes different gestures in detail and performs corresponding operations for each, enabling precise operation of various gestures, thereby achieving accurate definition and recognition of gestures and obtaining more accurate recognition data.
[0100] Figure 7 This is a schematic diagram of a calibration system according to an embodiment of the present invention. Figure 7 As shown, in an embodiment of the present invention, a calibration system includes a depth camera, a display, and a calibration rod.
[0101] The calibration rod is designed to measure the distance to four specified vertices of the screen and can be divided into contact and non-contact types. The contact calibration rod is a slender rod of the first color with a fixed length of L, and the end is marked with the second color. The non-contact calibration rod has a laser emitter at the front end.
[0102] The calibration of a contact calibration rod includes the following steps:
[0103] Step S61: Place the end of the calibration rod in contact with point P1 on the screen and continuously capture N depth images;
[0104] Step S62: Extract the spatial straight line equation L1 of the calibration rod, and calculate the spatial coordinates P1xN based on the length of the calibration rod;
[0105] Step S63: Repeat the above steps to touch the other 3 corner points on the screen to obtain L2, L3, L4; finally, 4 sets of images are obtained, totaling 4xn, and then the spatial coordinates of {P2, P3, P4}xN are obtained.
[0106] Step S64: Obtain the plane equation of the display based on {P1,P2,P3,P4}xN;
[0107] Step S65: Finally, based on the intersection of any straight line and the calibration plane, determine the positions of the four corner points {P1, P2, P3, P4}.
[0108] The following steps are included when extracting the calibration rod:
[0109] Step S71: Extract the first channel data from the depth image, extract the ROI of the calibration rod, and remove noise using morphological methods; wherein, the first channel can extract the first color;
[0110] Step S72: Fit a straight line using the Ransanc method, exclude exterior points, and output interior points_2d;
[0111] Step S73: Transform points_2d to points_3d / Make the camera coordinate system and world coordinate system coincide, and fit line_3d;
[0112] Step S74: Extract the second channel data from the depth image, extract the calibration rod startpointroi, substitute it into the 2d / 3d line equation, filter it, and finally obtain the most ideal startpoint; among them, the second channel can extract the second color.
[0113] This embodiment, through channel coordination, can effectively improve the identification and acquisition of calibration rods.
[0114] Figure 8 This is a schematic diagram of a non-contact calibration system according to an embodiment of the present invention. Figure 8 As shown, a non-contact calibration method in an embodiment of the present invention includes:
[0115] Step S81: Point the calibration rod to p1, change its position, and point it to p1 again. Repeat the above steps N times to obtain N depth images.
[0116] Step S82: Extract the spatial linear equations L1_0...L1_3...,N lines of the calibration rod;
[0117] Step S83: Find the same point p where different lines intersect, or two points whose distance is less than a certain threshold, and calculate the position p1;
[0118] Step S84: Repeat the above steps to obtain p2, p3, p4, and finally solve the equation of the display plane.
[0119] Figure 9 This is a schematic diagram of another non-contact calibration system according to an embodiment of the present invention. Figure 9 As shown, compared to the previous embodiments, this embodiment does not require a calibration rod. The calibration process includes the following steps:
[0120] Step S91: Extract the human eye position point_eye using an image algorithm;
[0121] Step S91: Extract the position of the fingertip, point_finger, using an image algorithm;
[0122] Step S91: Calculate the hand-eye pointing equation line_p1_0;
[0123] Step S91: Collect N images at each position p1, p2, p3, p4 to obtain a total of 4xNline_p;
[0124] Step S91: Find the same point p where different lines intersect, or two points whose distance is less than a certain threshold, and calculate the position p1;
[0125] Step S91: Repeat the aforementioned steps to obtain p2, p3, and p4, and finally solve the equation of the display plane.
[0126] Figure 10 This is a schematic diagram of a two-hand position recognition system according to an embodiment of the present invention. Figure 10 As shown, an embodiment of the present invention provides a two-hand position recognition system comprising:
[0127] The acquisition module is used to obtain 3D information of the left hand, right hand, and eyes using a depth camera;
[0128] A determination module is used to determine the position of the eyeballs, the center point of the face, and the hand anchor points; wherein the hand anchor points are determined by the left hand gesture, the right hand gesture, and the relative positions of the left and right hands;
[0129] The connection module is used to calculate the line connecting the eyeball position and the hand anchor point, and intersect the display at position a; calculate the line connecting the face center point and the hand anchor point, and intersect the display at position b;
[0130] The calculation module is used to assign weight values α and β to position a and position b respectively based on the hand anchor point and the distance of the eye from the display, and to calculate the weighted center point c of position a and position b; where α + β = 1.
[0131] Specifically, the acquisition module obtains 3D information about the left hand, right hand, and eyes. The determination module uses this information to perform 3D reconstruction and determine the eye positions, the face center point, and the hand anchor points. The connection module calculates the lines connecting the eye positions, the face center point, and the hand anchor points, as well as their intersections with the display. The calculation module assigns different weight values to different intersection points and calculates the final weighted center point.
[0132] This embodiment uses a depth camera to obtain three-dimensional information of the left hand, right hand, and eyes, calculates the line connecting the eyeball position and the hand anchor point, and intersects the display at position a. The final weighted center point c is determined by different weight values. It has the advantages of wide operability, high accuracy, low cost, simple operation, and flexible settings.
[0133] This invention also provides a two-hand position recognition device, including a processor and a memory storing executable instructions for the processor. The processor is configured to execute steps of a two-hand position recognition method by executing the executable instructions.
[0134] As shown above, this embodiment uses a depth camera to obtain three-dimensional information of the left hand, right hand and eyes, calculates the line connecting the eyeball position and the hand anchor point, and intersects the display at position a. The final weighted center point c is determined by different weight values. It has the advantages of large operability, high accuracy, low cost, simple operation and flexible settings.
[0135] Those skilled in the art will understand that various aspects of the present invention can be implemented as systems, methods, or program products. Therefore, various aspects of the present invention can be specifically implemented in the following forms: a completely hardware implementation, a completely software implementation (including firmware, microcode, etc.), or a combination of hardware and software aspects, collectively referred to herein as a "circuit," "module," or "platform."
[0136] Figure 11 This is a schematic diagram of the structure of a two-hand position recognition device according to an embodiment of the present invention. The following refers to... Figure 11 To describe an electronic device 600 according to this embodiment of the present invention. Figure 11 The electronic device 600 shown is merely an example and should not impose any limitations on the functionality and scope of use of the embodiments of the present invention.
[0137] like Figure 11 As shown, the electronic device 600 is presented in the form of a general-purpose computing device. The components of the electronic device 600 may include, but are not limited to: at least one processing unit 610, at least one storage unit 620, a bus 630 connecting different platform components (including storage unit 620 and processing unit 610), a display unit 640, etc.
[0138] The storage unit stores program code, which can be executed by the processing unit 610 to perform the steps described in the above-described section of this specification regarding a method for recognizing the position of both hands, according to various exemplary embodiments of the present invention. For example, the processing unit 610 can perform actions such as... Figure 1 The steps are shown in the figure.
[0139] Storage unit 620 may include a readable medium in the form of a volatile storage unit, such as random access memory (RAM) 6201 and / or cache memory 6202, and may further include a read-only memory (ROM) 6203.
[0140] Storage unit 620 may also include a program / utility 6204 having a set (at least one) program module 6205, such program module 6205 including but not limited to: operating system, one or more application programs, other program modules and program data, each or some combination of these examples may include an implementation of a network environment.
[0141] Bus 630 can represent one or more of several types of bus structures, including a memory cell bus or memory cell controller, a peripheral bus, a graphics acceleration port, a processing unit, or a local bus using any of the various bus structures.
[0142] Electronic device 600 can also communicate with one or more external devices 700 (e.g., keyboard, pointing device, Bluetooth device, etc.), and with one or more devices that enable a user to interact with electronic device 600, and / or with any device that enables electronic device 600 to communicate with one or more other computing devices (e.g., router, modem, etc.). This communication can be performed via input / output (I / O) interface 650. Furthermore, electronic device 600 can also communicate with one or more networks (e.g., local area network (LAN), wide area network (WAN), and / or public networks, such as the Internet) via network adapter 660. Network adapter 660 can communicate with other modules of electronic device 600 via bus 630. It should be understood that, although... Figure 11 As not shown in the diagram, other hardware and / or software modules may be used in conjunction with the electronic device 600, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage platforms.
[0143] This invention also provides a computer-readable storage medium for storing a program that, when executed, implements the steps of a two-hand position recognition method. In some possible implementations, various aspects of the invention can also be implemented as a program product comprising program code that, when run on a terminal device, causes the terminal device to perform the steps described in the foregoing section on a two-hand position recognition method according to various exemplary embodiments of the invention.
[0144] As shown above, this embodiment uses a depth camera to obtain three-dimensional information of the left hand, right hand and eyes, calculates the line connecting the eyeball position and the hand anchor point, and intersects the display at position a. The final weighted center point c is determined by different weight values. It has the advantages of large operability, high accuracy, low cost, simple operation and flexible settings.
[0145] Figure 12 This is a schematic diagram of the structure of a computer-readable storage medium according to an embodiment of the present invention. (Reference) Figure 12 As shown, a program product 800 for implementing the above-described method according to an embodiment of the present invention is described. This product may employ a portable compact disc read-only memory (CD-ROM) and include program code, and may run on a terminal device, such as a personal computer. However, the program product of the present invention is not limited thereto. In this document, the readable storage medium may be any tangible medium containing or storing a program that may be used by or in conjunction with an instruction execution system, apparatus, or device.
[0146] The program product may employ any combination of one or more readable media. A readable medium may be a readable signal medium or a readable storage medium. A readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples (a non-exhaustive list) of readable storage media include: electrical connections having one or more wires, portable disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.
[0147] Computer-readable storage media may include data signals propagated in baseband or as part of a carrier wave, carrying readable program code. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A readable storage medium may also be any readable medium other than a readable storage medium that can transmit, propagate, or transfer a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the readable storage medium may be transmitted using any suitable medium, including but not limited to wireless, wired, optical fiber, RF, etc., or any suitable combination thereof.
[0148] Program code for performing the operations of this invention can be written in any combination of one or more programming languages, including object-oriented programming languages such as Java and C++, and conventional procedural programming languages such as C or similar languages. The program code can execute entirely on the user's computing device, partially on the user's computing device, as a standalone software package, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server. In cases involving remote computing devices, the remote computing device can be connected to the user's computing device via any type of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computing device (e.g., via the Internet using an Internet service provider).
[0149] This embodiment uses a depth camera to obtain three-dimensional information of the left hand, right hand, and eyes, calculates the line connecting the eyeball position and the hand anchor point, and intersects the display at position a. The final weighted center point c is determined by different weight values. It has the advantages of wide operability, high accuracy, low cost, simple operation, and flexible settings.
[0150] The various embodiments described in this specification are presented in a progressive manner, with each embodiment focusing on its differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. The above description of the disclosed embodiments enables those skilled in the art to make or use the invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the invention. Therefore, the invention is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
[0151] The specific embodiments of the present invention have been described above. It should be understood that the present invention is not limited to the specific embodiments described above, and those skilled in the art can make various modifications or variations within the scope of the claims, which do not affect the essence of the present invention.
Claims
1. A method for recognizing the position of both hands, characterized in that, Includes the following steps: Step S1: Use a depth camera to obtain 3D information of the left hand, right hand, and eyes; Step S2: Determine the eye position, the center point of the face, and the hand anchor point; wherein, the hand anchor point is determined by the left hand gesture, the right hand gesture, and the relative positions of the left and right hands; Step S3: Calculate the line connecting the eyeball position and the hand anchor point, and intersect the display at position a; calculate the line connecting the face center point and the hand anchor point, and intersect the display at position b; Step S4: Based on the hand anchor point and the distance between the eye and the display, assign weight values α and β to position a and position b respectively, and calculate the weighted center point c of position a and position b; where α + β = 1; In step S2, the positions of the left and right pupils are determined respectively, which are used as the positions of the eyeballs; In step S3, a line is drawn from the position of the left pupil to the hand anchor point, intersecting the display at position a1, and assigned a weight value μ; a line is drawn from the position of the right pupil to the hand anchor point, intersecting the display at position a2, and assigned a weight value ν; and the weighted center point a of positions a1 and a2 is calculated; where μ + ν = 1, and the values of μ and ν are related to the distance difference between the left pupil and the right pupil and the display. Determining the hand anchor point includes: Step S21: Determine the valid hand based on the correspondence and positional relationship between the left hand, right hand and face; Step S22: Identify the key points of the effective hand and determine the hand state based on the distribution of the key points; Step S23: If the hand state is the first state, then take some key points of the effective hand as the left hand anchor point or the right hand anchor point; if the hand state is the second state, then take the center point of at least some key points of the effective hand as the left hand anchor point or the right hand anchor point; take the midpoint of the left hand anchor point or the right hand anchor point as the hand anchor point.
2. The method for recognizing the position of both hands according to claim 1, characterized in that, Step S21 includes: Step S211: Determine the first face and the first hand based on whether the hand is in front of the face; Step S212: Remove faces and hands whose facial normal vectors form an angle smaller than the first angle with the display, to obtain a second face and a second hand; Step S213: Compare the second hand with the preset hand posture and determine the hand with the highest similarity as the valid hand.
3. The method for recognizing the position of both hands according to claim 2, characterized in that, Step S23 includes: Step S231: If the hand state is the first state and the number of protruding fingers is 1, then take the key point of the fingertip of the protruding finger as the left hand anchor point or the right hand anchor point. Step S232: If the hand state is the first state and the number of protruding fingers is greater than 1, then take the center point of the key point of the fingertip of the protruding finger as the left hand anchor point or the right hand anchor point. Step S233: If the hand state is the second state, then the center point of at least some key points of the effective hand is determined as the left hand anchor point or the right hand anchor point; Step S234: Determine the midpoint of the left-hand anchor point or the right-hand anchor point as the hand anchor point.
4. The method for recognizing the position of both hands according to claim 1, characterized in that, The weighted center point c has a correction parameter ε relative to the actual position of the mouse on the display; wherein, the correction parameter ε is related to the position of the weighted center point c on the display.
5. A two-hand position recognition system for implementing the two-hand position recognition method according to any one of claims 1 to 4, characterized in that, include: The acquisition module is used to obtain 3D information of the left hand, right hand, and eyes using a depth camera; A determination module is used to determine the position of the eyeballs, the center point of the face, and the hand anchor points; wherein the hand anchor points are determined by the left hand gesture, the right hand gesture, and the relative positions of the left and right hands; The connection module is used to calculate the line connecting the eyeball position and the hand anchor point, and intersect the display at position a; calculate the line connecting the face center point and the hand anchor point, and intersect the display at position b; The calculation module is used to assign weight values α and β to position a and position b respectively based on the hand anchor point and the distance of the eye from the display, and to calculate the weighted center point c of position a and position b; where α+β=1; In the determining module, the positions of the left and right pupils are determined respectively, which are used as the eyeball positions; In the connection module, a line is drawn from the position of the left pupil to the hand anchor point, intersecting the display at position a1, and assigned a weight value μ; a line is drawn from the position of the right pupil to the hand anchor point, intersecting the display at position a2, and assigned a weight value ν; and the weighted center point a of positions a1 and a2 is calculated; where μ + ν = 1, and the values of μ and ν are related to the distance difference between the left pupil and the right pupil and the display. Determining the hand anchor point includes: Step S21: Determine the valid hand based on the correspondence and positional relationship between the left hand, right hand and face; Step S22: Identify the key points of the effective hand and determine the hand state based on the distribution of the key points; Step S23: If the hand state is the first state, then take some key points of the effective hand as the left hand anchor point or the right hand anchor point; if the hand state is the second state, then take the center point of at least some key points of the effective hand as the left hand anchor point or the right hand anchor point; take the midpoint of the left hand anchor point or the right hand anchor point as the hand anchor point.
6. A two-hand position recognition device, characterized in that, include: processor; A memory in which executable instructions of the processor are stored; The processor is configured to perform the steps of the hand position recognition method according to any one of claims 1 to 4 by executing the executable instructions.
7. A computer-readable storage medium for storing a program, characterized in that, When the program is executed, it implements the steps of the two-hand position recognition method according to any one of claims 1 to 4.