Hand tracking method and apparatus, and electronic device and storage medium

By constructing a feature vector based on the hand key point model and selecting the motion state area with the highest similarity as the tracking target, the problem of additional limitations in multi-objective tracking is solved, and efficient and accurate gesture tracking is achieved in complex environments.

WO2025138522A1PCT designated stage expired Publication Date: 2025-07-03SHENZHEN HONGHE INNOVATION INFORMATION TECH CO LTD
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/091743
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-27
Filing Date
2024-05-08
Publication Date
2025-07-03

AI Technical Summary

Technical Problem

In complex interactive environments, in multi-objective tracking gesture recognition, the prior art requires additional limitations to affect the degree of freedom of gesture interaction.

Method used

By obtaining the current frame, performing hand object detection, selecting the moving state target area with the highest similarity between the first hand feature vector and the tracking feature vector of the tracking target as the tracking result, and the feature vector is constructed based on the hand key point model.

Benefits of technology

Effectively reduce interference from stationary hands, quickly find tracking targets, realize movement tracking of specific hands, avoid additional restrictions, and improve tracking accuracy in multi-handed environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024091743_03072025_PF_FP_ABST
    Figure CN2024091743_03072025_PF_FP_ABST
Patent Text Reader

Abstract

The present application is applicable to the technical field of image processing. Provided are a hand tracking method and apparatus, and an electronic device and a storage medium. The method comprises: acquiring a current frame; performing target detection on a hand in the current frame, so as to obtain a plurality of target areas in the current frame; and from among the plurality of target areas in the current frame, selecting a target area, which is in a motion state and has the highest similarity between a first hand feature vector and a tracking feature vector of a tracking target, as a tracking result of the tracking target in the current frame, wherein the first hand feature vector is constructed on the basis of a hand key point model of the target area.
Need to check novelty before this filing date? Find Prior Art

Description

Hand tracking method and device, electronic device and storage medium

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS

[0002] This application claims priority to Chinese patent application No. 202311833403.X filed on December 27, 2023, the entire contents of which are incorporated herein by reference. Technical Field

[0003] The present application belongs to the field of image processing technology, and in particular relates to a hand tracking method and device, an electronic device, and a storage medium. Background Art

[0004] In recent years, with the development of computer vision technology, gesture interaction, that is, remote control of devices through gestures, has become an important new development direction in human-computer interaction. To achieve this function, accurate gesture recognition is necessary.

[0005] Gesture recognition generally uses computer vision technology to track the hand in a video containing the hand, that is, a sequence of images arranged in time, to obtain the hand's movement trajectory, and then identify the specific type of gesture in the video based on the movement trajectory.

[0006] In complex interactive environments, multiple hands may appear in the captured video at the same time, and some or all of these hands may be targets that need to be tracked. However, for the tracked target, other hands may interfere with it, making how to quickly and accurately achieve multi-target tracking of hands a difficult problem that needs to be solved.

[0007] Related technologies often distinguish hands by wearing iconic accessories, or by directly drawing an area in the image to inform the user that there is only one hand in the area. These methods all bring obvious additional restrictions and affect the freedom of gesture interaction.

[0008] Summary of the Invention

[0009] The embodiments of the present application provide a hand tracking method, an electronic device, and a storage medium, which can solve the problem in the related art that multi-target tracking of a hand requires additional restrictions.

[0010] In a first aspect, an embodiment of the present application provides a hand tracking method, which includes: obtaining a current frame; performing target detection on the hand in the current frame to obtain multiple target areas in the current frame; among the multiple target areas in the current frame, selecting a first hand feature vector having the highest similarity to a tracking feature vector of a tracking target and being a target area in a motion state as the tracking result of the tracking target in the current frame, wherein the first hand feature vector is constructed based on a hand key point model of the target area.

[0011] In second aspect, an embodiment of the present application provides a hand tracking device, which includes: an acquisition module for acquiring a current frame; a detection module for performing target detection on the hand in the current frame to obtain multiple target areas in the current frame; a tracking module for selecting, from multiple target areas in the current frame, a first hand feature vector having the highest similarity to a tracking feature vector of a tracking target, and a target area in a motion state as the tracking result of the tracking target in the current frame, wherein the first hand feature vector is constructed based on a hand key point model of the target area.

[0012] In a third aspect, an embodiment of the present application provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable by the processor, wherein the processor implements the hand tracking method described in the first aspect above when executing the computer program.

[0013] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the hand tracking method described in the first aspect above.

[0014] In a fifth aspect, an embodiment of the present application provides a computer program product. When the computer program product is run on an electronic device, the electronic device executes the hand tracking method described in the first aspect above.

[0015] Compared with the related art, the beneficial effects of the embodiments of the present application are: by obtaining the current frame; performing target detection on the hand in the current frame to obtain multiple target areas in the current frame; among the multiple target areas in the current frame, selecting the first hand feature vector with the highest similarity to the tracking feature vector of the tracking target, and the target area in the motion state as the tracking result of the tracking target in the current frame, wherein the first hand feature vector is constructed based on the hand key point model of the target area, and adding the motion state restriction can effectively reduce the interference of the stationary hand on the tracking target, and can quickly find the tracking result of the tracking target in the current frame while ensuring accuracy, so that no additional restrictions are required, thereby realizing motion tracking of a specific one or more hands in a complex environment with multiple hands. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the embodiments or related technical descriptions. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0017] FIG1 is a schematic structural diagram of an electronic device provided in one embodiment of the present application;

[0018] FIG2 is a flow chart of a hand tracking method provided in an embodiment of the present application;

[0019] FIG3 is a schematic diagram of a process of constructing a first hand feature vector for a target area in a hand tracking method provided in an embodiment of the present application;

[0020] FIG4 is an example diagram of a hand key point model in a hand tracking method provided in an embodiment of the present application;

[0021] FIG5 is a schematic diagram of a flow chart of determining whether each target area is in a motion state in a hand tracking method provided in an embodiment of the present application;

[0022] FIG6 is a schematic diagram of a specific process of S26 in FIG5 ;

[0023] FIG7 is a schematic diagram of a specific flow chart of a hand tracking method provided in an embodiment of the present application;

[0024] FIG8 is a schematic structural diagram of a hand tracking device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0025] In the following description, specific details such as specific system structures and techniques are provided for purposes of illustration rather than limitation to facilitate a thorough understanding of the embodiments of the present application. However, it will be apparent to those skilled in the art that the present application may be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid obscuring the description of the present application with unnecessary detail.

[0026] It should be understood that when used in the present specification and the appended claims, the term "comprising" indicates the presence of described features, integers, steps, operations, elements and / or components, but does not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or groups thereof.

[0027] It will also be understood that the term "and / or" used in this specification and the appended claims refers to and includes any and all possible combinations of one or more of the associated listed items.

[0028] As used in this specification and the appended claims, the term "if" can be interpreted as "when" or "upon" or "in response to determining" or "in response to detecting," depending on the context. Similarly, the phrase "if it is determined" or "if [described condition or event] is detected" can be interpreted as meaning "upon determination" or "in response to determining" or "upon detection of [described condition or event]" or "in response to detecting [described condition or event]," depending on the context.

[0029] In addition, in the description of the present application specification and the appended claims, the terms "first", "second", "third", etc. are only used to distinguish the descriptions and cannot be understood as indicating or implying relative importance.

[0030] References to "one embodiment" or "some embodiments" in this specification mean that a particular feature, structure, or characteristic described in conjunction with that embodiment is included in one or more embodiments of the present application. Thus, phrases such as "in one embodiment," "in some embodiments," "in other embodiments," and "in other embodiments" appearing in various places in this specification do not necessarily refer to the same embodiment, but rather mean "one or more but not all embodiments," unless otherwise specifically emphasized. The terms "including," "comprising," "having," and variations thereof all mean "including but not limited to," unless otherwise specifically emphasized.

[0031] The hand tracking method provided in the embodiments of this application can be applied to electronic devices, including but not limited to servers, server clusters, mobile phones, tablet computers, laptop computers, desktop computers, personal digital assistants, wearable devices, and other electronic devices with computing capabilities. The embodiments of this application do not impose any restrictions on the specific type of electronic device.

[0032] FIG1 is a block diagram showing a partial structure of an electronic device provided in accordance with an embodiment of the present application. Referring to FIG1 , the electronic device includes: a processor 10, a memory 20, a bus 30, an input device 40, an output device 50, and a communication device 60. The processor 10 and the memory 20 are connected to each other via the bus 30, and the input device 40, the output device 50, and the communication device 60 are also connected to the bus 30. It will be understood by those skilled in the art that the structure of the electronic device shown in FIG1 does not constitute a limitation of the electronic device, and may include more or fewer components than shown, or combine certain components, or arrange the components differently.

[0033] The following is a detailed introduction to the various components of the electronic device with reference to FIG1 :

[0034] The processor 10 is the control center of the electronic device, which can run the programs stored in the memory 20 to perform various functions and process data. The processor 10 can be a central processing unit (CPU), and the processor 10 can also be other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc. In some embodiments, the processor 10 may include an AI (Artificial Intelligence) processor, which is used to process computing operations related to machine learning.

[0035] The memory 20 is used to store an operating system, an application program, a boot loader (BootLoader), data, and other programs, such as the program code of a computer program. The memory 20 can also be used to temporarily store data required for and generated by executing the program. The memory 20 may include a high-speed random access memory, and may also include a non-volatile memory, such as a flash memory, a hard disk, a multimedia card, a card-type memory, etc. The memory 20 may include a storage unit provided inside the electronic device, such as the hard disk of the electronic device, and / or a removable external storage unit, such as a mobile hard disk, a USB flash drive, a smart memory card (Smart Media Card, SMC), a secure digital (Secure Digital, SD) card, etc.

[0036] The input device 40 may include at least one of a keyboard, a mouse, a touch panel, a joystick, etc., and is used to collect user input operations to generate corresponding input signals.

[0037] The output device 50 is used to output information to be provided to the user. The output device 50 generally includes a display. Optionally, a liquid crystal display (LCD), an organic light-emitting diode (OLED), etc. can be used. In addition, the output device can further include a speaker.

[0038] The communication device 60 may include a modem, a network card, etc., and is used to establish a network connection with other electronic devices and communicate with each other.

[0039] The hand tracking method provided in the embodiment of the present application can be implemented as a computer software program. For example, the embodiment of the present application provides a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program contains program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from the network through the communication device 60, and / or installed from a removable external storage unit. When the computer program is executed by the processor 10, the various functions defined in the hand tracking method provided in the embodiment of the present application are implemented.

[0040] FIG2 shows a schematic flowchart of a hand tracking method provided in an embodiment of the present application. As an example and not a limitation, the method can be applied to the above-mentioned electronic device, including the following S1-S4.

[0041] S1: Get the current frame.

[0042] The current frame may be a three-dimensional image, including depth information of objects in the environment, that is, distance information between the objects and the image sensor that collects the depth information.

[0043] Image sensors that collect depth information are called depth cameras and can use binocular stereo vision, structured light, time-of-flight, radar ranging, and other methods to collect depth information. Depth information can be stored and used as point clouds or depth images, and the two can be converted into each other. A point cloud contains the coordinates of multiple 3D points, each representing a point on the surface of an object in the environment. A depth image is a 2D image in which the pixel value of each pixel represents the depth information of that pixel.

[0044] In addition to image sensors that capture depth information, a common visible light camera can also be used to capture a conventional planar image. This planar image can then be fused with the directly or indirectly acquired depth image to create an RGB-D image. Each pixel in an RGB-D image has four channels: R, G, B, and D. The D channel's pixel value represents the pixel's depth information. Combined with the pixel's planar coordinates in the image, the corresponding three-dimensional coordinates can be calculated.

[0045] The current frame may be a point cloud, an RGB-D image, or a 3D model obtained by 3D reconstruction of the point cloud / RGB-D image, without limitation.

[0046] S2: Perform target detection on the hand in the current frame to obtain multiple target areas in the current frame.

[0047] To simplify the description, except for the parts specifically mentioned below, the target area refers to the target area in the current frame.

[0048] Target detection can be performed on the current frame itself, or on an intermediate image used to produce the current frame, such as a point cloud or a plane image. If the target detection object is a plane image or an RGB-D image, then a target area is represented by a rectangular area in the current frame, and the target detection algorithm determines that a hand exists in the rectangular area. If the target detection object is a three-dimensional model, then a target area is represented by a cuboid area in the current frame, and the target detection algorithm determines that a hand exists in the cuboid area. For ease of description, the following explanation will be given using the example of a target area represented by a rectangular area. By analogy, the case where the target area is represented by a cuboid area can be obtained.

[0049] Target detection can be performed using traditional image processing or neural networks. If a neural network is used, the output includes not only the target area but also the confidence level of the target area.

[0050] S3: Among the multiple target regions in the current frame, select the target region whose first hand feature vector has the highest similarity to the tracking feature vector of the tracking target and is in motion as the tracking result of the tracking target in the current frame.

[0051] The tracking target may include part or all of the detected hand. If it is a part, a pre-set rule, such as making a specified gesture with the hand, can be used to indicate the start of tracking. This embodiment mainly discusses how to continue tracking an already identified tracking target and does not discuss how to confirm whether there is a new target.

[0052] Before S3, a first hand feature vector may be constructed for the target area. The first hand feature vector is constructed based on the hand key point model of the target area. As shown in FIG3 , constructing the first hand feature vector for the target area may specifically include the following steps S21-S23.

[0053] S21: Perform hand key point detection on the target area to obtain a hand key point model of the target area.

[0054] Because the target region only represents an area surrounding the hand and cannot represent the specific shape of the hand, it is necessary to extract the specific hand information for subsequent use. To reduce resource consumption, hand keypoint detection is generally performed on the target region. The specific area occupied by the hand in the target region is converted into a small number of hand keypoints, forming a hand keypoint model for subsequent use. An example of a hand keypoint model is shown in Figure 4.

[0055] S22: Construct a first hand feature vector of the target area according to the distance between adjacent key points in the hand key point model.

[0056] Still referring to the example given in Figure 4, the first hand feature vector is a 21-dimensional vector, including the distance from point 0 to point 1, the distance from point 1 to point 2, the distance from point 2 to point 3, the distance from point 3 to point 4, the distance from point 0 to point 5, the distance from point 5 to point 6, the distance from point 6 to point 7, the distance from point 7 to point 8, the distance from point 5 to point 9, the distance from point 9 to point 10, the distance from point 10 to point 11, the distance from point 11 to point 12, the distance from point 9 to point 13, the distance from point 13 to point 14, the distance from point 14 to point 15, the distance from point 15 to point 16, the distance from point 13 to point 17, the distance from point 0 to point 17, the distance from point 17 to point 18, the distance from point 18 to point 19, and the distance from point 19 to point 20, a total of 21 components.

[0057] S23: Correct the first hand feature vector using the reference vector.

[0058] Optionally, a reference vector can be obtained based on statistical data, experimental data, theoretical models, etc. A first processing is performed on the reference vector and the first hand feature vector. Specifically, at least one of the reference vector and the first hand feature vector is scaled so that the sum of their components, also referred to as the total length, is the same. Since the first hand feature vector is used later, for ease of description, this example illustrates scaling only the reference vector. In actual applications, the scaled vector may be modified.

[0059] The sum of the components of the scaled reference vector B is the same as the sum of the components of the first hand feature vector. The components in the scaled reference vector are compared one by one with the components in the first hand feature vector to obtain the comparison results of each component. Here, the comparison result of a certain component can be the absolute value of the difference, the ratio, etc. of the component in the two vectors. When the comparison result of a certain component indicates that the difference between the two vectors is large, for example, the absolute value of the difference exceeds the preset value, the difference between the ratio and 1 exceeds the set range, etc., the component in the reference vector is used to correct the component in the first hand feature vector to reduce the impact of possible model abnormalities on subsequent calculations, otherwise no correction is performed. For example, for the component that needs to be corrected, the weighted average or arithmetic mean of the component in the reference vector and the component in the first hand feature vector can be calculated, and then the component in the first hand feature vector can be replaced with the calculated average value.

[0060] The distance between the first hand feature vector and the tracking feature vector can be calculated to represent the similarity between the two. The distance can include cosine distance, Euclidean distance, Hamilton distance, etc. If the selected distance is affected by the modulus value of the vector itself, the first hand feature vector can be normalized in advance, that is, the first hand feature vector can be scaled to modify its modulus value to a uniform value, such as 1. Taking the cosine distance as an example, the larger the cosine distance, the higher the similarity. The target area with the largest cosine distance between the first hand feature vector and the tracking feature vector and in motion can be selected as the tracking result of the tracking target in the current frame.

[0061] Before S3, it is possible to first determine whether each target region is in motion. This process and the aforementioned process of constructing the first hand feature vector are not restricted in order and can be performed in parallel. When determining whether each target region is in motion first and then constructing the first hand feature vector, to save resources, it is possible to select to construct the first hand feature vector only for the selected target regions in motion.

[0062] As shown in FIG5 , determining whether each target area is in motion may specifically include the following steps S25 - S26 .

[0063] S25: Match each target region with multiple target regions in the previous frame according to the principle of the highest intersection-over-union ratio, and obtain matching results for each target region.

[0064] Intersection over Union (IoU), also known as overlap, is a metric used to evaluate the degree of overlap between two regions in an image. It is calculated as the ratio of the intersection to the union of the two regions, with a value range of [0, 1]. Ideally, the two regions completely overlap, and the IoU reaches its maximum value of 1.

[0065] Specifically, for a target region A in the current frame, its intersection-over-union (IoU) with each target region in the previous frame can be calculated, and the target region with the largest IoU in the previous frame is selected as the matching result for target region A. The above steps are performed for each target region in the current frame to complete the matching of each target region in the current frame.

[0066] S26: Determine whether the target area is in motion based on the target area and its matching result.

[0067] Optionally, as shown in FIG6 , this step may specifically include the following parts S261 - S264 .

[0068] S261: Determine whether the target area moves based on the intersection-over-union ratio between the target area and its matching results.

[0069] Optionally, when the intersection-and-union ratio between the target area and its matching result is less than a first threshold, it means that there is little overlap between the target area and its matching result, then the position of the target area has changed significantly between the current frame and the previous frame, and it can be determined that the target area has moved; when the intersection-and-union ratio between the target area and its matching result is greater than or equal to the first threshold, it is determined that the target area has not moved.

[0070] When the target area moves, jump to S263; when the target area does not move, since the target area only represents an area surrounding the hand, it itself cannot represent the specific shape of the hand therein. There is a possibility that the target area does not move while the hand therein changes shape. Therefore, it is possible to further judge whether the shape of the hand has changed to improve the accuracy of the motion state judgment, thereby improving the accuracy of hand tracking, and jump to S262.

[0071] S262: Determine whether the target area has a shape change based on the target area and the hand key point model of the matching result.

[0072] Specifically, a second hand feature vector for the target area can be constructed based on the three-dimensional coordinates of each key point in the hand key point model of the target area, and a second hand feature vector for the matching result can be constructed based on the three-dimensional coordinates of each key point in the hand key point model of the matching result. To remove the impact of the displacement between the target area and the matching result on subsequent calculations, each hand key point model is generally normalized first, that is, the world coordinate system is translated, and the origin after translation is the specified key point in the hand key point model, such as point 0 in Figure 4. After normalization, the coordinates of the specified key point in each second hand feature vector are all (0,0,0).

[0073] When the similarity between the target region and the second hand feature vector of the matching result is less than a second threshold, it is determined that the target region has undergone a shape change. When the similarity between the target region and the second hand feature vector of the matching result is greater than or equal to the second threshold, it is determined that the target region has not undergone a shape change.

[0074] Similarly, the distance between the target region and the second hand feature vector of the matching result can be calculated to represent the similarity between the target region and the matching result. The distance can include cosine distance, Euclidean distance, Hamilton distance, etc. If the selected distance is affected by the modulus value of the vector itself, the second hand feature vector can be normalized in advance, that is, the second hand feature vector can be scaled to modify its modulus value to a uniform value, such as 1.

[0075] When the target area changes in shape, the process jumps to S263 ; when the target area does not change in shape, the process jumps to S264 .

[0076] S263: Determine that the target area is in motion.

[0077] S264: Determine that the target area is in a stationary state.

[0078] S4: Update the tracking feature vector using the tracking results.

[0079] Specifically, a weighted average of the tracking result and the tracking feature vector can be calculated, and then the tracking feature vector can be updated to the weighted average for use in the next round of calculations. When a neural network is used for target detection, the weight of the tracking result can be the confidence level output by the neural network. When traditional image processing is used for target detection, the weight of the tracking result can be a fixed value or a confidence level calculated using traditional image processing methods.

[0080] Through the implementation of this embodiment, among multiple target regions in the current frame, the target region whose first hand feature vector has the highest similarity to the tracking feature vector of the tracking target and is in motion is selected as the tracking result for the tracking target in the current frame. Adding the motion state restriction effectively reduces the interference of a stationary hand on the tracking target, while ensuring accuracy, and can quickly find the tracking result for the tracking target in the current frame, eliminating the need for additional restrictions and achieving motion tracking of one or more specific hands in complex environments with multiple hands.

[0081] The specific process of the hand tracking method is described below with reference to the accompanying drawings.

[0082] As shown in Figure 7, the hand tracking method provided in one embodiment of the present application specifically includes the following parts S31 to S43. This embodiment is a specific extension of the above embodiment, wherein the same / corresponding parts are not repeated.

[0083] S31: Get the current frame.

[0084] S32: Perform target detection on the hand in the current frame to obtain multiple target areas in the current frame.

[0085] S33: Perform hand key point detection on each target area to obtain a hand key point model of each target area.

[0086] S34: Constructing a first hand feature vector for each target area according to the distance between adjacent key points in the hand key point model of each target area.

[0087] S35: Using the reference vector, try to correct each first hand feature vector.

[0088] S36: Match each target region with multiple target regions in the previous frame according to the principle of the highest intersection-over-union ratio, and obtain matching results for each target region.

[0089] S37: Determine whether the target area moves according to the intersection-over-union ratio between the target area and its matching result.

[0090] When the intersection-over-union ratio between the target area and its matching result is less than the first threshold, it is determined that movement has occurred and the process jumps to S40. When the intersection-over-union ratio between the target area and its matching result is greater than or equal to the first threshold, it is determined that no movement has occurred and the process jumps to S38.

[0091] S38: Constructing second hand feature vectors of the target area and the matching result respectively according to the three-dimensional coordinates of each key point in the hand key point model of the target area and the matching result.

[0092] S39: Determine whether the target region has undergone a shape change based on the similarity between the target region and the second hand feature vector of the matching result.

[0093] When the similarity between the second hand feature vectors of the target area and its matching result is less than the second threshold, it is determined that a shape change has occurred and the process jumps to S40; when the similarity between the second hand feature vectors of the target area and its matching result is greater than or equal to the second threshold, it is determined that no shape change has occurred and the process jumps to S41.

[0094] S40: Determine that the target area is in motion.

[0095] S41: Determine that the target area is in a stationary state.

[0096] It is necessary to execute the S37-S41 portion outlined by the dotted line in the figure for each target area to determine whether each target area is in a motion state, and jump to S42 after completion.

[0097] S42: Among the multiple target regions in the current frame, select a target region whose first hand feature vector has the highest similarity to the tracking feature vector of the tracking target and is in motion as a tracking result of the tracking target in the current frame.

[0098] S43: Update the tracking feature vector using the tracking result.

[0099] FIG8 shows a schematic structural diagram of a hand tracking device provided in an embodiment of the present application. The hand tracking device includes an acquisition module 11 , a detection module 12 and a tracking module 13 .

[0100] The acquisition module 11 is used to acquire the current frame.

[0101] The detection module 12 is configured to perform target detection on the hand in the current frame to obtain a plurality of target regions in the current frame.

[0102] The tracking module 13 is used to select, from among the multiple target areas in the current frame, a target area whose first hand feature vector has the highest similarity with the tracking feature vector of the tracking target and is in a moving state as the tracking result of the tracking target in the current frame, wherein the first hand feature vector is constructed based on a hand key point model of the target area.

[0103] It should be noted that the information interaction, execution process, etc. between the above-mentioned devices / modules / units are based on the same concept as the method embodiment of this application. Their specific functions and technical effects can be found in the method embodiment section and will not be repeated here.

[0104] It should be understood that the size of the serial numbers of the steps in the above embodiments does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.

[0105] Those skilled in the art can clearly understand that, for the convenience and brevity of description, only the division of the above-mentioned functional units and modules is used as an example for illustration. In actual applications, the above-mentioned functions can be distributed and completed by different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiment can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of software functional units. In addition, the specific names of the functional units and modules are only for the convenience of distinguishing each other, and are not used to limit the scope of protection of this application. The specific working process of the units and modules in the above-mentioned system can refer to the corresponding process in the aforementioned method embodiment, and will not be repeated here.

[0106] An embodiment of the present application further provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps in the above-mentioned various method embodiments are implemented.

[0107] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the present application implements all or part of the process in the above-mentioned embodiment method by instructing the relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, it can implement the steps of each of the above-mentioned method embodiments. The computer program includes computer program code, which can be in source code form, object code form, executable file or some intermediate form. The computer-readable medium can at least include: any entity or device that can carry the computer program code to the camera / electronic device, recording medium, computer memory, read-only memory (ROM), random access memory (RAM), electric carrier signal, telecommunication signal and software distribution medium. For example, a USB flash drive, a mobile hard disk, a magnetic disk or an optical disk. In some jurisdictions, according to legislation and patent practice, computer-readable media cannot be electric carrier signals and telecommunication signals.

[0108] In the above embodiments, the description of each embodiment has its own focus. For parts that are not described or recorded in detail in a certain embodiment, reference can be made to the relevant description of other embodiments.

[0109] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0110] In the embodiments provided in this application, it should be understood that the disclosed devices / network equipment and methods can be implemented in other ways. For example, the device / network equipment embodiments described above are merely illustrative. For example, the division of the modules or units is merely a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.

[0111] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.

[0112] The above-described embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present application, and should all be included in the scope of protection of the present application.

Claims

1. A hand tracking method, the method comprising: Obtaining a current frame; Performing object detection on the hand in the current frame to obtain a plurality of target regions in the current frame; Among the plurality of target regions in the current frame, selecting a target region with the highest similarity between the first hand feature vector and the tracking feature vector of the tracking target and being in a motion state as the tracking result of the tracking target in the current frame, wherein the first hand feature vector is constructed according to the hand key point model of the target region.

2. The method according to claim 1, before selecting, among the plurality of target regions in the current frame, a target region with the highest similarity between the first hand feature vector and the tracking feature vector of the tracking target and being in a motion state as the tracking result of the tracking target in the current frame, further comprising: Matching each of the target regions with the plurality of target regions in the previous frame respectively according to the principle of the highest intersection over union to obtain the matching result of each target region; Judging whether the target region is in a motion state according to the target region and its matching result.

3. The method according to claim 2, wherein The judging whether the target region is in a motion state according to the target region and its matching result includes: Judging whether the target region moves according to the intersection over union between the target region and its matching result; When the target region moves, determining that the target region is in a motion state, and when the target region does not move, judging whether the target region changes in shape according to the hand key point model of the target region and its matching result; When the target region changes in shape, determining that the target region is in a motion state, and when the target region does not change in shape, determining that the target region is in a stationary state.

4. The method according to claim 3, wherein The judging whether the target region moves according to the intersection over union between the target region and its matching result includes: When the intersection over union between the target region and its matching result is less than a first threshold, determining that the target region moves; When the intersection over union between the target region and its matching result is greater than or equal to the first threshold, determining that the target region does not move.

5. The method according to claim 3, wherein The judging whether the target region changes in shape according to the hand key point model of the target region and its matching result includes: Constructing a second hand feature vector of the target region according to the three-dimensional coordinates of each key point in the hand key point model of the target region, and constructing a second hand feature vector of the matching result according to the three-dimensional coordinates of each key point in the hand key point model of the matching result; When the similarity between the second hand feature vectors of the target region and its matching result is less than a second threshold, determining that the target region changes in shape; When the similarity between the second hand feature vectors of the target region and its matching result is greater than or equal to the second threshold, determining that the target region does not change in shape.

6. The method according to claim 1, before selecting, from the multiple target regions in the current frame, the target region with the highest similarity between the first hand feature vector and the tracking feature vector of the tracking target and being in a motion state as the tracking result of the tracking target in the current frame, further includes: Performing hand key point detection on the target region to obtain a hand key point model of the target region; Constructing a first hand feature vector of the target region according to the distances between adjacent key points in the hand key point model; Correcting the first hand feature vector using a reference vector.

7. The method according to claim 1, the method further includes: Updating the tracking feature vector using the tracking result.

8. A hand tracking device, the device includes: An acquisition module, configured to acquire a current frame; A detection module, configured to perform target detection on the hand in the current frame to obtain multiple target regions in the current frame; A tracking module, configured to select, from the multiple target regions in the current frame, the target region with the highest similarity between the first hand feature vector and the tracking feature vector of the tracking target and being in a motion state as the tracking result of the tracking target in the current frame, where the first hand feature vector is constructed according to the hand key point model of the target region.

9. The device according to claim 8, before the device selects, from the multiple target regions in the current frame, the target region with the highest similarity between the first hand feature vector and the tracking feature vector of the tracking target and being in a motion state as the tracking result of the tracking target in the current frame, the device is further configured to match each of the target regions with the multiple target regions in the previous frame respectively according to the principle of the highest intersection over union to obtain the matching result of each target region; Judging whether the target region is in a motion state according to the target region and its matching result.

10. The apparatus according to claim 9, wherein, The device is configured to: Judging whether the target region moves according to the intersection over union between the target region and its matching result; When the target region moves, determining that the target region is in a motion state, and when the target region does not move, judging whether the target region changes in shape according to the hand key point model of the target region and its matching result; When the target region changes in shape, determining that the target region is in a motion state, and when the target region does not change in shape, determining that the target region is in a stationary state.

11. The device according to claim 10, wherein, The device is configured to: When the intersection over union between the target region and its matching result is less than a first threshold, determining that the target region moves; When the intersection over union between the target region and its matching result is greater than or equal to the first threshold, determining That the target region does not move.

12. The device according to claim 10, wherein, The device is configured to: Constructing a second hand feature vector of the target region according to the three-dimensional coordinates of each key point in the hand key point model of the target region, and constructing a second hand feature vector of the matching result according to the three-dimensional coordinates of each key point in the hand key point model of the matching result; When the similarity between the target region and the second hand feature vector of its matching result is less than the second threshold, it is determined that the shape of the target region has changed; When the similarity between the target region and the second hand feature vector of its matching result is greater than or equal to the second threshold, it is determined that the shape of the target region has not changed.

13. The device according to claim 8, before selecting, among the multiple target regions in the current frame, the target region with the highest similarity between the first hand feature vector and the tracking feature vector of the tracking target and being in a motion state as the tracking result of the tracking target in the current frame, the device is configured to perform hand key point detection on the target region to obtain a hand key point model of the target region; Construct the first hand feature vector of the target region according to the distances between adjacent key points in the hand key point model; Use a reference vector to correct the first hand feature vector.

14. The device according to claim 8, the device is configured to update the tracking feature vector using the tracking result.

15. An electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable by the processor, where when the processor executes the computer program, the method according to any one of claims 1 to 7 is implemented.

16. A computer-readable storage medium, the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.

17. A computer program product, when the computer program product runs on an electronic device, enables the electronic device to implement the method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Multi-target tracking method based on region overlap and template matching

    CN109410243A

  • Hand tracking method, device, equipment and computer storage medium

    CN114067426A

  • Automatic thunder vision calibration method based on target tracking

    CN115685102A

  • Tracking method and device of manipulator, equipment and storage medium

    CN115686207A

  • Target object tracking method, related device, equipment and storage medium

    CN116152289A