Finger joint marking method and device thereof, and electronic device
By collecting multi-view image sets for key point detection and responsiveness filtering, the problem of low accuracy in finger joint annotation was solved, achieving higher 3D positioning accuracy and reliability.
Patent Information
- Application Number
- CN202511251694.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-03
- Publication Date
- 2025-11-25
- Estimated Expiration
- 2045-09-03
AI Technical Summary
The accuracy of finger joint annotation in existing technologies is low, which cannot meet the needs of gesture interaction in the XR field.
A collection of multi-view images is acquired at the same time frame. Multiple cameras are used to capture images simultaneously within a common viewing area. Key point detection is performed to generate potential pixel locations and potential bounding boxes. Corner points are extracted and 3D positions are determined based on responsivity. Finally, finger joints are labeled.
By using multi-view fusion and responsiveness filtering, the accuracy of finger joint annotation and spatial positioning reliability are improved, thus enhancing 3D positioning accuracy.
Smart Images

Figure CN120747965B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of computer vision technology, and in particular to a method, apparatus, and electronic device for annotating finger joints. Background Technology
[0002] In the field of XR (Extended Reality), gestures are crucial for interaction. Real-time estimation of the spatial positions of human finger joints is challenging, often heavily reliant on the camera equipment used and the application scenario. Currently available datasets annotated finger joints are insufficient to meet the needs of research and development. Therefore, achieving accurate finger joint annotation is a pressing issue that requires resolution. Summary of the Invention
[0003] This invention provides a method, apparatus, and electronic device for annotating finger joints, which addresses the shortcomings of low accuracy in finger joint annotation in existing technologies and improves the accuracy of 3D finger joint annotation.
[0004] This invention provides a method for marking finger joints, comprising the following steps:
[0005] A collection of multi-view images is acquired within the same time frame; the collection of multi-view images is obtained by multiple cameras simultaneously capturing images within a shared field of view.
[0006] Keypoint detection is performed on the multi-view image set to obtain the potential pixel position of each finger joint, and a potential bounding box centered on the potential pixel position is generated;
[0007] The corner points of each potential bounding box are extracted, and the 3D position of each finger joint is determined based on the responsivity of each corner point; the corner points are image feature points related to the finger joint.
[0008] Each finger joint is labeled according to its 3D position.
[0009] According to a finger joint annotation method provided by the present invention, determining the 3D position of each finger joint based on the responsivity of each corner point includes:
[0010] For each finger joint, determine the number of corner points where the responsivity is greater than the responsivity threshold;
[0011] Based on the number of corner points and camera parameters, determine the 3D position of each finger joint in the reference camera coordinate system.
[0012] According to a finger joint annotation method provided by the present invention, the step of determining the 3D position of each finger joint in a reference camera coordinate system based on the number of corner points and camera parameters includes:
[0013] When the number of corner points is 2, determine the first target camera corresponding to each of the two corner points, and the first pixel position of each of the two corner points;
[0014] Based on the camera parameters of each of the first target cameras, construct the projection matrix of each of the first target cameras;
[0015] Based on the projection matrices and the first pixel positions of the two corner points, a triangulated system of linear equations is established.
[0016] By solving the system of linear equations, the 3D position of each finger joint in the reference camera coordinate system is obtained.
[0017] According to a finger joint annotation method provided by the present invention, the step of determining the 3D position of each finger joint in a reference camera coordinate system based on the number of corner points and camera parameters includes:
[0018] When the number of corner points is greater than 2, determine the second target camera corresponding to each corner point and the second pixel position of each corner point;
[0019] The 3D position of the finger joint in the coordinate system of the reference camera is used as the optimization variable;
[0020] For each of the second target cameras, the projection coordinates of the optimization variables on the image plane of the second target camera are determined based on the optimization variables and the camera parameters of the second target camera.
[0021] Based on the projection coordinates and the positions of the second pixels, a target function is constructed;
[0022] The objective function is solved by nonlinear optimization to obtain the 3D position of each finger joint in the reference camera coordinate system.
[0023] According to a finger joint annotation method provided by the present invention, the step of performing keypoint detection on the multi-view image set to obtain the potential pixel position of each finger joint includes:
[0024] The multi-view image set is input into the key point detection model to detect key points, and the detection results output by the key point detection model are obtained.
[0025] The detection results include the potential pixel position of each finger joint and the relative distance of each finger joint; the relative distance is used to measure the distance between the finger joint and the average joint position.
[0026] According to a finger joint annotation method provided by the present invention, the extraction of corner points of each potential bounding box includes:
[0027] Determine the overlap of potential frames corresponding to the same finger joint point under different shooting angles;
[0028] If the overlap is greater than the overlap threshold, the potential boxes are filtered according to the relative distance of each finger joint to obtain multiple candidate boxes;
[0029] Corner detection is performed on each candidate box to extract the corner points corresponding to each candidate box.
[0030] According to a finger joint annotation method provided by the present invention, determining the overlap degree of potential bounding boxes corresponding to the same finger joint includes:
[0031] The boundaries of each potential box are determined based on the location of each potential pixel.
[0032] Based on the boundaries of each potential box, determine the intersection region and union region corresponding to each potential box;
[0033] The overlap of potential boxes corresponding to the same finger joint is determined based on the intersection region and the union region.
[0034] According to a method for annotating finger joints provided by the present invention, the step of annotating each finger joint according to its 3D position includes:
[0035] Based on the intrinsic parameters of multiple reference cameras, each of the 3D positions is projected onto the image plane of each of the reference cameras to obtain the pixel position corresponding to each of the 3D positions;
[0036] If the pixel position corresponding to the 3D position is on the image plane of the reference camera, then the 3D position is determined to be a valid 3D position.
[0037] Each of the finger joints is labeled based on the effective 3D position.
[0038] The present invention also provides a finger joint marking device, comprising the following modules:
[0039] The acquisition module is used to acquire a set of multi-view images in the same time frame; the set of multi-view images is obtained by multiple cameras simultaneously capturing images within a common viewing area;
[0040] The key point detection module is used to perform key point detection on the multi-view image set, obtain the potential pixel position of each finger joint, and generate a potential bounding box centered on the potential pixel position.
[0041] A 3D position determination module is used to extract the corner points of each potential bounding box and determine the 3D position of each finger joint based on the responsivity of each corner point; the corner points are image feature points related to the finger joint.
[0042] The annotation module is used to annotate each of the finger joints according to the 3D positions.
[0043] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the finger joint annotation method as described above.
[0044] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the finger joint annotation method as described above.
[0045] The present invention also provides a computer program product, including a computer program that, when executed by a processor, implements the finger joint annotation method as described above.
[0046] The finger joint annotation method, apparatus, and electronic device provided by this invention acquire a set of multi-view images within the same time frame; the multi-view image set is obtained by multiple cameras simultaneously capturing images within a shared field of view; keypoint detection is performed on the multi-view image set to obtain the potential pixel position of each finger joint, and a potential bounding box centered on the potential pixel position is generated; corner points of each potential bounding box are extracted, and the 3D position of each finger joint is determined based on the responsivity of each corner point; corner points are image feature points related to the finger joint; each finger joint is annotated based on its 3D position. This invention overcomes the limitations of single-view by multi-view fusion, improving the reliability of spatial positioning. Simultaneously, through potential bounding box and responsivity filtering, it improves feature reliability and 3D positioning accuracy, thereby enhancing the accuracy of 3D finger joint annotation. Attached Figure Description
[0047] To more clearly illustrate the technical solutions in this invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0048] Figure 1 This is an application environment diagram of the finger joint annotation method provided by the present invention.
[0049] Figure 2 This is a flowchart illustrating the finger joint annotation method provided by the present invention.
[0050] Figure 3 This is a schematic diagram of the camera layout provided by the present invention.
[0051] Figure 4 This is a schematic diagram of finger joint markers provided by the present invention.
[0052] Figure 5 This is a schematic diagram of the finger joint marking device provided by the present invention.
[0053] Figure 6 This is a schematic diagram of the structure of the electronic device provided by the present invention. Detailed Implementation
[0054] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.
[0055] The following is combined with Figures 1-6 The present invention describes a finger joint marking method, apparatus, and electronic device.
[0056] The finger joint marking method provided by this invention can be applied to, for example... Figure 1 In the application environment shown, terminal 20 communicates with server 10 via a network. A data storage system can store the data that server 10 needs to process; this system can be integrated onto server 10 or located in the cloud or on another network server. Terminal 20 can acquire a set of multi-view images within the same time frame; perform keypoint detection on the multi-view image set to obtain the potential pixel position of each finger joint and generate a potential bounding box centered on the potential pixel position; extract the corner points of each potential bounding box, and determine the 3D position of each finger joint based on the responsivity of each corner point; and annotate each finger joint based on its 3D position, thereby improving the efficiency and accuracy of 3D finger joint annotation. Alternatively, terminal 20 can send the acquired multi-view image set to server 10, where server 10 determines the 3D position of each finger joint, and server 10 returns the 3D position of each finger joint to terminal 20, which then annotates each finger joint based on its 3D position.
[0057] It should be noted that the execution subject of the finger joint annotation method provided by the present invention is a server or computer device, such as a mobile phone, tablet computer, laptop computer, handheld computer, vehicle electronic device, wearable device, etc.
[0058] Figure 2 This is a flowchart illustrating the finger joint annotation method provided by the present invention, as shown below. Figure 2 As shown, the method includes the following:
[0059] Step 101: Collect a set of multi-view images at the same time frame.
[0060] The same time frame refers to images captured by multiple cameras having completely identical timestamps or minimal errors. The purpose of acquiring images within the same time frame is to avoid motion blur or positional shifts caused by time differences in capture time (for example, when a finger moves, asynchronous images will record different postures, resulting in ghosting or errors in subsequent 3D reconstruction). This can be achieved through hardware synchronization (such as camera trigger cable connections) or software synchronization, such as GPS (Global Positioning System) timestamps or NTP (Network Time Protocol), ensuring a consistent time base across multiple cameras.
[0061] A multi-view image set is obtained by simultaneously capturing images from multiple cameras within a shared field of view. Specifically, it is a dataset consisting of images taken by at least two or more cameras, with each camera corresponding to a viewpoint. It should be understood that the differences in viewpoints arise from different camera placements (e.g., different orientations around the subject) and different shooting angles (e.g., front, side, or obliquely above). For example, using eight cameras to capture images of the same finger from the front, side, and obliquely behind the left hand, the eight images together form a multi-view image set.
[0062] The common field of view is the area where the fields of view of all cameras overlap, meaning the subject (e.g., a finger) needs to appear in the images of every camera simultaneously. In each camera's image, the same target (e.g., a finger joint) must be clearly visible, and its position within the valid area (not a blurred edge area). It should be understood that only targets within the common field of view can have their 3D position calculated using the parallax of multi-view images.
[0063] Synchronous shooting refers to the combined requirements of the same time frame and shared viewing area: multiple cameras capture images of a target within the same shared viewing area at the same time. Synchronization (time) and shared viewing (space) are both indispensable; if they are not synchronized, the target's movement will result in different states at different times; if there is no shared viewing, the target will be invisible in some cameras, making multi-view cross-verification impossible.
[0064] In the same time frame, multiple cameras simultaneously capture images of a target (such as a hand) within a shared field of view, resulting in multiple images. These multiple images are then combined to form a multi-view image set.
[0065] In one embodiment, images are acquired using eight cameras to generate a multi-view image set. (Reference) Figure 3 Eight cameras are placed within a 3m × 1.5m environment. These cameras require time-stamp synchronization during recording. The system needs a master synchronization controller to generate square wave pulse signals with the same frequency as the recording frequency. These pulse signals are connected to all cameras via a star-shaped topology signal trigger line. All cameras will operate in external trigger mode, triggering recording via the pulse signal. For photos captured by the same pulse signal, the acquisition base station will uniformly add a timestamp based on either the rising or falling edge of the pulse signal. The selected camera module model must have external trigger recording capability, and to ensure a sufficiently wide field of view, the FOV (Field of View) must be at least 120 degrees. Cameras 1 and 2 will simulate glasses cameras, with a baseline distance of approximately 13cm. Larger baseline spaces can be reserved between cameras in other locations to achieve greater parallax.
[0066] For a properly configured camera synchronization system, the intrinsic and extrinsic parameters of the cameras need to be calibrated. This can be done using tools such as Matlab stereo calibration tools or Kalibr (an open-source calibration tool). Due to viewing angle limitations, multiple cameras can be calibrated simultaneously using a single set of calibrations. The calibration results yield the intrinsic parameters of each camera and its extrinsic parameters relative to other cameras. For example, a fixed method can be used for extrinsic parameter transfer, with the specific transfer and calibration sequence as follows:
[0067] 1) Record data on the movement of the calibration board within the common field of view of cameras 1, 2, 3, and 4, and obtain the intrinsic and extrinsic parameters of cameras 1, 2, 3, and 4 using calibration tools. External reference External reference ;
[0068] 2) Record data on the movement of the calibration board within the common field of view of cameras 1, 7, and 8, and obtain the intrinsic and extrinsic parameters of cameras 1, 7, and 8 using calibration tools. External reference ;
[0069] 3) Record data on the movement of the calibration board within the common field of view of cameras 3, 4, 5, and 6, and obtain the intrinsic and extrinsic parameters of cameras 3, 4, 5, and 6 using calibration tools. External reference External reference ;
[0070] 4) The transformation relationship between cameras 5 and 6 and camera 1 was calculated using parameter passing. The specific calculation method is as follows:
[0071] ;
[0072] ;
[0073] in, Indicates from Camera coordinate system to The transformation relationship of the camera coordinate system is physically significant because it transforms the coordinate system of the camera into a coordinate system that is 10 ... Points in the camera coordinate system, transformed to The transformation in the camera coordinate system, where... , .
[0074] Currently, many hand image acquisition solutions require the gesture collector to wear special gloves. However, gloves alter the appearance of the hand, such as obscuring real skin texture (natural features like palm lines and knuckle creases), covering real skin tone, and reflecting light characteristics (skin is flesh-colored and changes in brightness with lighting; gloves have different colors and materials that reflect light differently). Therefore, developing gesture detection algorithms using images of gloved hands differs significantly from reality. The model parameters cannot reflect the true condition of a real hand because the distribution of training data (gloved hands) differs greatly from that of actual application scenarios (real, bare hands). The parameters tuned using "fake hand data" cannot adapt to the characteristics of a real hand, leading to errors during detection (such as missing joints or misjudging gestures), ultimately resulting in insufficient accuracy.
[0075] Based on this, in this embodiment of the invention, the collector affixes black or white marker dots to the finger joints, with the radius of the marker dots being less than or equal to 1 mm. The positions of the markers are as follows: Figure 4 As shown, the marker is placed on the back of the hand. Figure 4In this text, "y:up" indicates the y-axis upwards, "x:right" indicates the x-axis to the right, "z:back" indicates the z-axis backwards, "Thumb" represents the thumb, "Index" represents the index finger, "Middle" represents the middle finger, "Ring" represents the ring finger, "Little" represents the little finger, "Tip" represents the fingertip, "Distal" represents the distal phalanx (such as the joint corresponding to the distal phalanx, close to the fingertip), "Intermediate" represents the intermediate phalanx (present in some fingers, such as the middle joint of the middle finger), "Proximal" represents the proximal phalanx (the phalanx close to the palm), "Metacarpal" represents the metacarpal joint (the metacarpal joint of the finger in the palm area), "Wrist" represents the wrist joint, and "Palm" represents the palm (marked in the palm area, indicating that the corresponding joint belongs to the palm-related structure).
[0076] It should be understood that the color of the marker (black / white) forms a strong contrast with the skin, making it very easy to identify in the image. Whether it is manually labeled or automatically detected by the algorithm, it can quickly locate the joint position, which is more accurate than finding natural skin texture (such as wrinkles) and greatly shortens the labeling time.
[0077] During the process of capturing images or recording data, the person recording rests their head between cameras 1 and 2, and then extends their hands to make any hand gestures in the space in front of them. At the same time, the eight cameras record data at a frequency of 5Hz per second (the frequency can be set as needed).
[0078] Step 102: Perform key point detection on the multi-view image set to obtain the potential pixel position of each finger joint and generate a potential bounding box centered on the potential pixel position.
[0079] Keypoint detection refers to the process of identifying and locating key feature points with specific semantic meanings in an image using algorithms. It should be understood that keypoints are feature points in an image that play a decisive role in the structure and posture of a target (such as a hand). For example, taking a hand as an example, keypoints include the fingertips and joints (such as the tip of the thumb, the middle joint of the index finger, and the base of the palm; typically, a hand has 26 key joints). The purpose of keypoint detection is to extract the locations of these key feature points from the image, serving as the basis for subsequent analysis (such as 3D reconstruction and pose recognition).
[0080] The potential pixel location is the pixel position (i.e., coordinates) of the keypoint in the image predicted by the model after keypoint detection. However, this location is preliminary and may contain errors; it is not the final, precise location. Due to interference such as blurring, occlusion, and changes in lighting, the model's prediction is not always 100% accurate. For example, if the actual fingertip is at pixel (320, 240), the model may predict it as (318, 239). Therefore, this coordinate is only "potential" (it may be correct, or it may be biased).
[0081] A latent bounding box is a rectangular (or square) region centered on a potential pixel location, used to enclose the actual area where key points might exist. The latent bounding box is typically determined by its center coordinates (i.e., the potential pixel location (u, v)) and dimensions (such as width and height). For example, a box centered at (320, 240) with dimensions of 20×20 pixels covers an area of u±10, v±10 pixels. It should be understood that the purpose of the latent bounding box is to: narrow the search area: subsequent steps (such as corner extraction and exact matching) only need to be processed within the bounding box, without traversing the entire image, thus improving efficiency; and filter interference: exclude irrelevant pixels outside the bounding box (such as background and other objects), reducing the impact of noise on subsequent processing.
[0082] Keypoint detection algorithms or models are used to detect keypoints in a multi-view image set, obtaining the potential pixel positions of each finger joint in each image. For each image, if a hand exists and the potential pixel positions of each finger joint are recorded, then all possible potential bounding boxes for each joint are drawn, centered on the potential pixel positions.
[0083] Step 103: Extract the corner points of each potential box, and determine the 3D position of each finger joint based on the responsiveness of each corner point.
[0084] Corner points are image feature points related to finger joints. Specifically, corner points are pixels in an image where grayscale changes drastically or where two edges intersect. They possess local uniqueness (with significant differences between surrounding pixels) and are often represented by the corners of objects or the intersections of edges. In hand joint scenes, corner points can be inflection points of joint edges (such as the intersection of the contours at the bend of the knuckle) or intersections of skin folds and joint boundaries. These points are strongly correlated with the physical location of the joint.
[0085] It should be understood that hand joints (such as knuckles and wrist joints) are turning points of skin, bones, and tendons. When a camera captures an image, the bending of bones creates a corner in the surface contour (e.g., when a knuckle bends, the skin surface transitions from extension to bending). During joint movement, skin folds intersect with the joint boundary (e.g., when making a fist, the skin folds at the metacarpophalangeal joints intersect with the joint contour line). These physiological structural transitions cause drastic grayscale changes in the image (the grayscale differences between bones, skin, and background are amplified), naturally becoming a "high-response area" for corner detection algorithms. For example, taking the corner at the bend of a knuckle as an example: within the potential bounding box (the area around the joint), the grayscale changes of the corner are significantly different from the surrounding pixels (e.g., at the corner of a joint edge, one side is skin texture, and the other side is the background or joint shadow); even if multiple corners exist, the geometric position and grayscale pattern of each corner are unique (e.g., the corners on both sides of a knuckle are symmetrical in position but have opposite directions of grayscale change), which ensures the uniqueness of corners with the same name (the same physical corner from different perspectives) in multi-view matching.
[0086] It should be understood that the purpose of extracting corner points is to ensure that the 3D reconstruction of finger joints relies on multi-view feature matching: for the same physical finger joint, corresponding feature points (corner points) need to be found in images from different viewpoints in order to calculate the 3D position of the finger joint.
[0087] Corner points of each potential bounding box can be extracted using corner detection algorithms, such as Harris corner detection, SIFT feature points, and Shi-Tomasi corner detection.
[0088] Corner responsivity is a quantitative metric used by corner detection algorithms to measure whether a pixel is a corner and the strength of its corner features. Simply put, responsivity is a confidence metric output by the corner detection algorithm, used to filter out false corners caused by noise interference. For example, the Harris algorithm calculates responsivity by analyzing the grayscale variation trends around a pixel.
[0089] It should be understood that a higher responsivity indicates a more dramatic change in grayscale and a stronger local uniqueness at that corner point. For example, a true corner point at the bend of a finger may have a responsivity of 5000 because the grayscale changes are significant in both directions (one side is skin, and the other side is the joint shadow); slight textures in the background (such as clothing wrinkles) may only have a weak change in one direction, and the responsivity may only be 800. Even if it exceeds the responsivity threshold, it is still considered a "weak corner point".
[0090] After calculating the responsivity of each corner point using a corner detection algorithm, the 3D position of each finger joint is determined based on the responsivity of each corner point, such as 3D coordinate points or 3D spatial coordinates.
[0091] Step 104: Mark each finger joint point according to the 3D position.
[0092] After determining the 3D position of each finger joint, each finger joint is labeled according to its 3D position. Specifically, a one-to-one correspondence is established between the calculated 3D spatial coordinates and the specific finger joint (such as the fingertip of the left index finger and the middle joint of the right middle finger), and the coordinates are recorded or presented in a standardized form.
[0093] It should be understood that the 3D position of each finger joint is a set of spatial coordinates, which does not contain semantic information. The essence of annotation is to bind the 3D coordinates with the semantic name of the finger joint. For example, (50, 30, 200) is marked as the fingertip of the left index finger, ensuring that each finger joint has a unique corresponding 3D coordinate, and that the coordinate is consistent with the physical position of the joint.
[0094] For example, the correspondence between 3D locations and the semantic names of finger joints can be established in the following way:
[0095] 1) Association of 2D pixels from multiple perspectives: During the keypoint detection stage, the joint semantics corresponding to each 2D pixel position have been identified by the model (e.g., the model marks the potential pixel position of "left index fingertip" in the image). The 3D position is calculated based on these 2D pixels, so the semantics of the 3D coordinates can be inferred by "2D-3D projection correspondence". For example, the 3D coordinates calculated from the 2D pixels of the left index fingertip naturally correspond to the left index fingertip.
[0096] 2) Joint topological constraints: Finger joints have fixed connections in space, such as the "proximal index finger joint" being located between the "middle index finger joint" and the palm. If the position of a certain 3D coordinate violates the topological relationship (such as the "fingertip 3D position" being located behind the "proximal joint 3D position"), the corresponding relationship needs to be corrected.
[0097] Optionally, the labeled data of 3D finger joints can provide the algorithm with supervision information on "hand structure and movement patterns" for training models such as 3D gesture detection and joint tracking.
[0098] The finger joint annotation method provided in this invention involves acquiring a multi-view image set within the same time frame; the multi-view image set is obtained by multiple cameras simultaneously capturing images within a shared field of view; keypoint detection is performed on the multi-view image set to obtain the potential pixel position of each finger joint, and a potential bounding box centered on the potential pixel position is generated; corner points of each potential bounding box are extracted, and the 3D position of each finger joint is determined based on the responsivity of each corner point; corner points are image feature points related to the finger joint; and each finger joint is annotated based on its 3D position. This invention overcomes the limitations of single-view by multi-view fusion, improving the reliability of spatial positioning. Simultaneously, through potential bounding box and responsivity filtering, it improves feature reliability and 3D positioning accuracy, thereby enhancing the accuracy of 3D finger joint annotation.
[0099] Based on the above embodiments, step 103 includes:
[0100] For each finger joint, determine the number of corner points where the responsivity is greater than the responsivity threshold;
[0101] Based on the number of corner points and camera parameters, determine the 3D position of each finger joint in the reference camera coordinate system.
[0102] It should be understood that corner points exceeding the responsivity threshold will be initially identified as valid or high-quality corner points, while corner points below the responsivity threshold will be directly filtered out (this could be due to edges, flat areas, or noise). For example, suppose there are two points within the potential bounding box of a finger joint: point A, located at the corner where the knuckle bends, exhibits drastic grayscale changes in both the horizontal and vertical directions, with a responsivity R=6000; and point B, located in a flat skin area near the joint, exhibits only slight texture changes, with a responsivity R=900. If the responsivity threshold is set to 1000, point A will be identified as a high-quality corner point, and point B will be filtered out because point A, with its high responsivity, better represents the physical characteristics of the joint and is suitable for subsequent 3D reconstruction.
[0103] The more corner points with a responsivity greater than the responsivity threshold, the richer the features around the joint (such as wrinkles and clear edges), and the more sufficient the "constraints" for subsequent 3D calculations. For example, 5 corner points provide more projection relationships than 2 corner points, and errors can be reduced through nonlinear optimization (such as least squares). In addition, in multi-view images (such as 8 images / frame), the number of corner points of the same joint reflects the consistency of features under different viewpoints. A large number of corner points with a responsivity greater than the responsivity threshold means that the finger joint has clear features under multiple viewpoints, and the 3D reconstruction is more reliable.
[0104] It should be understood that the core of 3D reconstruction is to infer 3D coordinates from 2D corner points from different perspectives. 3D reconstruction requires at least two 2D points from different perspectives (such as the principle of binocular triangulation) to calculate depth information (Z-axis). Therefore, if the number of corner points with a responsivity greater than the responsivity threshold is less than 2, the gesture captured in the current frame is considered incomplete (e.g., severe target occlusion, motion blur, insufficient lighting), 3D reconstruction cannot be completed, and the calculation for that frame is abandoned.
[0105] The reference camera coordinate system is a 3D coordinate system defined with a certain camera as the reference (the origin is usually at the camera's optical center, the X-axis is to the right, the Y-axis is downward, and the Z-axis points to the scene along the optical axis). The 3D position of all finger joints is described with reference to this coordinate system, which facilitates the unified processing of data from multiple cameras or multiple perspectives.
[0106] Camera parameters include intrinsic and extrinsic parameters. Intrinsic parameters describe the camera's own imaging characteristics, such as focal length, principal point coordinates, and distortion coefficients, which determine how 3D points are projected onto the 2D image plane. Extrinsic parameters describe the camera's position and orientation in the reference coordinate system (rotation matrix R, translation vector T), and are used for coordinate transformation between different camera viewpoints.
[0107] Based on the number of corner points and camera parameters, the 3D position of each finger joint in the reference camera coordinate system is determined. Specifically, using the selected effective corner points (i.e., corner points with a responsivity greater than the responsivity threshold) and camera imaging patterns, the 3D spatial coordinates are inferred from the 2D image features. For example, using the 2D coordinates of the effective corner points and the camera's intrinsic and extrinsic parameters, the 3D position is inferred through a perspective projection model, ultimately unified to the reference coordinate system. It should be understood that if multiple cameras are used for shooting (such as stereo cameras or multi-view arrays), the 3D coordinates calculated by different cameras need to be transformed to the reference camera coordinate system through camera extrinsic parameters to achieve unified annotation.
[0108] The embodiments of this invention ensure the reliability of features by using a high number of corner points with high responsiveness, and achieve accurate 2D to 3D mapping by using camera parameters. The combination of the two reduces noise interference and improves spatial positioning accuracy by using multi-view redundancy. The resulting 3D position annotation is closer to the physical spatial state of real gestures, providing high-quality and high-reliability training data for subsequent gesture detection algorithms.
[0109] Based on the above embodiments, determining the 3D position of each finger joint in the reference camera coordinate system according to the number of corner points and camera parameters includes:
[0110] When the number of corner points is 2, determine the first target camera corresponding to each of the two corner points, and the first pixel position of each of the two corner points;
[0111] Based on the camera parameters of each of the first target cameras, construct the projection matrix of each of the first target cameras;
[0112] Based on the projection matrices and the first pixel positions of the two corner points, a triangulated system of linear equations is established.
[0113] By solving the system of linear equations, the 3D position of each finger joint in the reference camera coordinate system is obtained.
[0114] With two corner points, the 3D position of the joint in the reference camera coordinate system can be recovered using a binocular triangulation algorithm combined with camera extrinsic parameters. It should be understood that the essence of binocular triangulation is to find the unique 3D intersection point by utilizing the intersection of the projected rays from both views. The camera's intrinsic and extrinsic parameters provide spatial direction constraints for the rays, and the optimal solution is ultimately found using linear algebra methods. Specifically, when the number of corner points is two (i.e., the same finger joint point is detected as a high-response corner point in the images from two cameras), the core of recovering the 3D position using the binocular triangulation algorithm is to use the intrinsic and extrinsic parameters of the two cameras and the 2D corner point coordinates to inversely deduce the unique 3D point that satisfies the dual-view projection constraints.
[0115] Determine the 2D pixel coordinates (e.g., pixel coordinates) of the first target camera and the corner points corresponding to the two corner points respectively. Based on the intrinsic and extrinsic parameters of the two first target cameras, construct the projection matrix of the two cameras. The projection matrix describes the projection relationship from 3D point to 2D image point and is composed of the intrinsic and extrinsic parameters. Based on the projection matrix and the 2D pixel coordinates of the corner points, establish a triangulated linear equation system. By solving the linear equation system (e.g., singular value decomposition), obtain the 3D position of each finger joint in the reference camera coordinate system. It should be understood that the essence of the triangulated linear equation system is to minimize the error between the pixel position of the 3D point projected onto the two cameras and the actual detected corner pixel position. This constraint is directly related to the geometric relationship of physical space, rather than empirical speculation; therefore, the calculated 3D position is closer to the actual physical position.
[0116] For example, given the i-th finger joint... High-response corner points (coordinates) were detected in two images (m and n). and ), combined with camera extrinsic parameters (such as (i.e., pose transformation from camera m to camera 1), finger joints are calculated through triangulation. 3D coordinates in the coordinate system of camera #1. For example, Describe the rigid transformation from camera m to camera 1 (including rotation R and translation t), if the corner point of camera m... The corresponding 3D points are (In the coordinate system of camera m), the 3D point in the coordinate system of camera 1 is: Similarly, the corner point of camera n The corresponding 3D points are ,satisfy: Triangulation finds the equations that simultaneously satisfy the projection relationships of both viewpoints by solving these two equations. .
[0117] The embodiments of the present invention utilize the principle of binocular vision to achieve three-dimensional positioning, which not only ensures the calculability of 3D position, but also takes into account both accuracy and efficiency.
[0118] Based on the above embodiments, determining the 3D position of each finger joint in the reference camera coordinate system according to the number of corner points and camera parameters includes:
[0119] When the number of corner points is greater than 2, determine the second target camera corresponding to each corner point and the second pixel position of each corner point;
[0120] The 3D position of the finger joint in the coordinate system of the reference camera is used as the optimization variable;
[0121] For each of the second target cameras, the projection coordinates of the optimization variables on the image plane of the second target camera are determined based on the optimization variables and the camera parameters of the second target camera.
[0122] Based on the projection coordinates and the positions of the second pixels, a target function is constructed;
[0123] The objective function is solved by nonlinear optimization to obtain the 3D position of each finger joint in the reference camera coordinate system.
[0124] The image plane refers to the physical plane inside the camera used to record light signals and convert them into electrical signals. It is the plane where the camera sensor is located and is the carrier for forming a 2D image after an object in 3D space is projected through the lens.
[0125] When the number of corner points is greater than 2 (i.e., multiple corner points of the same finger joint are detected under multiple views), the 3D position can be solved by nonlinear optimization algorithms (such as bundle adjustment). The core logic is: with the goal of minimizing the error between the projection of the 3D point on each camera and the coordinates of the detected 2D corner points, the optimal 3D coordinates are determined through iterative optimization.
[0126] When the number of corner points is greater than 2, determine the second target camera corresponding to each corner point, and the second pixel position (e.g., pixel coordinates) of each corner point; use the 3D position (e.g., 3D coordinates) of the finger joint in the reference camera coordinate system as the optimization variable, denoted as X; for each second target camera, determine the projection coordinates of the optimization variable on the image plane of the second target camera based on the optimization variable X, the intrinsic and extrinsic parameters of the second target camera; construct an objective function based on each projection coordinate and each second pixel position. For example, define the projection error as the difference between the calculated projection coordinates and the detected 2D pixel coordinates of the corner point, and construct the objective function as the sum of squares of all projection errors; use a nonlinear optimization algorithm (e.g., the Levenberg-Marquardt algorithm) to iteratively solve the objective function, update the optimization variable, until the objective function converges to the minimum value. The converged optimization variable is the 3D position of the finger joint in the reference camera coordinate system.
[0127] In one embodiment, a nonlinear optimization method is used for calculation. :
[0128] ;
[0129] in, This represents the 3D position of the finger joint to be solved. The parameter represents the parameter that makes the function reach its minimum value. Indicates the camera index. This represents the set of all cameras that saw the finger joint. Indicates the first Image coordinates of the key points under camera number 1 Indicates from the first Camera No. 1 to No. 1 The transformation relationship of camera number 1 Indicates the first Spatial location of the key points under camera number 1 Indicates the first Camera intrinsic parameter projection model of camera number 1 This represents the square of the single-view projection error.
[0130] The optimization goal is to minimize the sum of squared errors between the coordinates of 3D points projected onto each image and the actual coordinates of the detected corner points, so that the 3D points simultaneously satisfy the projection constraints of all viewpoints.
[0131] The embodiments of this invention utilize multi-view projection constraints, construct an error objective function, and solve nonlinear optimization. The core of this invention is to leverage the redundancy and constraints of multi-source data to further improve the accuracy and robustness of 3D positioning through mathematical optimization.
[0132] Based on the above embodiments, step 102 includes:
[0133] The multi-view image set is input into a keypoint detection model for keypoint detection, and the detection results output by the keypoint detection model are obtained. The detection results include the potential pixel position of each finger joint and the relative distance of each finger joint. The relative distance is used to measure the distance between the finger joint and the average joint position.
[0134] The keypoint detection model is pre-trained and can be a MediaPipe, MegaTrack, or Umetrack model. By learning from a large amount of labeled data, this model can automatically identify hand features (such as contours and joint textures) in images and output structured results.
[0135] A multi-view image set is input into a keypoint detection model for keypoint detection, and the detection results output by the keypoint detection model are obtained. For example, to speed up annotation and achieve fully automatic annotation, for each frame of 8 images, a pre-trained keypoint detection model is used to pre-label potential keypoints in the images. All 8 images of each frame are fed into the model for detection, and the detection results output by the keypoint detection model are obtained. The detection results may include:
[0136] (1) Presence and classification of hands: Determine whether there is a hand in each image, how many hands there are, and whether it is a left or right hand, and record the number of hands and their numbers in each image. The purpose is to filter out images without hands (to avoid invalid calculations later). The left and right hands are distinguished because the semantics of the finger joints are different, such as "left thumb" and "right thumb" corresponding to different 3D positions.
[0137] (2) Preliminary localization of finger joints: If a hand exists, record the potential pixel position of each finger joint. Its function is to provide the initial position for subsequent potential box generation and corner point extraction, and it is a bridge from model prediction to accurate 3D reconstruction.
[0138] (3) Relative distance of each finger joint: The relative distance of each finger joint is defined as the ratio of the difference between the position of each finger joint and the average joint position in the camera coordinate system to the size of the palm.
[0139] ;
[0140] in, Indicates the first The relative distance between the finger joints Indicates the first The spatial coordinates of each finger joint. This indicates the average joint position (the average depth of all joints, representing the depth of the center of the palm). This represents the size of the hand (such as hand width or length, used for normalization). The average joint position is the average coordinate of all joints of the same hand in the 3D camera coordinate system, representing the center reference point of the hand.
[0141] Relative distance is used to measure the distance between a finger joint and the average joint position, that is, to measure the distance between a finger joint and the center of the palm, such as the fingertip being farther than the center of the palm. It is positive; the heel of the palm is closer to the center of the hand. It is negative. Its function is to provide depth constraints for subsequent 3D reconstruction. If the relative distance of the joints predicted by the model does not conform to the human body structure (such as the fingertip being less than the palm base), it can be judged as an incorrect prediction and filtered in advance.
[0142] This invention utilizes a keypoint detection model to detect keypoints in multi-view image sets, thereby accelerating annotation speed, meeting the needs of real-time or large-scale data processing, and improving detection accuracy.
[0143] Based on the above embodiments, step 103 further includes:
[0144] Determine the overlap of potential frames corresponding to the same finger joint point under different shooting angles;
[0145] If the overlap is greater than the overlap threshold, the potential boxes are filtered according to the relative distance of each finger joint to obtain multiple candidate boxes;
[0146] Corner detection is performed on each candidate box to extract the corner points corresponding to each candidate box.
[0147] In multi-view images, the same physical finger joint (such as the tip of the left index finger) can be captured by different cameras. Each camera outputs a potential box (a rectangular area surrounding the joint). By calculating the overlap between the potential boxes, it can be verified whether the potential boxes from different views belong to the same finger joint.
[0148] Relative distance is the difference between the finger joint position in the camera coordinate system and the average joint position (a normalized value reflecting the depth relationship of the joints). Using relative distance filtering can eliminate outliers caused by background noise or model misjudgment.
[0149] For potential bounding boxes that meet the overlap criteria, their predicted relative distances are extracted, and boxes whose relative distances conform to the laws of human anatomy are selected. For example, the relative distance between the fingertip and the metacarpophalangeal joint should be greater than that between the fingertips (because the fingertips are farther away); if the relative distance of a potential bounding box obviously violates the law (such as the fingertip being closer than the palm base), it is judged as an abnormal bounding box and filtered out.
[0150] For example, when the overlap exceeds a threshold (e.g., 0.5), it indicates that the two bounding boxes have the same height and likely correspond to the same 3D joint. In this case, the potential bounding box corresponding to the joint with the smaller relative distance is selected as the candidate bounding box. Since the relative distance is the normalized distance from the joint to the average joint position of the hand (the smaller the distance, the closer the joint is to the center of the hand, which is more in line with the natural structure), if two bounding boxes have a high overlap, but the relative distance of one of the finger joints is large, it is more likely to be an occluded or misjudged point (e.g., background noise is misjudged as a joint by the model, usually deviating from the center of the hand).
[0151] Corner detection algorithms are used to detect corners in each candidate box in order to extract the corners corresponding to each candidate box.
[0152] In this embodiment of the invention, potential bounding boxes are screened by overlap and relative distance before corner detection, which improves the reliability and efficiency of feature extraction from the source, reduces false detections and noise interference, and makes subsequent corner analysis and 3D localization more accurate and efficient.
[0153] Based on the above embodiments, determining the overlap of the potential boxes corresponding to the same finger joint point includes:
[0154] The boundaries of each potential box are determined based on the location of each potential pixel.
[0155] Based on the boundaries of each potential box, determine the intersection region and union region corresponding to each potential box;
[0156] The overlap of potential boxes corresponding to the same finger joint is determined based on the intersection region and the union region.
[0157] In real 3D finger joints, although the 2D projections of these joints in multi-view images are located at different positions, the overlapping areas of their potential bounding boxes are more concentrated (because the projection pattern is constrained by the camera pose). Conversely, finger joints misidentified by the model will have very low overlap in their multi-view potential bounding boxes (e.g., randomly guessed points will have almost no overlap in their bounding boxes across different views). Therefore, by calculating the overlap of these potential bounding boxes, the reliability of the model's predictions can be determined (the higher the overlap, the more likely it is to be realistic).
[0158] It should be understood that the intersection region represents the area of the overlapping portion of two potential boxes (i.e., the area that simultaneously belongs to both boxes); the union region represents the total area covered by the two potential boxes. Overlap can be measured by the intersection-union ratio (IU). The closer the overlap is to 1, the higher the overlap between the two potential boxes, and the more likely they correspond to the same physical incisor (incisor positions are consistent from different viewpoints, so potential boxes naturally overlap); the closer the overlap is to 0, the less overlap the two potential boxes have, and they likely correspond to different incisors (or one of the boxes is predicted incorrectly).
[0159] Assume finger joint point A is The selected size is Finger joint point B is The selected size is Then the left boundary of the potential bounding box of finger joint point A is The right boundary is The lower boundary is The upper boundary is Similarly, the left boundary of the potential bounding box of finger joint point B can be obtained. right boundary lower boundary and upper boundary .
[0160] For the potential bounding boxes of finger joint A and finger joint B, the overlap can be calculated using the following steps:
[0161] (1) Calculate the horizontal overlap width : Take the minimum value of the right boundary and the maximum value of the left boundary of the two boxes:
[0162] ;
[0163] like This indicates that there is no overlap in the horizontal direction. ;otherwise, It represents the number of overlapping pixels in the horizontal direction;
[0164] (2) Calculate the vertical overlap height : Take the minimum value of the upper boundary and the maximum value of the lower boundary of the two boxes:
[0165] ;
[0166] like This indicates that there is no overlap in the vertical direction. ;otherwise, It represents the number of overlapping pixels in the vertical direction;
[0167] (3) Calculate the intersection region (i.e., the overlapping area). :
[0168] ;
[0169] (4) Calculate the union region (i.e., the area of the union). Both boxes have a side length of . A square, the area of which is Therefore, the area of the union is:
[0170] ;
[0171] (5) Calculate the degree of overlap :
[0172] ;
[0173] The closer to 1, the more the two boxes overlap, and the more likely they correspond to the same real joint point, because the projection boxes of the same 3D joint should have reasonable overlap from different viewpoints.
[0174] For example, suppose:
[0175] Potential frame of camera 1: center (320, 240), size 20×20, boundary [310, 330, 230, 250];
[0176] Potential frame of camera 2: center (322, 243), size 20×20, boundary [312, 332, 233, 253];
[0177] Intersection boundary: [312, 330, 233, 250], area 18 × 17 = 306;
[0178] Union area: 20×20+20×20-306=494;
[0179] Overlap ratio: IoU = 306 / 494 ≈ 0.62.
[0180] This invention uses overlap analysis to filter out the most likely real key point locations from potential bounding boxes, supporting fully automated 3D annotation, thereby reducing manual intervention and improving the efficiency and accuracy of 3D annotation.
[0181] Based on the above embodiments, step 104 includes:
[0182] Based on the intrinsic parameters of multiple reference cameras, each of the 3D positions is projected onto the image plane of each of the reference cameras to obtain the pixel position corresponding to each of the 3D positions;
[0183] If the pixel position corresponding to the 3D position is on the image plane of the reference camera, then the 3D position is determined to be a valid 3D position.
[0184] Each of the finger joints is labeled based on the effective 3D position.
[0185] It should be understood that in a multi-camera system, the field of view (the area of the image that can be captured) of each camera is finite. Even if 3D coordinates are calculated through mathematical transformations, if the pixel position projected onto a certain camera exceeds the image boundary (e.g., the x-coordinate is -10 or 1000, while the image resolution is 640×480), it indicates that the 3D coordinates are physically impossible to observe by that camera (or there is an error in the calculation process). Therefore, the process verifies in two steps: checking whether the pixel position projected onto the camera image by the 3D coordinates is within the image range; retaining the results of valid projections (pixels within the image) and filtering out invalid calculations.
[0186] It should be understood that the reference camera is the preset main camera, such as... Figure 3 Cameras 1 and 2 in the diagram can be used as reference cameras.
[0187] Verification is performed using projection from a first reference camera (e.g., camera 1): The pixel position of a 3D coordinate point on the image of camera 1 can be obtained through intrinsic projection from camera 1. If the calculated pixel coordinates are within the image, the calculated pixel is considered valid and recorded. Specifically, camera intrinsic parameters (such as focal length, principal point, and distortion coefficients) describe the projection relationship from "3D spatial point to 2D pixel." Assuming the 3D coordinates of a certain joint are known (in the coordinate system of camera 1, or transformed to the coordinate system of camera 1 through extrinsic parameters), substituting them into the intrinsic projection model (e.g., pinhole model + distortion correction), the pixel coordinates (u, v, typically ranging from [0, image width] × [0, image height]) of that point on the image of camera 1 can be calculated. If the projected pixel coordinates satisfy 0≤u≤image width and 0≤v≤image height, it means that the 3D point can theoretically be captured by camera 1, and the calculation result is valid; if it exceeds the range (e.g., u=-5 or v=800, while the image width is 640), it means that there is a problem with the 3D coordinates or projection model (e.g., incorrect external parameter transformation or inaccurate internal parameter calibration), and the result is invalid.
[0188] Verification via projection from a second reference camera (such as camera number 2): Then via... The 3D position in the coordinate system of camera 2 can be calculated. By projecting the camera's intrinsic parameters, the pixel position of the 3D coordinate point on the image of camera 2 can be obtained. If the calculated pixel coordinates are within the image, the calculated pixel is considered valid and recorded. Specifically, This is the extrinsic transformation matrix from camera 2 to camera 1 (including rotation and translation, describing the relative pose of the two cameras). If the 3D coordinates of the joints in camera 1's coordinate system are known... ,pass (Matrix transformation) can be used to obtain the 3D coordinates of the point in the coordinate system of camera 2. Using the intrinsic parameters of camera number 2 (independently calibrated focal length, principal point, etc.), Projecting the image onto the image plane of camera 2, calculate the pixel coordinates (u', v'). Similarly, check if (u', v') is within the image range of camera 2. If valid, it means that the 3D coordinates can be observed by both cameras 1 and 2, and are reliable data from multi-view fusion; if invalid (e.g., only valid on camera 1, with projection from camera 2 outside the range), then troubleshooting is needed (e.g., external parameters). (Incorrect calibration, or the 3D coordinates themselves are unreasonable).
[0189] The validity verification of 3D coordinates to 2D pixels provided in this embodiment of the invention uses intrinsic parameter projection combined with image range judgment to ensure that 3D joints can be physically observed by the camera, thereby improving the accuracy and reliability of 3D annotation of finger joints.
[0190] The finger joint marking device provided by the present invention will be described below. The finger joint marking device described below can be referred to in correspondence with the finger joint marking method described above.
[0191] refer to Figure 5 The finger joint annotation device provided by the present invention includes a data acquisition module 501, a key point detection module 502, a 3D position determination module 503, and an annotation module 504.
[0192] The acquisition module 501 is used to acquire a set of multi-view images in the same time frame; the set of multi-view images is obtained by multiple cameras synchronously capturing images within a common viewing area;
[0193] The key point detection module 502 is used to perform key point detection on the multi-view image set, obtain the potential pixel position of each finger joint, and generate a potential bounding box centered on the potential pixel position.
[0194] The 3D position determination module 503 is used to extract the corner points of each potential bounding box and determine the 3D position of each finger joint based on the responsivity of each corner point; the corner points are image feature points related to the finger joint.
[0195] The annotation module 504 is used to annotate each of the finger joints according to the 3D positions.
[0196] The finger joint annotation device provided in this invention collects a set of multi-view images within the same time frame; these images are simultaneously captured by multiple cameras within a shared field of view; keypoint detection is performed on the multi-view images to obtain the potential pixel positions of each finger joint, and a potential bounding box centered on the potential pixel positions is generated; corner points of each potential bounding box are extracted, and the 3D position of each finger joint is determined based on the responsivity of each corner point; corner points are image feature points related to the finger joint; and each finger joint is annotated based on its 3D position. This invention overcomes the limitations of single-view fusion through multi-view fusion, improving the reliability of spatial positioning. Simultaneously, through potential bounding box and responsivity filtering, it improves feature reliability and 3D positioning accuracy, thereby enhancing the accuracy of 3D finger joint annotation.
[0197] In one embodiment, the 3D position determination module 503 is further configured to:
[0198] For each finger joint, determine the number of corner points where the responsivity is greater than the responsivity threshold;
[0199] Based on the number of corner points and camera parameters, determine the 3D position of each finger joint in the reference camera coordinate system.
[0200] In one embodiment, the 3D position determination module 503 is further configured to:
[0201] When the number of corner points is 2, determine the first target camera corresponding to each of the two corner points, and the first pixel position of each of the two corner points;
[0202] Based on the camera parameters of each of the first target cameras, construct the projection matrix of each of the first target cameras;
[0203] Based on the projection matrices and the first pixel positions of the two corner points, a triangulated system of linear equations is established.
[0204] By solving the system of linear equations, the 3D position of each finger joint in the reference camera coordinate system is obtained.
[0205] In one embodiment, the 3D position determination module 503 is further configured to:
[0206] When the number of corner points is greater than 2, determine the second target camera corresponding to each corner point and the second pixel position of each corner point;
[0207] The 3D position of the finger joint in the coordinate system of the reference camera is used as the optimization variable;
[0208] For each of the second target cameras, the projection coordinates of the optimization variables on the image plane of the second target camera are determined based on the optimization variables and the camera parameters of the second target camera.
[0209] Based on the projection coordinates and the positions of the second pixels, a target function is constructed;
[0210] The objective function is solved by nonlinear optimization to obtain the 3D position of each finger joint in the reference camera coordinate system.
[0211] In one embodiment, the key point detection module 502 is further configured to:
[0212] The multi-view image set is input into the key point detection model to detect key points, and the detection results output by the key point detection model are obtained.
[0213] The detection results include the potential pixel position of each finger joint and the relative distance of each finger joint; the relative distance is used to measure the distance between the finger joint and the average joint position.
[0214] In one embodiment, the 3D position determination module 503 is further configured to:
[0215] Determine the overlap of potential frames corresponding to the same finger joint point under different shooting angles;
[0216] If the overlap is greater than the overlap threshold, the potential boxes are filtered according to the relative distance of each finger joint to obtain multiple candidate boxes;
[0217] Corner detection is performed on each candidate box to extract the corner points corresponding to each candidate box.
[0218] In one embodiment, the 3D position determination module 503 is further configured to:
[0219] The boundaries of each potential box are determined based on the location of each potential pixel.
[0220] Based on the boundaries of each potential box, determine the intersection region and union region corresponding to each potential box;
[0221] The overlap of potential boxes corresponding to the same finger joint is determined based on the intersection region and the union region.
[0222] In one embodiment, the annotation module 504 is further configured to:
[0223] Based on the intrinsic parameters of multiple reference cameras, each of the 3D positions is projected onto the image plane of each of the reference cameras to obtain the pixel position corresponding to each of the 3D positions;
[0224] If the pixel position corresponding to the 3D position is on the image plane of the reference camera, then the 3D position is determined to be a valid 3D position.
[0225] Each of the finger joints is labeled based on the effective 3D position.
[0226] Figure 6 An example is a schematic diagram of the physical structure of an electronic device, such as... Figure 6 As shown, the electronic device may include a processor 610, a communications interface 620, a memory 630, and a communication bus 640, wherein the processor 610, communications interface 620, and memory 630 communicate with each other via the communication bus 640. The processor 610 can call logical instructions in the memory 630 to execute a finger joint annotation method. This method includes: acquiring a set of multi-view images in the same time frame; the multi-view image set is obtained by multiple cameras simultaneously capturing images within a common viewing area; performing keypoint detection on the multi-view image set to obtain the potential pixel position of each finger joint, and generating a potential bounding box centered on the potential pixel position; extracting the corner points of each potential bounding box, and determining the 3D position of each finger joint based on the responsivity of each corner point; the corner points are image feature points related to the finger joint; and annotating each finger joint based on its 3D position.
[0227] Furthermore, the logical instructions in the aforementioned memory 630 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0228] On the other hand, the present invention also provides a computer program product, which includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer is able to execute the finger joint annotation method provided by the above methods. The method includes: acquiring a set of multi-view images in the same time frame; the set of multi-view images is obtained by multiple cameras simultaneously capturing images within a common viewing area; performing keypoint detection on the set of multi-view images to obtain the potential pixel position of each finger joint, and generating a potential bounding box centered on the potential pixel position; extracting the corner points of each potential bounding box, and determining the 3D position of each finger joint based on the responsivity of each corner point; the corner points are image feature points related to the finger joint; and annotating each finger joint based on the 3D position.
[0229] In another aspect, the present invention also provides a non-transitory computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the finger joint annotation method provided by the above methods. The method includes: acquiring a set of multi-view images within the same time frame; the multi-view image set being simultaneously captured by multiple cameras within a shared viewing area; performing keypoint detection on the multi-view image set to obtain the potential pixel position of each finger joint, and generating a potential bounding box centered on the potential pixel position; extracting corner points of each potential bounding box, and determining the 3D position of each finger joint based on the responsivity of each corner point; the corner points being image feature points associated with the finger joint; and annotating each finger joint based on its 3D position.
[0230] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.
[0231] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.
[0232] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for marking finger joints, characterized in that, include: A collection of multi-view images is acquired within the same time frame; the collection of multi-view images is obtained by multiple cameras simultaneously capturing images within a shared field of view. Keypoint detection is performed on the multi-view image set to obtain the potential pixel position of each finger joint, and a potential bounding box centered on the potential pixel position is generated; The corner points of each potential bounding box are extracted, and the 3D position of each finger joint is determined based on the responsivity of each corner point; the corner points are image feature points related to the finger joint. Each finger joint is labeled according to its 3D position; The step of performing keypoint detection on the multi-view image set to obtain the potential pixel position of each finger joint includes: The multi-view image set is input into the key point detection model to detect key points, and the detection results output by the key point detection model are obtained. The detection results include the potential pixel position of each finger joint and the relative distance of each finger joint; the relative distance is used to measure the distance between the finger joint and the average joint position.
2. The finger joint marking method according to claim 1, characterized in that, Determining the 3D position of each finger joint based on the responsiveness of each corner point includes: For each finger joint, determine the number of corner points where the responsivity is greater than the responsivity threshold; Based on the number of corner points and camera parameters, determine the 3D position of each finger joint in the reference camera coordinate system.
3. The finger joint marking method according to claim 2, characterized in that, The step of determining the 3D position of each finger joint in the reference camera coordinate system based on the number of corner points and camera parameters includes: When the number of corner points is 2, determine the first target camera corresponding to each of the two corner points, and the first pixel position of each of the two corner points; Based on the camera parameters of each of the first target cameras, construct the projection matrix of each of the first target cameras; Based on the projection matrices and the first pixel positions of the two corner points, a triangulated system of linear equations is established. By solving the system of linear equations, the 3D position of each finger joint in the reference camera coordinate system is obtained.
4. The finger joint marking method according to claim 2 or 3, characterized in that, The step of determining the 3D position of each finger joint in the reference camera coordinate system based on the number of corner points and camera parameters includes: When the number of corner points is greater than 2, determine the second target camera corresponding to each corner point and the second pixel position of each corner point; The 3D position of the finger joint in the coordinate system of the reference camera is used as the optimization variable; For each of the second target cameras, the projection coordinates of the optimization variables on the image plane of the second target camera are determined based on the optimization variables and the camera parameters of the second target camera. Based on the projection coordinates and the positions of the second pixels, a target function is constructed; The objective function is solved by nonlinear optimization to obtain the 3D position of each finger joint in the reference camera coordinate system.
5. The finger joint marking method according to claim 1, characterized in that, The extraction of the corner points of each of the potential bounding boxes includes: Determine the overlap of potential frames corresponding to the same finger joint point under different shooting angles; If the overlap is greater than the overlap threshold, the potential boxes are filtered according to the relative distance of each finger joint to obtain multiple candidate boxes; Corner detection is performed on each candidate box to extract the corner points corresponding to each candidate box.
6. The finger joint marking method according to claim 5, characterized in that, Determining the overlap of potential boxes corresponding to the same finger joint includes: The boundaries of each potential box are determined based on the location of each potential pixel. Based on the boundaries of each potential box, determine the intersection region and union region corresponding to each potential box; The overlap of potential boxes corresponding to the same finger joint is determined based on the intersection region and the union region.
7. The finger joint marking method according to claim 1, characterized in that, The step of labeling each finger joint point according to its 3D position includes: Based on the intrinsic parameters of multiple reference cameras, each of the 3D positions is projected onto the image plane of each of the reference cameras to obtain the pixel position corresponding to each of the 3D positions; If the pixel position corresponding to the 3D position is on the image plane of the reference camera, then the 3D position is determined to be a valid 3D position. Each of the finger joints is labeled based on the effective 3D position.
8. A finger joint marking device, characterized in that, include: The acquisition module is used to acquire a set of multi-view images in the same time frame; the set of multi-view images is obtained by multiple cameras simultaneously capturing images within a common viewing area; The key point detection module is used to perform key point detection on the multi-view image set, obtain the potential pixel position of each finger joint, and generate a potential bounding box centered on the potential pixel position. A 3D position determination module is used to extract the corner points of each potential bounding box and determine the 3D position of each finger joint based on the responsivity of each corner point; the corner points are image feature points related to the finger joint. The annotation module is used to annotate each of the finger joints according to the 3D positions. The key point detection module is further configured to input the multi-view image set into the key point detection model for key point detection, and obtain the detection result output by the key point detection model; the detection result includes the potential pixel position of each finger joint and the relative distance of each finger joint; the relative distance is used to measure the distance between the finger joint and the average joint position.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that, When the processor executes the computer program, it implements the finger joint annotation method as described in any one of claims 1 to 7.
Citation Information
Patent Citations
Method for automatically generating annotation data of hand and method for calculating skeleton length
CN112767300A
Rapid human body posture estimation method and estimation system
CN119169659A