Hand recognition method and device based on multi-view camera and storage medium
By acquiring hand images through multi-cameras, extracting meshes and masks, converting them into feature maps and fusing them, the problem of the significant influence of hand posture is solved, and convenient hand recognition in any posture is achieved.
Patent Information
- Application Number
- CN202510836336.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-21
- Publication Date
- 2025-10-03
- Estimated Expiration
- 2045-06-21
AI Technical Summary
Existing hand recognition technology is easily affected by hand posture and requires users to deliberately cooperate with the prescribed hand posture. It has great limitations and is not convenient enough.
Multi-cameras are used to capture hand images, extract the hand mesh and visibility mask, convert them into feature maps through a neural network, and perform splicing and fusion to reduce the influence of posture and achieve recognition in any posture.
The convenience of hand recognition is improved, and users can recognize through any hand gesture, reducing the impact of hand gesture on recognition.
Smart Images

Figure CN120748006A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of biometric recognition technology, and in particular to a hand recognition method, device and storage medium based on a multi-camera. Background Art
[0002] Hand recognition technology is based on the human hand, leveraging its characteristics to achieve identity recognition. These characteristics primarily include hand veins. Although the structure of the human hand is relatively simple, it possesses significant diversity and ambiguity. Furthermore, hand features remain relatively stable throughout adulthood, making them excellent biometric features for identity recognition.
[0003] Existing hand recognition technology is greatly affected by hand posture. Hands with different angles and gestures are difficult to recognize each other. During recognition, the user's hand posture needs to be strictly constrained, and the user is required to make a prescribed hand posture (such as an open palm) before recognition can be performed. That is, the existing hand recognition technology is easily affected by hand posture, and the user needs to deliberately cooperate with the prescribed hand posture to complete hand recognition. It has great limitations and is not convenient enough. Summary of the Invention
[0004] Embodiments of the present invention provide a multi-camera-based hand recognition method, device, and storage medium, which can reduce the impact of hand posture in existing hand recognition technologies and improve the convenience of hand recognition.
[0005] An embodiment of the present invention provides a multi-camera-based hand recognition method comprising:
[0006] Acquire a plurality of hand images of a hand to be identified captured by a multi-camera; wherein each hand image is captured by one camera in the multi-camera;
[0007] For each hand image, extract the hand mesh of the hand image and the hand mask used to characterize the visibility of each hand feature point in the hand mesh;
[0008] Convert each hand image into a corresponding feature map; for each feature map, extract the eigenvalues of each hand feature point in the corresponding hand mesh in the feature map, and concatenate them to obtain the low-level features of the hand corresponding to each feature map;
[0009] The low-level hand features of each feature map are fused with the corresponding hand mask to obtain the low-level hand features after fusion of each hand feature point, and the final hand features of the hand to be identified are determined based on the low-level hand features after fusion of the hand feature points;
[0010] Determine the hand visibility of the hand to be identified based on all hand masks;
[0011] The hand to be identified is identified according to the hand visibility and the final hand features of the hand to be identified.
[0012] Furthermore, the hand mesh extracted from the hand image and the hand mask used to characterize the visibility of each hand feature point in the hand mesh include:
[0013] Detecting a hand frame in the hand image, and extracting a partial hand image located within the hand frame according to the hand frame;
[0014] According to the partial hand image, the hand Mesh and the hand Mask are identified; wherein the hand Mask is a Boolean vector. When the hand feature point in the hand Mesh is a visible hand feature point, the corresponding value of the hand feature point in the hand Mask is 1; when the hand feature point in the hand Mesh is an invisible hand feature point, the corresponding value of the hand feature point in the hand Mask is 0.
[0015] Furthermore, each hand image is converted into a corresponding feature map; for each feature map, the feature value of each hand feature point in the corresponding hand mesh in the feature map is extracted, and the low-level features corresponding to each feature map are obtained by splicing, including:
[0016] The hand image is converted into the corresponding feature map through a preset neural network;
[0017] For each feature map, use the horizontal coordinate value and vertical coordinate value of each hand feature point in the hand mesh as the index to extract the feature value of each hand feature point in the feature map;
[0018] The eigenvalues of all hand feature points in the hand mesh in the feature map are concatenated to obtain the low-level hand features corresponding to each feature map.
[0019] Furthermore, the fused low-level features of the hand are determined by the following formula:
[0020]
[0021] Among them, E merge [j] is the low-level hand feature after the j-th hand feature point is fused; w i [j] is the weight coefficient corresponding to the j-th hand feature point in the i-th feature map; E i [j] is the eigenvalue corresponding to the j-th hand feature point in the i-th feature map; eps is a preset value to prevent the denominator from being zero; n is the total number of feature maps; k is the total number of hand feature points; Mask i[j] is the value of the j-th hand feature point in the hand mask corresponding to the i-th feature map; d i [j] is the depth value of the j-th hand feature point in the hand Mesh corresponding to the i-th feature map.
[0022] Furthermore, the hand visibility is determined by the following formula:
[0023]
[0024] Among them, V1[j] is the hand visibility of the j-th hand feature point in the hand to be identified.
[0025] Furthermore, based on the hand visibility and the final hand features of the hand to be identified, the hand to be identified is identified, including:
[0026] Calculating the hand visibility of the hand to be identified and the similarity score of the hand visibility in the pre-stored data to obtain a first similarity score;
[0027] Calculating a similarity score between the final hand feature of the hand to be identified and the final hand feature in the pre-stored data to obtain a second similarity score;
[0028] When it is determined that both the first similarity score and the second similarity score are greater than corresponding preset thresholds, it is determined that the hand to be identified and the hand corresponding to the pre-stored data are the same hand.
[0029] Furthermore, the first similarity score is calculated by the following formula:
[0030]
[0031] Among them, S v is the first similarity score; V2[j] is the hand visibility of the j-th hand feature point in the hand corresponding to the pre-stored data.
[0032] Furthermore, the second similarity score is calculated by the following formula:
[0033]
[0034] Among them, S f is the second similarity score; F1 is the final hand feature of the hand to be identified; F1 is the final hand feature of the hand corresponding to the pre-stored data.
[0035] Based on the above method embodiment, the present invention provides a corresponding device embodiment;
[0036] An embodiment of the present invention provides a hand recognition device based on a multi-camera, comprising: an image acquisition module, a hand detection module, a feature extraction module, a feature fusion module, a visibility determination module, and a comparison module;
[0037] The image acquisition module is used to obtain a plurality of hand images of the hand to be identified acquired by the multi-camera; wherein each hand image is acquired by one camera in the multi-camera;
[0038] The hand detection module is used to extract a hand mesh of each hand image and a hand mask for representing the visibility of each hand feature point in the hand mesh;
[0039] The feature extraction module is used to convert each hand image into a corresponding feature map; for each feature map, the feature value of each hand feature point in the corresponding hand mesh in the feature map is extracted, and the feature values are concatenated to obtain the low-level hand features corresponding to each feature map;
[0040] The feature fusion module is used to fuse the low-level hand features of each feature map with the corresponding hand mask to obtain the low-level hand features after fusion of each hand feature point, and determine the final hand features of the hand to be identified based on the low-level hand features after fusion of the hand feature points;
[0041] The visibility determination module is used to determine the hand visibility of the hand to be identified based on all hand masks;
[0042] The comparison module is used to identify the hand to be identified based on the hand visibility and final hand features of the hand to be identified.
[0043] Based on the above method embodiment, the present invention provides a corresponding storage medium embodiment;
[0044] An embodiment of the present invention provides a storage medium, which includes a stored computer program, wherein when the computer program is running, the device where the storage medium is located is controlled to execute any one of the multi-camera based hand recognition methods described in the present invention.
[0045] The following beneficial effects are achieved by implementing the embodiments of the present invention:
[0046] An embodiment of the present invention provides a hand recognition method, device and storage medium based on a multi-camera. The method first collects several hand images of a hand to be identified; for each hand image, extracts a hand mesh of the hand image and a hand mask used to characterize the visibility of each hand feature point in the hand mesh; converts each hand image into a corresponding feature map; for each feature map, extracts the feature value of each hand feature point in the corresponding hand mesh in the feature map, and splices them to obtain the low-order hand features corresponding to each feature map; fuses the low-order hand features of each feature map with the corresponding hand mask to obtain the fused low-order hand features of each hand feature point, and determines the final hand features of the hand to be identified based on the fused low-order hand features of the hand feature points; determines the hand visibility of the hand to be identified based on all hand masks; and identifies the hand to be identified based on the hand visibility and the final hand features of the hand to be identified. Compared with the prior art, the present invention extracts hand Mesh vertices with three-dimensional coordinates for each hand image, and the feature values in the feature map can be mapped to the Mesh vertices through projection relationships. When the hand posture changes, the three-dimensional coordinates of the Mesh vertices will change, but the present application deforms and aligns the features so that the Mesh vertices of different postures are transformed into a standard space, so that the features of different postures can be compared in the same coordinate system. At this time, the feature difference is only determined by the hand identity (such as hand veins), not the posture, thereby reducing the influence of the hand posture, allowing users to identify hands through any hand posture, improving the convenience of hand recognition. BRIEF DESCRIPTION OF THE DRAWINGS
[0047] Figure 1 The present invention provides a flow chart of a multi-camera-based hand recognition method.
[0048] Figure 2 The figure is a schematic diagram of the principle of a hand recognition method based on multi-cameras provided by one embodiment of the present invention.
[0049] Figure 3 This is a schematic diagram of the principle of feature extraction in a hand recognition method based on multiple cameras provided by one embodiment of the present invention.
[0050] Figure 4 The figure is a schematic structural diagram of a hand recognition device based on a multi-camera according to an embodiment of the present invention. DETAILED DESCRIPTION
[0051] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.
[0052] Figure Figure 1 As shown, an embodiment of the present invention provides a hand recognition method based on a multi-camera, which includes at least the following steps:
[0053] Step S1: Acquire a plurality of hand images of the hand to be identified captured by a multi-camera; wherein each hand image is captured by a camera in the multi-camera.
[0054] In the present invention, a hand is placed within the working range of the multi-camera device, and a static image of the hand is captured to obtain a number of hand images: {P1, P2...Pn}, where n is the number of cameras in the multi-camera device (it can also be the total number of subsequent feature maps);
[0055] Reference Figure 2 Schematically, in one embodiment of the present invention, a multi-eye device (i.e., the multi-eye camera described above) is provided with four infrared cameras (i.e., four infrared cameras). The four infrared cameras are arranged in four different directions: up, down, left, and right. Each camera is 90 degrees apart, and the parameters set for each camera are consistent. The user places the hand to be identified in the collection area of the multi-eye camera, and the four infrared cameras respectively capture the hand to be identified, obtaining four hand images: P1, P2, P3, and P4. Of course, the number of infrared cameras only needs to be two or more, and the specific number is not limited. Compared with general visible light-based cameras, infrared cameras can capture richer hand vein information and can also reduce interference from information such as hand lines and stains. By deploying four infrared cameras in different directions, information on palm veins, finger veins, and dorsal veins can be collected. Compared with single palm vein and finger vein algorithms, the accuracy is higher and recognition will not be difficult due to insufficient information on certain local veins. In addition, because multiple images of the hand at different angles are used simultaneously, the three-dimensional information of the hand is naturally included in the recognition features. Generally, two-dimensional printed photos cannot be successfully attacked, and the method has strong anti-counterfeiting capabilities.
[0056] Step S2: For each hand image, extract the hand Mesh of the hand image and the hand Mask used to characterize the visibility of each hand feature point in the hand Mesh.
[0057] In a preferred embodiment, the extraction of the hand mesh from the hand image and the hand mask used to characterize the visibility of each hand feature point in the hand mesh include:
[0058] Detecting a hand frame in the hand image, and extracting a partial hand image located within the hand frame according to the hand frame;
[0059] According to the partial hand image, the hand Mesh and the hand Mask are identified; wherein the hand Mask is a Boolean vector. When the hand feature point in the hand Mesh is a visible hand feature point, the corresponding value of the hand feature point in the hand Mask is 1; when the hand feature point in the hand Mesh is an invisible hand feature point, the corresponding value of the hand feature point in the hand Mask is 0.
[0060] Specifically, according to the image {P1, P2...Pn}, the hand Mesh and the hand Mask, {Mesh1, Mesh2...Meshn} and {Mask1, Mask2...Maskn}.
[0061] Schematically, in the present invention, the hand mesh is in the form of 3D coordinates of k points, where k is determined by the hand mesh model. This embodiment uses the MANO model, so k = 778; the hand mask corresponds to the visibility of k points in the hand mesh (its form is a k-dimensional Boolean vector, with a value of 1 when visible and 0 when invisible).
[0062] Specifically, we first use the YOLO public target detection algorithm to detect the hand frame in the hand image; then extract the local hand image based on the hand frame, and further use the MANO hand mesh detection algorithm to detect the hand mesh and hand mask.
[0063] Step S3: Convert each hand image into a corresponding feature map; for each feature map, extract the eigenvalues of each hand feature point in the corresponding hand Mesh in the feature map, and splice them to obtain the low-level hand features corresponding to each feature map.
[0064] In a preferred embodiment, each hand image is converted into a corresponding feature map; for each feature map, the feature value of each hand feature point in the corresponding hand mesh in the feature map is extracted, and the low-level features corresponding to each feature map are obtained by splicing, including:
[0065] The hand image is converted into the corresponding feature map through a preset neural network;
[0066] For each feature map, use the horizontal coordinate value and vertical coordinate value of each hand feature point in the hand mesh as the index to extract the feature value of each hand feature point in the feature map;
[0067] The eigenvalues of all hand feature points in the hand mesh in the feature map are concatenated to obtain the low-level hand features corresponding to each feature map.
[0068] Specifically, feature extraction is performed on {P1, P2...Pn} and {Mesh1, Mesh2...Meshn} to obtain the low-order hand features {E1, E2...En} of the hand to be identified;
[0069] like Figure 3 As shown in the figure, taking an input hand image as an example: for a hand image P and its corresponding hand mesh, the hand low-level features E are extracted. Where E is the low-level hand feature corresponding to the mesh, which is a tensor of dimension (k, s). s is the algorithm default value. In this embodiment, k = 778 and s = 64. The specific steps are as follows:
[0070] Use the U-net neural network to convert the hand image P into a feature map E-Map. The dimension of P is (640, 400, 3), and the dimension of E-Map is (640, 400, 64).
[0071] In the E-Map, deformation alignment is performed based on the mesh coordinates to obtain the low-level hand features E. The deformation alignment operation is as follows: a. For the j-th point coordinate (uj, vj, dj) in the mesh, the index (uj, vj) in the E-Map has the eigenvalue ej, where ej is an s-dimensional vector; b. {e1, e2...ek} are concatenated together to obtain E. uj is the horizontal coordinate value of the j-th hand feature point, vj is the vertical coordinate value of the j-th hand feature point, and dj represents the depth value of the j-th hand feature point in three-dimensional space, reflecting the distance of the j-th hand feature point relative to the camera (larger values generally indicate that the vertex is farther from the camera).
[0072] Schematically, after the above feature extraction, the four hand images P1, P2, P3 and P4 obtain low-level hand features E1, E2, E3 and E4 respectively.
[0073] In this embodiment of the present invention, by deforming and aligning the features, the Mesh vertices of different postures are transformed into a standard space, so that the features of different postures can be compared in the same coordinate system. At this time, the feature difference is only determined by the hand identity (such as hand veins) rather than the posture, thereby reducing the influence of the hand posture. When performing hand recognition, the user can recognize the hand through any hand posture, thereby improving the convenience of hand recognition.
[0074] Step S4: Fuse the low-level hand features of each feature map with the corresponding hand Mask to obtain the low-level hand features after fusion of each hand feature point, and determine the final hand features of the hand to be identified based on the low-level hand features after fusion of the hand feature points.
[0075] In a preferred embodiment, the fused low-level features of the hand are determined by the following formula:
[0076]
[0077] Among them, E merge [j] is the low-level hand feature after the j-th hand feature point is fused; w i [j] is the weight coefficient corresponding to the j-th hand feature point in the i-th feature map; E i [j] is the eigenvalue corresponding to the j-th hand feature point in the i-th feature map; eps is a preset value to prevent the denominator from being zero; n is the total number of feature maps; k is the total number of hand feature points; Mask i [j] is the value of the j-th hand feature point in the hand mask corresponding to the i-th feature map; d i [j] is the depth value of the j-th hand feature point in the hand Mesh corresponding to the i-th feature map.
[0078] Schematically, multi-eye fusion is performed on {E1, E2...En} and {Mask1, Mask2...Maskn} to obtain the hand feature F1 after fusion of the hand to be identified; F1 is a t-dimensional unit vector, t is the algorithm preset value, in this embodiment k=778, t=64.
[0079] The specific steps are as follows:
[0080] First, calculate the weights {w1, w2...wn}; calculate the weight of each hand feature point in the corresponding feature map based on the depth value of the hand feature point in the hand mesh (the vertex closer to the camera has a larger weight, and the weight of invisible points is 0), n = 4, k = 778, eps = 10-5, to prevent the denominator from being 0:
[0081]
[0082] The low-order hand feature Emerge[j] of the j-th hand feature point after fusion is:
[0083]
[0084] The low-level hand features of each hand feature point are input into a pre-trained neural network GCN (Graph Convolutional Network) that can accept graph structure input to obtain the hand feature F1 of the hand to be identified.
[0085] Step S5: Determine the hand visibility of the hand to be identified based on all hand masks.
[0086] In a preferred embodiment, hand visibility is determined by the following formula:
[0087]
[0088] Among them, V1[j] is the hand visibility of the j-th hand feature point in the hand to be identified.
[0089] Indicatively, n=4, k=778; then, the hand visibility of the j-th hand feature point in the hand to be identified is:
[0090] Step S6: Identify the hand to be identified based on the hand visibility and final hand features of the hand to be identified.
[0091] In a preferred embodiment, identifying the hand to be identified based on the hand visibility and the final hand features of the hand to be identified includes:
[0092] Calculating the hand visibility of the hand to be identified and the similarity score of the hand visibility in the pre-stored data to obtain a first similarity score;
[0093] Calculating a similarity score between the final hand feature of the hand to be identified and the final hand feature in the pre-stored data to obtain a second similarity score;
[0094] When it is determined that both the first similarity score and the second similarity score are greater than corresponding preset thresholds, it is determined that the hand to be identified and the hand corresponding to the pre-stored data are the same hand.
[0095] In a preferred embodiment, the first similarity score is calculated by the following formula:
[0096]
[0097] Among them, S v is the first similarity score; V2[j] is the hand visibility of the j-th hand feature point in the hand corresponding to the pre-stored data.
[0098] In a preferred embodiment, the second similarity score is calculated by the following formula:
[0099]
[0100] Among them, S f is the second similarity score; F1 is the final hand feature of the hand to be identified; F2 is the final hand feature of the hand corresponding to the pre-stored data.
[0101] Specifically, in this embodiment, the threshold corresponding to the first similarity is T v ; The threshold corresponding to the first similarity is T f ; Preferably, set: T v =0.8; T f =0.5;
[0102] When Sv≤Tv, it is impossible to determine whether the hand to be identified and the hand corresponding to the pre-stored data are the same hand of the same person;
[0103] If Sv>Tv, and Sf>Tf, then it is determined that the hand to be identified and the hand corresponding to the pre-stored data are the same hand of the same person;
[0104] If Sv>Tv, and Sf≤Tf, it is determined that the hand to be identified and the hand corresponding to the pre-stored data are not the same hand of the same person.
[0105] like Figure 4 As shown, an embodiment of the present invention provides a hand recognition device based on a multi-camera, comprising: an image acquisition module, a hand detection module, a feature extraction module, a feature fusion module, a visibility determination module, and a comparison module;
[0106] The image acquisition module is used to obtain a plurality of hand images of the hand to be identified acquired by the multi-camera; wherein each hand image is acquired by one camera in the multi-camera;
[0107] The hand detection module is used to extract a hand mesh of each hand image and a hand mask for representing the visibility of each hand feature point in the hand mesh;
[0108] The feature extraction module is used to convert each hand image into a corresponding feature map; for each feature map, the feature value of each hand feature point in the corresponding hand mesh in the feature map is extracted, and the feature values are concatenated to obtain the low-level hand features corresponding to each feature map;
[0109] The feature fusion module is used to fuse the low-level hand features of each feature map with the corresponding hand mask to obtain the low-level hand features after fusion of each hand feature point, and determine the final hand features of the hand to be identified based on the low-level hand features after fusion of the hand feature points;
[0110] The visibility determination module is used to determine the hand visibility of the hand to be identified based on all hand masks;
[0111] The comparison module is used to identify the hand to be identified based on the hand visibility and final hand features of the hand to be identified.
[0112] It should be noted that the device embodiment described above corresponds to the method embodiment of the present invention, which can implement the hand recognition method based on multi-cameras described in any of the above method embodiments of the present invention. In addition, the device described above in the present invention is merely schematic, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the scheme of this embodiment. In addition, in the drawings of the device embodiment provided by the present invention, the connection relationship between the modules indicates that there is a communication connection between them, which can be specifically implemented as one or more communication buses or signal lines. A person of ordinary skill in the art can understand and implement it without paying any creative work.
[0113] Based on the above method embodiment, the present invention provides a corresponding terminal device embodiment.
[0114] An embodiment of the present invention provides a terminal device including a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor. When the processor executes the computer program, the hand recognition method based on multi-cameras described in any embodiment of the present invention is implemented.
[0115] The terminal device may be a computing device such as a desktop computer, a notebook, a PDA, or a cloud server.
[0116] The processor may be a central processing unit (CPU), other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor, etc. The processor is the control center of the terminal device, connecting various parts of the entire terminal device using various interfaces and lines.
[0117] The memory may be used to store the computer program, and the processor implements various functions of the terminal device by running or executing the computer program stored in the memory and calling the data stored in the memory. In addition, the memory may include a high-speed random access memory and may also include a non-volatile memory, such as a hard disk, a memory, a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, at least one disk storage device, a flash memory device, or other volatile solid-state storage device.
[0118] The computer program includes computer program code, which may be in source code form, object code form, executable file, or some intermediate form.
[0119] Based on the above method embodiment, the present invention provides a corresponding storage medium embodiment.
[0120] An embodiment of the present invention provides a storage medium, wherein the storage medium includes a stored computer program, wherein when the computer program is running, the device where the storage medium is located is controlled to perform the hand recognition method of a multi-camera as described in any one of the present inventions.
[0121] The storage medium is a computer-readable medium, which may include: any entity or device that can carry the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal and software distribution medium, etc.
[0122] The above is a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present invention. These improvements and modifications are also considered to be within the scope of protection of the present invention.
Claims
1. A hand recognition method based on multi-camera, characterized in that: include: Acquire a plurality of hand images of a hand to be identified captured by a multi-camera; wherein each hand image is captured by one camera in the multi-camera; For each hand image, extract the hand mesh of the hand image and the hand mask used to characterize the visibility of each hand feature point in the hand mesh; Convert each hand image into a corresponding feature map; for each feature map, extract the eigenvalues of each hand feature point in the corresponding hand mesh in the feature map, and concatenate them to obtain the low-level features of the hand corresponding to each feature map; The low-level hand features of each feature map are fused with the corresponding hand mask to obtain the low-level hand features after fusion of each hand feature point, and the final hand features of the hand to be identified are determined based on the low-level hand features after fusion of the hand feature points; Determine the hand visibility of the hand to be identified based on all hand masks; The hand to be identified is identified according to the hand visibility and the final hand features of the hand to be identified.
2. The hand recognition method based on multi-camera according to claim 1, characterized in that: The hand mesh extracted from the hand image and the hand mask used to characterize the visibility of each hand feature point in the hand mesh include: Detecting a hand frame in the hand image, and extracting a partial hand image located within the hand frame according to the hand frame; According to the partial hand image, the hand Mesh and the hand Mask are identified; wherein the hand Mask is a Boolean vector. When the hand feature point in the hand Mesh is a visible hand feature point, the corresponding value of the hand feature point in the hand Mask is 1; when the hand feature point in the hand Mesh is an invisible hand feature point, the corresponding value of the hand feature point in the hand Mask is 0.
3. The hand recognition method based on multi-camera according to claim 2, characterized in that: Each hand image is converted into a corresponding feature map; for each feature map, the feature value of each hand feature point in the corresponding hand mesh in the feature map is extracted, and the low-level features corresponding to each feature map are obtained by splicing, including: The hand image is converted into the corresponding feature map through a preset neural network; For each feature map, use the horizontal coordinate value and vertical coordinate value of each hand feature point in the hand mesh as the index to extract the feature value of each hand feature point in the feature map; The eigenvalues of all hand feature points in the hand mesh in the feature map are concatenated to obtain the low-level hand features corresponding to each feature map.
4. The hand recognition method based on multi-camera according to claim 3, characterized in that: The fused low-level features of the hand are determined by the following formula: Among them, E merge [j] is the low-level hand feature after the j-th hand feature point is fused; w i [j] is the weight coefficient corresponding to the j-th hand feature point in the i-th feature map; E i [j] is the eigenvalue corresponding to the j-th hand feature point in the i-th feature map; eps is a preset value to prevent the denominator from being zero; n is the total number of feature maps; k is the total number of hand feature points; Mask i [j] is the value of the j-th hand feature point in the hand mask corresponding to the i-th feature map; d i [j] is the depth value of the j-th hand feature point in the hand Mesh corresponding to the i-th feature map.
5. The hand recognition method based on multi-camera according to claim 4, characterized in that: Hand visibility is determined by the following formula: Among them, V1[j] is the hand visibility of the j-th hand feature point in the hand to be identified.
6. The hand recognition method based on multi-camera according to claim 5, characterized in that: The hand to be identified is identified based on the hand visibility and final hand features, including: Calculating the hand visibility of the hand to be identified and the similarity score of the hand visibility in the pre-stored data to obtain a first similarity score; Calculating a similarity score between the final hand feature of the hand to be identified and the final hand feature in the pre-stored data to obtain a second similarity score; When it is determined that both the first similarity score and the second similarity score are greater than corresponding preset thresholds, it is determined that the hand to be identified and the hand corresponding to the pre-stored data are the same hand.
7. The hand recognition method based on multi-camera according to claim 6, characterized in that: The first similarity score is calculated by the following formula: Among them, S v is the first similarity score; V2[j] is the hand visibility of the j-th hand feature point in the hand corresponding to the pre-stored data.
8. The hand recognition method based on multi-camera according to claim 7, characterized in that: The second similarity score is calculated by the following formula: Among them, S f is the second similarity score; F1 is the final hand feature of the hand to be identified; F1 is the final hand feature of the hand corresponding to the pre-stored data.
9. A hand recognition device based on a multi-camera, characterized in that: include: Image acquisition module, hand detection module, feature extraction module, feature fusion module, visibility determination module and comparison module; The image acquisition module is used to obtain a plurality of hand images of the hand to be identified acquired by the multi-camera; wherein each hand image is acquired by one camera in the multi-camera; The hand detection module is used to extract a hand mesh of each hand image and a hand mask for representing the visibility of each hand feature point in the hand mesh; The feature extraction module is used to convert each hand image into a corresponding feature map; for each feature map, the feature value of each hand feature point in the corresponding hand mesh in the feature map is extracted, and the feature values are concatenated to obtain the low-level hand features corresponding to each feature map; The feature fusion module is used to fuse the low-level hand features of each feature map with the corresponding hand mask to obtain the low-level hand features after fusion of each hand feature point, and determine the final hand features of the hand to be identified based on the low-level hand features after fusion of the hand feature points; The visibility determination module is used to determine the hand visibility of the hand to be identified based on all hand masks; The comparison module is used to identify the hand to be identified based on the hand visibility and final hand features of the hand to be identified.
10. A storage medium, characterized in that: The storage medium includes a stored computer program, wherein, when the computer program is running, the device where the storage medium is located is controlled to execute the hand recognition method for a multi-camera according to any one of claims 1 to 8.
Citation Information
Patent Citations
Vehicle-mounted HUD man-machine interaction system based on gesture recognition
CN111158457A
Target detection system, method and terminal based on improved YOLO-V3
CN111553406A
Virtual reality input system based on gesture recognition
CN116909393A
Feature map processing method, image recognition method and related device
CN117152454A