Palm region detection method and apparatus

Through two-level network structure and image pyramid technology, the problems of low accuracy and slow detection of palm area detection are solved, and high-precision and rapid detection under palm posture changes are achieved.

WO2025139304A1PCT designated stage expired Publication Date: 2025-07-03GRG BANKING EQUIPMENT CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/127442
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-29
Filing Date
2024-10-25
Publication Date
2025-07-03

AI Technical Summary

Technical Problem

The existing palm area detection methods have low detection accuracy and slow inference speed, which cannot meet the speed requirements of palm recognition.

Method used

Using a two-level network structure, the first-level network performs feature extraction and candidate corner point detection of initial palm images, and the second-level network performs corner point type recognition and screening, combining image pyramids and non-maximum suppression algorithms to improve detection accuracy and speed.

Benefits of technology

In the case of changing palm posture, the target palm area can be accurately detected, which improves the accuracy and accuracy of detection, and at the same time accelerates the network's inference speed and meets the speed requirements of palm recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024127442_03072025_PF_FP_ABST
    Figure CN2024127442_03072025_PF_FP_ABST
Patent Text Reader

Abstract

A palm region detection method and apparatus. The method comprises: performing feature extraction on an acquired initial palm image on the basis of a first-level network, so as to obtain a plurality of first images, wherein the first images comprise candidate corner point features; performing type recognition on the candidate corner point features on the basis of a second-level network, and screening the plurality of first images on the basis of corner point types obtained by means of recognition, so as to obtain a plurality of second images, wherein the second images comprise target corner point features and target position information corresponding to the target corner point features, the target corner point features corresponding to the corner point types; and performing region localization on the initial palm image on the basis of the corner point types and the target position information, which correspond to the target corner point features, so as to determine a target palm region, wherein the target palm region is a region in a palm that comprises at least some palmprint features. The palm region detection method is relatively high in detection precision and accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Palm area detection method and device

[0001] Related applications

[0002] This application claims priority to Chinese patent application No. 2023118743436, filed on December 29, 2023, entitled “Palm Area Detection Method and Device,” which is hereby incorporated by reference in its entirety. Technical Field

[0003] The present application belongs to the field of image processing technology, and in particular relates to a palm area detection method and device. Background Art

[0004] Palm biometric recognition requires identifying biometric features based on the palm region of interest within the palm image. Related technologies include methods for palm region detection based on neural networks. Common palm region detection methods suffer from low accuracy and often lead to false detections. Furthermore, the inference speed of common neural networks is slow, making them incapable of meeting the speed requirements for palm recognition.

[0005] Summary of the Invention

[0006] According to various embodiments of the present application, a palm area detection method and apparatus are provided.

[0007] The technical solutions adopted in the embodiments of this application are as follows:

[0008] In a first aspect, the present application provides a palm area detection method, the method comprising:

[0009] Performing feature extraction on the acquired initial palm image based on the first-level network to obtain a plurality of first images; the first images include candidate corner features; the first-level network is trained based on the first training set;

[0010] Based on the second-level network, each candidate corner feature is identified as a type, and the plurality of first images are screened based on the identified corner type to obtain a plurality of second images; the second images include target corner features and target position information corresponding to the target corner features, and the target corner features correspond to the corner type; the second-level network is trained based on the second training set;

[0011] Based on the corner point type corresponding to each target corner point feature and the target position information, the initial palm image is region-located to determine a target palm region; the target palm region is a region of the palm that includes at least part of the palm print feature.

[0012] In a second aspect, the present application provides a palm area detection device, the device comprising:

[0013] A first processing module is configured to extract features from the acquired initial palm image based on a first-level network to obtain a plurality of first images; the first images include candidate corner features; and the first-level network is trained based on a first training set;

[0014] a second processing module configured to identify the type of each candidate corner feature based on a second-level network, and filter the plurality of first images based on the identified corner type to obtain a plurality of second images; the second images including target corner features and target position information corresponding to the target corner features, the target corner features corresponding to the corner type; the second-level network being trained based on a second training set;

[0015] The third processing module is used to perform regional positioning on the initial palm image based on the corner point type corresponding to each target corner point feature and the target position information to determine the target palm area; the target palm area is the area in the palm that includes at least part of the palm print features.

[0016] In a third aspect, the present application provides an electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the palm area detection method as described in the first aspect above is implemented.

[0017] In a fourth aspect, the present application provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the palm area detection method as described in the first aspect above.

[0018] In a fifth aspect, the present application provides a computer program product, including a computer program, which, when executed by a processor, implements the palm area detection method as described in the first aspect above.

[0019] Details of one or more embodiments of the present application are set forth in the following drawings and description. Other features, objects, and advantages of the present application will become apparent from the description, drawings, and claims, or will be understood through practice of the present application. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the conventional technology, the following briefly introduces the drawings required for use in the description of the embodiments or conventional technology. Obviously, the drawings described below are only embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on the disclosed drawings without creative work. In order to better describe and illustrate the embodiments and / or examples of the inventions disclosed herein, one or more drawings may be referenced. The additional details or examples used to describe the drawings should not be considered as limiting the scope of the disclosed invention, the embodiments and / or examples currently described, and any of the best modes of these inventions currently understood.

[0021] FIG1 is a flowchart of a palm area detection method according to some embodiments.

[0022] FIG. 2 is a schematic diagram showing a palm region detection method according to some embodiments.

[0023] FIG3 is a second schematic diagram of the principles of the palm area detection method according to some embodiments.

[0024] FIG4 is a third schematic diagram of the principles of the palm area detection method according to some embodiments.

[0025] FIG. 5 is a fourth schematic diagram of the principles of a palm area detection method according to some embodiments.

[0026] FIG6 is a fifth schematic diagram of the principles of the palm area detection method according to some embodiments.

[0027] FIG. 7 is a sixth schematic diagram of the principles of the palm area detection method according to some embodiments.

[0028] FIG8 is a schematic structural diagram of a palm area detection device according to some embodiments.

[0029] FIG9 is a schematic structural diagram of an electronic device according to some embodiments.

[0030] FIG10 is a seventh schematic diagram of the principles of a palm area detection method according to some embodiments. DETAILED DESCRIPTION

[0031] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0032] The terms "first," "second," and the like in the specification and claims of this application are used to distinguish similar objects, and are not used to describe a specific order or precedence. It should be understood that the terms used in this manner are interchangeable where appropriate, so that the embodiments of this application can be implemented in an order other than that illustrated or described herein, and that the objects distinguished by "first," "second," and the like are generally of the same type, and do not limit the number of objects; for example, the first object can be one or more. In addition, the term "and / or" in the specification and claims refers to at least one of the connected objects, and the character " / " generally indicates that the objects connected are in an "or" relationship.

[0033] The palm area detection method, palm area detection device, electronic device, and readable storage medium provided in the embodiments of the present application are described in detail below with reference to the accompanying drawings through specific embodiments and their application scenarios.

[0034] The palm area detection method may be applied to a terminal, and may be specifically executed by hardware or software in the terminal.

[0035] The terminal includes, but is not limited to, a portable communication device such as a mobile phone or tablet computer having a touch-sensitive surface (e.g., a touch screen display and / or a touch pad). It should also be understood that, in some embodiments, the terminal may not be a portable communication device, but a desktop computer having a touch-sensitive surface (e.g., a touch screen display and / or a touch pad).

[0036] In the following embodiments, a terminal including a display and a touch-sensitive surface is described. However, it should be understood that the terminal may include one or more other physical user interface devices such as a physical keyboard, a mouse, and a joystick.

[0037] The palm area detection method provided in the embodiments of the present application may be executed by an electronic device or a functional module or functional entity in the electronic device that can implement the palm area detection method. The electronic devices mentioned in the embodiments of the present application include but are not limited to mobile phones, tablet computers, computers, cameras, and wearable devices. The palm area detection method provided in the embodiments of the present application is described below using an electronic device as an example of the execution entity.

[0038] As shown in FIG1 , the palm area detection method includes: operation 110, operation 120, and operation 130. It should be understood that although the steps in the flowchart shown in FIG1 are displayed in sequence as indicated by the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise specified herein, there is no strict order restriction for the execution of these steps, and they can be executed in other orders. Moreover, at least part of the steps in FIG1 may include multiple sub-steps or multiple stages, and these sub-steps or stages are not necessarily executed at the same time, but can be executed at different times, and their execution order is not necessarily sequential, but can be executed in turn or alternating with other steps or at least part of the sub-steps or stages of other steps.

[0039] Operation 110 : performing feature extraction on the acquired initial palm image based on the first-level network to obtain a plurality of first images; the first images include candidate corner features; and the first-level network is trained based on the first training set.

[0040] In this operation, when performing biometric recognition on the palm, it is necessary to collect the user's palm image. For example, when performing palm print or palm vein recognition on the palm, it is necessary to obtain the initial palm image based on the collection terminal.

[0041] For example, when a user passes through a subway gate by swiping his palm, the collection terminal may be an identification terminal installed at the gate.

[0042] The collection terminal may also be a mobile collection terminal, for example, the collection terminal may be a user's mobile phone or other device.

[0043] The initial palm image may include palm print features and corner point features of the user's palm.

[0044] By performing feature extraction on the initial palm image, multiple first images can be obtained.

[0045] The first image includes candidate corner features, and the first image is an image of potential corner targets.

[0046] Candidate corner features may include corner features and background features.

[0047] The multiple first images may include: an image containing all corner features and an image containing most corner features and a small portion of background features.

[0048] As shown in Figure 2, the corner features are shown as A, B, C and D, and the background features are shown as N.

[0049] As shown in FIG3 , the first-level network may be a convolutional neural network or any other neural network that can implement this function, and this application does not limit this.

[0050] When the first-stage network is a convolutional neural network, the first-stage network may include a maximum pooling (Max Pool) layer, a convolution (conv) layer, a classification layer, and a coordinate regression layer.

[0051] The output of the max pooling layer is connected to the input of the convolutional layer.

[0052] The output of the convolutional layer is connected to the input of the classification layer and the coordinate regression layer respectively.

[0053] During the training of the first-level network, the classification layer can be trained using the softmax activation function, and the convolutional layer can be trained using the relu activation function.

[0054] As shown in FIG3 , during the training process of the first-level network, the training samples of the first-level network need to be scaled to 16*16.

[0055] For example, multiple sample images can be obtained from a sample palm image, where the sample images include sample corner features or background features. The sample images are then scaled to 16*16*1, and then input into the maximum pooling layer. The output image of the maximum pooling layer is 7*7*10 in size.

[0056] Then input the image output by the maximum pooling layer into the first convolutional layer, and the image output by the first convolutional layer can be obtained, with a size of 5*5*16;

[0057] The image output by the first convolutional layer is input into the second convolutional layer to obtain the output image of the second convolutional layer, which has a size of 3*3*32.

[0058] In the actual implementation process, the first-level network may include three convolutional layers, as shown in Figure 4, and a channel attention module may be set between every two adjacent convolutional layers.

[0059] For example, an average pooling layer, a full connection layer, and a scale layer can be set between two convolutional layers.

[0060] The output of the average pooling layer is connected to the input of the fully connected layer.

[0061] The output of the fully connected layer is connected to the input of the scaling layer.

[0062] The channel attention module can be used to focus on different types of features of an image. It can learn a weight for each channel and then assign the corresponding weight to each channel.

[0063] In this application, by setting a channel attention module between each two adjacent convolutional layers, the channel attention module can automatically pay attention to the features of the corner area. When the palm is closed, the false detection rate of the finger area features is low, thereby improving the detection target recall rate in scenarios where the palm posture changes such as palm rotation or tilt.

[0064] The first training set is sample data used to train the first-level network.

[0065] The first training set may include sample images and label information corresponding to the sample images.

[0066] The sample images may include positive sample images, negative sample images, and partial sample images.

[0067] The labeling information corresponding to the positive sample image may include "is a corner point", etc., the labeling information corresponding to the negative sample image may include "is a background point", etc., and the labeling information corresponding to the partial sample image may include the actual offset between the partial sample image and the positive sample image.

[0068] In the actual implementation process, the classification layer of the first-level network can be trained based on the positive sample images and the label information corresponding to the positive sample images, and the negative sample images and the label information corresponding to the negative sample images.

[0069] The coordinate regression layer of the first-level network can be trained based on the partial sample images and the label information corresponding to the partial sample images. For example, the partial sample images and the actual offset between the partial sample images and the positive sample images can be input into the first-level network to obtain the predicted offset between the partial sample images and the positive sample images predicted by the first-level network, and then the first-level network can be trained based on the predicted offset and the actual offset.

[0070] In some embodiments, operation 110 may include:

[0071] Based on the initial palm image and the first-level network, position prediction is performed on multiple first candidate corner points in the initial palm image to obtain first predicted pixel positions and first predicted offset information corresponding to each first candidate corner point;

[0072] Based on the plurality of first prediction offset information and target parameters corresponding to the first-level network, the first predicted pixel positions are corrected to obtain a plurality of first images; the target parameters include: at least one of a downsampling multiple and an input size.

[0073] In this embodiment, the initial palm image includes a plurality of first candidate corner points.

[0074] The first candidate corner point may be a corner point, or the first candidate corner point may be a background point.

[0075] By performing position prediction on multiple first candidate corner points, the first predicted pixel positions of the first candidate corner points on the initial palm image can be obtained. The offset between the coordinates of the upper left corner and the lower right corner of the area where the first candidate corner points are located and the upper left corner and the lower right corner of the actual corner point area can also be predicted.

[0076] The target parameters corresponding to the first-level network may include the downsampling multiple of the first-level network and the input size of the first-level network.

[0077] The downsampling factor corresponding to the first-level network is the ratio between the input size and output size of the first-level network.

[0078] The input size of the first-level network can be 16, or can be other values, which is not limited in this application.

[0079] In some embodiments, based on the initial palm image and the first-level network, position prediction is performed on multiple first candidate corner points in the initial palm image to obtain first predicted pixel positions and first predicted offset information corresponding to each first candidate corner point, which may include:

[0080] Based on multiple scaling ratios, the initial palm image is scaled to obtain multiple first palm images of different sizes;

[0081] Inputting a target palm image from the plurality of first palm images into a first-stage network, obtaining a plurality of first candidate corner point images output by the first-stage network, first predicted offset information corresponding to the first candidate corner points in each of the first candidate corner point images, and first predicted pixel positions corresponding to each of the first candidate corner points; each of the first candidate corner point images has a first probability corresponding thereto;

[0082] Correcting the first predicted pixel positions based on the plurality of first prediction offset information and target parameters corresponding to the first-level network to obtain the plurality of first images may include:

[0083] When the first probability is greater than the first threshold, determining the first candidate corner point corresponding to the first probability as the first corner point;

[0084] Based on the downsampling multiple corresponding to the first-level network, the scaling ratio corresponding to the target palm image, the first prediction offset information corresponding to each first corner point, and the input size corresponding to the first-level network, the first predicted pixel position corresponding to each first corner point is corrected to obtain a second predicted pixel position of each first corner point on the initial palm image;

[0085] Based on the multiple second predicted pixel positions, deduplication processing is performed on images corresponding to the multiple first corner points to obtain multiple first images.

[0086] In this embodiment, the first palm image is an image obtained by scaling the initial palm image. The scaling ratios corresponding to the first palm images are different, and the image sizes corresponding to the first palm images may also be different.

[0087] Based on the same scaling ratio, the initial palm image can be scaled into one or more first palm images, which can be user-defined and is not limited in this application.

[0088] When scaling the initial palm image, the initial palm image can be scaled by 2-3 scales. For example, when the size of the initial palm image is 100*100, the initial palm image can be scaled to 60*60, 50*50 and 40*40 respectively, or the initial palm image can be scaled to other ratios, which is not limited in this application.

[0089] An image pyramid can be constructed based on multiple first palm images of different sizes. Each layer of the image pyramid may include multiple first palm images. Multiple sample first palm images in the same layer have the same size. The length and width of the smallest image in the image pyramid may be greater than 64.

[0090] In the present application, the initial palm image is scaled based on multiple scaling ratios to obtain multiple first palm images of different sizes. In actual applications, when the distance between the palm and the acquisition terminal changes and the proportion of the palm in the palm image is different, the collected initial palm image is scaled and the multiple scaled first palm images are input into the first-level network, so that the first-level network can accurately output potential corner point targets in the palm image, thereby improving the prediction precision and accuracy of the first-level network.

[0091] The target palm image is any one of the multiple first palm images. In actual application, the multiple first palm images can be input into the first-level network respectively.

[0092] The first candidate corner point image includes the first candidate corner point. A first palm image is input into the first-level network, and multiple first candidate corner point images output by the first-level network can be obtained.

[0093] The first candidate corner image may include all corner features, or may include most corner features and a small portion of background features.

[0094] Each first candidate corner point image corresponds to a first probability.

[0095] The first probability is used to represent the probability that the first candidate corner image is a corner image. For example, when the first candidate corner image includes most corner features and a small part of background features, the first probability corresponding to the first candidate corner image may be 90%.

[0096] The first prediction offset information is used to represent the offset between the coordinates of the upper left corner and lower right corner of the area where the first candidate corner point is located predicted by the first-level network and the upper left corner and lower right corner of the actual corner point area.

[0097] The first palm image is input into the first-level network, and the predicted pixel position of the first candidate corner point output by the first-level network on the first palm image, that is, the first predicted pixel position, can also be obtained.

[0098] The first predicted pixel position may include coordinate information of the upper left corner and the lower right corner of the region where the first candidate corner point is located.

[0099] The first threshold can be used to classify the first candidate corner image. For example, when the first probability of the first candidate corner image is greater than the first threshold, the first candidate corner image can be determined as a corner image; when the first probability of the first candidate corner image is not greater than the first threshold, the first candidate corner image can be determined as a background image.

[0100] The first threshold value may be 85% or 90%, etc., and may be user-defined, and is not limited in this application.

[0101] When the first probability is greater than the first threshold, the first candidate corner point in the first candidate corner point image corresponding to the first probability may be determined as the first corner point.

[0102] Multiple first corner points can be obtained by screening the multiple first candidate corner points based on the first threshold.

[0103] The second predicted pixel position is the pixel position of the first corner point output by the first-level network on the initial palm image.

[0104] When the first predicted offset information corresponding to the first corner point is obtained, the first predicted pixel position corresponding to the first corner point may be corrected to obtain a second predicted pixel position of the first corner point on the initial palm image.

[0105] Corner point classification and coordinate prediction are performed on the first palm images included in each layer of the image pyramid, so as to obtain second predicted pixel positions corresponding to the first corner points in each first palm image.

[0106] Based on the multiple second predicted pixel positions, the images corresponding to the multiple first corner points can be deduplicated using a non-maximum suppression (NMS) algorithm, or deduplication can be performed based on other algorithms, which is not limited in this application.

[0107] In actual execution, the reasoning process of the first-level network is as follows:

[0108] During the actual inference process, the input to the first-level network is a whole image with a size larger than 16*16. The output size of the first-level network is determined based on the input size.

[0109] 1) Based on the scaling ratio r i , the initial palm image is scaled to obtain multiple first palm images of different sizes, and an image pyramid P is constructed based on the multiple first palm images i , where the length of the first palm image with the smallest size is W i and width H i Both should be greater than 64.

[0110] 2) The first palm image P i Input to the first-level network, and you can get the classification feature plane F output by the first-level network cls , whose size is Among them, 2 represents the number of categories (corner points and background points), that is, the first-level network can output multiple candidate corner point images and multiple background images.

[0111] The first candidate corner image is the candidate corner image output by the first-level network.

[0112] The first-level network can also output the coordinate feature plane F reg , whose size is Among them, 4 represents the offset of the coordinates of the upper left corner and lower right corner of the target box where each corner point and background point are located relative to the real box.

[0113] The first position offset information is the offset of the coordinates of the upper left corner and lower right corner of the target box where the corner point is located relative to the real box.

[0114] 3) The first threshold can be set to TH, and the first candidate corner point corresponding to the first probability greater than the first threshold is determined as the first corner point, and the coordinates of the first corner point (X i ,Y i ), the first position offset information corresponding to the first corner point is (X1 offset ,Y1 offset ,X2 offset ,Y2 offset), and then the second predicted pixel position of the first corner point on the palm image is calculated based on the following formula: X1=(stride*X i ) / scale Y1=(stride*Yi i ) / scale X2=(stride*X i +16) / scale Y2=(stride*Yi i +16) / scale W=X2-X1+1 H=Y2-Y1+1 X3=X1+W*X1 offset Y3=Y1+H*Y1 offset X4=X2+W*X2 offset Y4=Y2+H*Y2 offset

[0115] Among them, (X1, Y1) is the initial upper left corner coordinate of the first corner point predicted by the first level network, (X2, Y2) is the initial lower right corner coordinate of the first corner point predicted by the first level network, (X3, Y3) is the upper left corner coordinate of the first corner point on the palm image after correction, (X4, Y4) is the lower right corner coordinate of the first corner point on the palm image after correction, W is the length of the target box where the first corner point is located, H is the width of the target box where the first corner point is located, (X1 offset ,Y1 offset ) is the predicted offset information of the upper left corner coordinate corresponding to the first corner point, (X2 offset ,Y2 offset ) is the predicted offset information of the lower right corner coordinate corresponding to the first corner point, stride is the downsampling multiple of the first-level network, scale is the scaling ratio of the image corresponding to the first corner point, 16 is the input size of the first-level network, and can also be other values, which can be customized by the user.

[0116] 4) Taking the deduplication process of images corresponding to multiple first corner points based on the NMS algorithm as an example, the deduplication operation is as follows:

[0117] Based on the first probability corresponding to the image corresponding to each first corner point, the multiple images are sorted in reverse order, and the set obtained by the reverse sorting is recorded as V;

[0118] Put the image with the highest probability in set V into set D and delete it from set V;

[0119] Traverse each image V in the set V i , get V i and each image D in set D i The intersection-over-union (IOU) ratio between them, if there is at least one IOU greater than the third threshold, V iDelete from the set V, if multiple IOUs are not greater than the third threshold, V i Put it into set D and delete V i .

[0120] The above two operations can be repeated until the set V is empty.

[0121] 5) Perform the above operations 2) to 4) on the multiple first palm images included in each layer of the image pyramid to deduplicate the images corresponding to the multiple first corner points output by all the first palm images, thereby obtaining multiple first images. The multiple first images can be recorded as a set S.

[0122] According to the palm area detection method provided in the embodiments of the present application, the initial palm image is scaled using different scaling ratios to construct an image pyramid. In practical applications, even when the distance between the palm and the acquisition terminal changes and the proportion of the palm in the palm image varies, the first-level network can still accurately output potential corner point targets in the palm image, thereby improving the prediction precision and accuracy of the first-level network. By performing non-maximum suppression and deduplication processing on the multiple first palm images included in each layer of the image pyramid, the number of inputs to the next-level network can be reduced, reducing data redundancy and thereby improving the network's inference speed.

[0123] Operation 120: Based on the second-level network, each candidate corner point feature is identified by type, and multiple first images are screened and processed based on the identified corner point types to obtain multiple second images; the second images include target corner point features and target position information corresponding to the target corner point features, and the target corner point features correspond to the corner point types; the second-level network is trained based on the second training set.

[0124] In this operation, the corner point types may include five categories, as shown in A, B, C, D, and N in FIG2 .

[0125] The plurality of first images are screened based on the corner point type, and first images of type N may be filtered out to obtain a plurality of second images.

[0126] The second image includes target corner point features and target position information corresponding to the target corner point features.

[0127] The target position information is used to represent the position of the target corner feature in the initial palm image.

[0128] The corner point type corresponding to the target corner point feature can be A, B, C or D, which can be identified based on the second-level network.

[0129] The second training set is sample data used to train the second-level network.

[0130] The second training set may include sample images and label information corresponding to the sample images.

[0131] The number of sample images can be multiple.

[0132] The label information corresponding to the sample image may include the corner point category corresponding to each feature in the sample image.

[0133] As shown in FIG. 2 , the label information corresponding to the sample image may include: A, B, C, D or N.

[0134] The input terminal of the second-stage network can be connected to the output terminal of the first-stage network.

[0135] As shown in FIG5 , the second-level network may be a convolutional neural network or any other neural network that can implement this function, and this application does not limit this.

[0136] In the case where the second-stage network is a convolutional neural network, the second-stage network may include a maximum pooling layer, a convolution (conv) layer, a fully connected layer (Fully Connected Layer) classification layer and a coordinate regression layer.

[0137] The output of the max pooling layer is connected to the input of the convolutional layer.

[0138] The output of the convolutional layer is connected to the input of the fully connected layer.

[0139] The output of the fully connected layer is connected to the input of the classification layer and the coordinate regression layer respectively.

[0140] During the training of the second-level network, the classification layer can be trained using the softmax activation function, and the convolutional layer can be trained using the relu activation function.

[0141] As shown in Figure 5, during the training of the second-level network, the training samples of the second-level network need to be scaled to 64*64. The size of the first sample image can be scaled to 64*64*1, and the first sample image is input to the first-level maximum pooling layer. The output image of the first-level maximum pooling layer can be obtained, and the size is 31*31*32.

[0142] Then the image output by the first-level maximum pooling layer is input to the second-level maximum pooling layer, and the image output by the second-level maximum pooling layer can be obtained, with a size of 14*14*64;

[0143] Then the image output by the second-level maximum pooling layer is input to the third-level maximum pooling layer, and the image output by the third-level maximum pooling layer can be obtained, with a size of 6*6*64;

[0144] Input the image output by the third-level maximum pooling layer into the convolution layer to obtain the image output by the convolution layer, with a size of 4*4*128;

[0145] The image output by the convolutional layer is input into the fully connected layer to obtain the output image of the fully connected layer with a size of 256.

[0146] In the actual implementation process, as shown in Figure 4, a channel attention module can be set between every two adjacent convolutional layers.

[0147] For example, an average pooling layer, a full connection layer, and a scale layer can be set between two convolutional layers.

[0148] The output of the average pooling layer is connected to the input of the fully connected layer.

[0149] The output of the fully connected layer is connected to the input of the scaling layer.

[0150] In some embodiments, operation 120 may include:

[0151] Based on the input size of the second-stage network, scaling the plurality of first images to obtain a plurality of fourth images;

[0152] Inputting the plurality of fourth images into the second-level network, obtaining at least one second candidate corner point image corresponding to each corner point type output by the second-level network and second predicted offset information corresponding to the second candidate corner point in each second candidate corner point image; each second candidate corner point image corresponds to a second probability;

[0153] When the second probability is greater than the second threshold, determining the second candidate corner point corresponding to the second probability as the second corner point;

[0154] The second predicted pixel position corresponding to the second corner point is corrected based on the second predicted offset information corresponding to the second corner point, and multiple second images are obtained based on the image corresponding to the corrected second corner point; the second predicted pixel position is obtained by processing the initial palm image based on the first-level network.

[0155] In this embodiment, the first image can be cut out from the initial palm image and then scaled to the input size of the second-level network to obtain the fourth image.

[0156] Each corner point type corresponds to at least one second candidate corner point image.

[0157] The second probability is used to represent the probability that the second candidate corner point image is of the corner point type.

[0158] For example, a corner point of type A may correspond to at least one third candidate corner point image, and a corner point of type B may correspond to at least one fourth candidate corner point image. The second probability corresponding to the third candidate corner point image can be used to characterize the probability that the image is a corner point of type A, and the second probability corresponding to the fourth candidate corner point image can be used to characterize the probability that the image is a corner point of type B.

[0159] The second prediction offset information can be used to characterize the offset between the coordinates of the upper left corner and the lower right corner of the area where the second candidate corner point is located predicted by the second-level network and the coordinates of the upper left corner and the lower right corner of the actual corner point area.

[0160] Based on the second prediction offset information, the second predicted pixel position corresponding to the second corner point predicted by the first-level network can be corrected.

[0161] The second predicted pixel position is obtained by processing the initial palm image based on the first-level network.

[0162] The second threshold can be used to determine the corner point category to which the second candidate corner point image belongs. For example, when the second probability between the second candidate corner point image and the Class A corner point is greater than the second threshold, the second candidate corner point image can be determined as the image corresponding to the Class A corner point.

[0163] When the second probability is greater than the second threshold, the second candidate corner point may be determined as the second corner point.

[0164] The corner point types corresponding to the multiple second corner points may be different.

[0165] After the coordinates of the position information of the second corner point are corrected, a plurality of corner points ultimately used for region extraction can be obtained, and then the region image where the corner points are located is determined as the fifth image.

[0166] During actual execution, the first image may be cut out from the palm image and then scaled to the input size of 64 of the second-level network to obtain at least one fourth image.

[0167] Input at least one fourth image into the second-level network to obtain the classification feature plane F cls2 , whose size is 5*1, and the classification feature plane includes at least one second candidate corner point image corresponding to 5 categories.

[0168] For example, by inputting at least one fourth network into the second-level network, the second candidate corner point images corresponding to the A, B, C, D and N categories output by the second-level network can be obtained.

[0169] At least one fourth image is input into the second-level network, and the coordinate feature plane F output by the second-level network can also be obtained. reg2, whose size is 4*1, and the coordinate feature plane includes the offset of the coordinates of the upper left corner and the lower right corner of the target box where each second candidate corner point is located relative to the real box.

[0170] The second threshold is set to TH2, and the second candidate corner point corresponding to the second probability greater than the second threshold is determined as the second corner point.

[0171] Get the second position offset information corresponding to the second corner point (XL offset ,YL offset ,XR offset ,YR offset ), and based on the second prediction offset information, the coordinates of the second predicted pixel position corresponding to the second corner point are corrected: W=X4-X3+1 H=Y4-Y3+1 X5=X3+W*XL offset Y5=Y3+H*YL offset X6=X t +W*XR offset Y6=Y4+H*YR offset

[0172] Among them, (X3, Y3) is the upper left corner coordinate of the second corner point output by the first level network on the palm image, (X4, Y4) is the lower right corner coordinate of the second corner point output by the first level network on the palm image, (X5, Y5) is the upper left corner coordinate of the second corner point on the palm image after correction, (X6, Y6) is the lower right corner coordinate of the second corner point on the palm image after correction, W is the length of the target box where the second corner point is located, H is the width of the target box where the second corner point is located, (XL offset ,YL offset ) is the predicted offset information of the upper left corner coordinate corresponding to the second corner point, (XR offset ,YR offset ) is the predicted offset information of the lower right corner coordinates corresponding to the second corner point.

[0173] Based on the coordinate information of the second corner point on the palm image, the final points A, B, C and D used for region extraction are obtained.

[0174] Operation 130 : Based on the corner point type corresponding to each target corner point feature and the target position information, the initial palm image is region-localized to determine the target palm region; the target palm region is a region of the palm that includes at least part of the palm print feature.

[0175] In this operation, the target palm area is an area of ​​the palm that includes at least part of the palm print features. For example, the target palm area may be an area where palm prints are concentrated, as shown in the area ROI in FIG6 .

[0176] In practical applications, the target palm area can be identified to complete the palm biometric recognition.

[0177] Each target corner feature corresponds to target position information, and the target position information is used to represent the coordinate information of the target corner feature in the initial palm image.

[0178] Based on the corner point type corresponding to the target corner point feature and the target position information, the region of interest (ROI) in the initial palm image can be located, and the ROI area is determined as the target palm area.

[0179] In this application, both the first-level network and the second-level network are lightweight neural networks. For example, the volume of the first-level network can be 12k, the volume of the second-level network can be 500k, and the embedded platform can be controlled within 10ms.

[0180] As shown in Figure 7, the intermediate layer convolutions of the first-level network and the second-level network can both use channel-separated convolutions. On the basis of ensuring the prediction accuracy of the network, it also improves the network's inference speed, further improving the user experience when performing palm biometric recognition.

[0181] In the actual implementation process, the first-level network and the second-level network can be trained separately, and the same training parameters can be used to train the first-level network and the second-level network.

[0182] For example, during training, the number of samples selected for one training (batchSize) can be 60, the process of training all training samples once (epoch) can be 10, the initial learning rate (learning rate) can be 0.01, and MomentumOptimizer can be used as the optimizer, with the learning rate decreasing by an order of magnitude at [5, 7, 9] epochs respectively.

[0183] The classification layer losses in the first-level network and the second-level network can both use cross entropy loss:

[0184] Among them, y is the input label label (one-hot form), is the classification probability (first probability or second probability) output by the network, and n is the number of classifications.

[0185] The coordinate regression layer losses of the first-level network and the second-level network can both use squared error loss:

[0186] Among them, b is the offset between the box coordinates of the input sample and the labeled box coordinates, and b^ is the network output value.

[0187] The loss function of the final network is the weighted sum of the classification layer loss and the coordinate regression layer loss:

[0188] In the process of training the first-level network, α can be 0.7 and β can be 0.3; in the process of training the second-level network, α can be 0.5 and β can be 0.5; is the coordinate regression layer loss, is the classification layer loss.

[0189] In this application, by setting different parameters to train the loss functions of the first-level network and the second-level network respectively, the first-level network can be trained to find more potential corner targets from the initial palm image, and the second-level network can be enabled to accurately classify multiple corner targets and accurately regress to the position coordinates of the corner targets.

[0190] According to the palm area detection method provided in the embodiment of the present application, the initial palm image is feature extracted through the first-level network, and multiple potential corner point targets can be roughly detected in the entire palm image, and the multiple potential corner point targets are input into the second-level network. The second-level network can accurately judge and adjust the positions of the multiple potential corner point targets to obtain multiple target corner points, so that the target palm area can be extracted from the initial palm image based on the target corner points. The target palm area is detected by the first-level network and the second-level network, and the target palm area can be accurately detected even when the palm posture changes. The detection precision and accuracy are high, and while ensuring the accuracy, the network reasoning speed is also improved. The palm area can be extracted from the palm image more quickly, thereby meeting the speed requirements of palm recognition.

[0191] In some embodiments, operation 130 may include:

[0192] Based on the corner point type corresponding to each target corner point feature and the target position information corresponding to each target corner point feature, a fitting process is performed on multiple first corner points of the target type to obtain a target straight line;

[0193] Based on a perpendicular line between a second corner point other than the first corner points in the plurality of target corner point features and the target straight line, a rotation angle between the perpendicular line and a longitudinal axis corresponding to the initial palm image is calculated;

[0194] Correcting the initial palm image based on the rotation angle to obtain a second palm image, corrected target corner features in the second palm image, and target position information corresponding to the corrected target corner features;

[0195] Based on the target position information corresponding to the corrected target corner point features, the initial palm image is region-localized to determine the target palm area.

[0196] In this embodiment, the target types may include: a corner point between the index finger and the middle finger, a corner point between the middle finger and the ring finger, and a corner point between the ring finger and the little finger.

[0197] By fitting the three first corner points, the target straight line can be obtained.

[0198] The least squares method may be used to perform straight line fitting on the three first corner points, or the fitting may be performed based on other methods, which are not limited in this application.

[0199] The second corner point may be the corner point closest to the base of the thumb.

[0200] The second palm image is an image obtained by correcting the initial palm image. For example, the initial palm image can be rotated toward the target direction based on the rotation angle to obtain the second palm image. In the second palm image, the perpendicular line between the second corner point and the target straight line is parallel to the longitudinal axis corresponding to the second palm image.

[0201] Based on the position information of the second corner point, the left-right attribute of the palm can be determined. For example, if the palm is a right palm, the target direction can be clockwise.

[0202] In actual implementation, as shown in FIG6 , the multiple first corner points may be point B, point C, and point D.

[0203] The target straight line can be expressed as y=bx+a, where b is the slope of the target straight line and a is the intercept of the target straight line.

[0204] Substitute the coordinates of points B, C, and D into the error equation And by minimizing the error, we can get the slope b and intercept a.

[0205] Draw a perpendicular line l from point A to the line y, and you can calculate the angle θ between the perpendicular line l and the x-axis, which is the rotation angle of the palm.

[0206] Based on the angle θ, the initial palm image is rotated to obtain the second palm image P θ , and record each corner point in the second palm image as

[0207] Point-based and point The location information of the point is calculated and point The distance Dis between them.

[0208] Based on the distance Dis, the coordinates of the upper left corner of the target palm area are obtained: in, for point The coordinates of point Dis and point the distance between them;

[0209] And the coordinates of the lower right corner of the target palm area: in, for point The coordinates of point Dis and point The distance between them.

[0210] The sample contents of the first training set and the second training set may be the same, but the sizes of the pictures included in the two training sets are different. The following takes obtaining the first training set as an example to specifically describe the method of obtaining the training set.

[0211] In some embodiments, obtaining the first training set may include:

[0212] Get a sample palm image;

[0213] Manually labeling the sample palm images to obtain a plurality of first positive sample images and a plurality of first negative sample images; the first positive sample images include sample corner features and first labels corresponding to the sample corner features; the first negative sample images include background features and second labels corresponding to the background features;

[0214] Randomly determine the area to be marked from the sample palm image;

[0215] Calculating the area intersection-and-union ratios between the region to be marked and each first positive sample image respectively to obtain a plurality of first area intersection-and-union ratios; calculating the area intersection-and-union ratios between the region to be marked and each first negative sample image respectively to obtain a plurality of second area intersection-and-union ratios;

[0216] Based on multiple first area intersection-to-union ratios, multiple second area intersection-to-union ratios, and a target ratio threshold, each area to be marked is marked to obtain multiple second positive sample images, multiple second negative sample images, and multiple intermediate sample images; the second positive sample images include sample corner features and first labels corresponding to the sample corner features; the second negative sample images include background features and second labels corresponding to the background features; and the intermediate sample images include sample corner features and background features;

[0217] A first training set is obtained based on a plurality of first positive sample images, a plurality of first negative sample images, a plurality of second positive sample images, a plurality of second negative sample images, and a plurality of intermediate sample images; the image sizes of the plurality of first positive sample images, the plurality of first negative sample images, the plurality of second positive sample images, the plurality of second negative sample images, and the plurality of intermediate sample images are the same as the input size of the first-level network.

[0218] In this embodiment, the sample palm image is a palm image in the sample data.

[0219] The marking information corresponding to the sample palm image is marking information corresponding to each feature on the sample palm image. For example, the marking information corresponding to the sample palm image may include corner feature marks and background feature marks.

[0220] There can be multiple sample palm images.

[0221] As shown in FIG2 , the sample palm image is manually marked. Multiple features may be marked on the sample palm image. For example, a first mark may be performed on the corner feature, and a second mark may be performed on the background feature.

[0222] As shown in FIG2 , a second mark may be performed on the area where black and white alternate.

[0223] During the manual marking process, a second mark may be manually performed on features such as the finger gaps of a closed palm that have a high similarity to the corner point area.

[0224] The first positive sample image may include the sample corner feature and a first label corresponding to the sample corner feature. For example, the first positive sample image may include the sample corner feature and a “is a corner” label.

[0225] The first negative sample image may include background features and a second label corresponding to the background features. For example, the first negative sample image may include background features and a “not a corner point” label.

[0226] The area to be marked is an area that needs to be marked, and the area to be marked needs to be marked with the first mark or the second mark.

[0227] There can be multiple areas to be marked.

[0228] A square with a side length of not less than 64 pixels may be randomly generated from the sample palm image, and the square is determined as the area to be marked.

[0229] The Intersection of Union (IOU) is used to characterize the degree of overlap between the true box and the predicted box. The value range of IOU is [0,1]. The larger the IOU, the higher the degree of overlap between the true box and the predicted box.

[0230] For example, as shown in Figure 10, the true box can be recorded as region A, the predicted box can be recorded as region B, and the overlapping area between the predicted box and the true box can be recorded as region C. The area intersection ratio between region A and region B can be determined based on the following formula:

[0231] Where IOU is the intersection-over-union ratio between region A and region B.

[0232] The first area intersection-to-union ratio is an area intersection-to-union ratio between the to-be-marked region and the first positive sample image.

[0233] The second area intersection-to-union ratio is an area intersection-to-union ratio between the to-be-marked region and the first negative sample image.

[0234] The target ratio threshold can be 0.3, 0.7, or other values, and can be user-defined and is not limited in this application. By calculating the area intersection-over-union ratio between the area to be marked and the first positive sample image, multiple second positive sample images, multiple second negative sample images, and multiple intermediate sample images can be automatically generated near the manually annotated first positive sample image.

[0235] By calculating the intersection-over-union ratio between the area to be labeled and the first negative sample image, a plurality of second negative sample images can be automatically generated near the manually labeled first negative sample image.

[0236] The plurality of second positive sample images and the plurality of first positive sample images may be used as positive sample data for training the first-level network.

[0237] The plurality of second negative sample images and the plurality of first negative sample images may be used as negative sample data for training the first-level network.

[0238] Multiple intermediate sample images can be used as partial sample data for training a first-level network. The intermediate sample images include sample corner features and background features. The first-level network can be trained based on offset information between the intermediate sample images and positive samples corresponding to the sample corner features, so that first position offset information is output based on the first-level network. In the first training set, the number of positive samples, the number of negative samples, and the number of intermediate samples are equal, or the difference between the three is less than a target number threshold, and the numbers of the three types of samples are in a balanced state.

[0239] In some embodiments, based on multiple first area intersection-to-union ratios, multiple second area intersection-to-union ratios, and a target ratio threshold, marking each to-be-marked region to obtain multiple second positive sample images, multiple second negative sample images, and multiple intermediate sample images may include:

[0240] When the multiple first area intersection-to-union ratios are not less than the first ratio threshold, determining the mark corresponding to the to-be-marked area as the first mark to obtain a second positive sample image;

[0241] When the first area intersection-to-union ratios are all less than the second ratio threshold, the mark corresponding to the area to be marked is determined as the second mark to obtain a second negative sample image; the second ratio threshold is less than the first ratio threshold;

[0242] When the plurality of first area intersection-to-union ratios are not less than the third ratio threshold and less than the first ratio threshold, the area to be marked is determined as an intermediate sample image; the third ratio threshold is greater than the second ratio threshold and less than the first ratio threshold;

[0243] When the multiple second area intersection-to-union ratios are not less than the first ratio threshold, the mark corresponding to the to-be-marked area is determined as the second mark to obtain a second negative sample image.

[0244] In this embodiment, the first ratio threshold, the second ratio threshold, and the third ratio threshold may all be different.

[0245] The third ratio threshold is greater than the second ratio threshold, and the third ratio threshold is less than the first ratio threshold.

[0246] The values ​​of the first ratio threshold, the second ratio threshold and the third ratio threshold can all be user-defined. For example, the first ratio threshold can be 0.7, the second ratio threshold can be 0.3, the third ratio threshold can be 0.64, or can be other values, which are not limited in this application.

[0247] When the first area intersection-union ratios between the region to be marked and each first positive sample image are not less than the first ratio threshold, it indicates that the degree of overlap between the region to be marked and the first positive sample image is high. The mark corresponding to the region to be marked can be determined as the first mark, and the second positive sample image can be obtained based on the region to be marked and the first mark.

[0248] When the first area intersection-union ratios between the region to be marked and each first positive sample image are all less than the second ratio threshold, it indicates that the degree of overlap between the region to be marked and the first positive sample image is low. The mark corresponding to the region to be marked can be determined as the second mark, and the second negative sample image can be obtained based on the region to be marked and the second mark.

[0249] When the first area intersection-union ratios between the region to be marked and each first positive sample image are not less than the third ratio threshold and less than the first ratio threshold, it means that the region to be marked contains most corner features, and an intermediate sample image can be obtained based on the region to be marked.

[0250] A second area intersection-under-union ratio between the region to be marked and each first negative sample image can also be calculated. When multiple second area intersection-under-union ratios are not less than the first ratio threshold, it indicates that the degree of overlap between the region to be marked and the first negative sample image is high. The mark corresponding to the region to be marked can be determined as the second mark, and the second negative sample image can be obtained based on the region to be marked and the second mark.

[0251] In the actual implementation process, as shown in Figure 2, the four corner points can be marked as A, B, C and D respectively, the four corner points and their corresponding marks are determined as positive samples, and the areas outside the positive samples where there is alternating black and white are marked as negative samples.

[0252] For each sample palm image, based on the annotation file, multiple first positive sample images, recorded as PositiveRects, and multiple first negative sample images, recorded as NegativeRects, can be obtained.

[0253] By intercepting the area where PositiveRects is located and scaling it to 16*16 size, we can obtain the second positive sample image, which is determined as the positive sample training data of the first-level network; by scaling the area where PositiveRects is located to 64*64 size, we can obtain the third positive sample image, which is determined as the positive sample training data of the second-level network.

[0254] By intercepting the area where NegativeRects is located and scaling it to 16*16 size, we can get the second negative sample image, which is determined as the negative sample training data of the first-level network; by scaling the area where NegativeRects is located to 64*64 size, we can get the third negative sample image, which is determined as the negative sample training data of the second-level network.

[0255] A square with a side length of not less than 64 pixels can be randomly generated in the sample palm image, recorded as Rect1, and the IOU between Rect1 and each first positive sample image in the positive sample rectangular area set PositiveRects is calculated;

[0256] The second ratio threshold is set to 0.3. When IOU is less than 0.3, the area where Rect1 is located can be intercepted and scaled to 16*16 to obtain the second negative sample image. The second negative sample image is used as the negative sample data for training the first-level network; the area is scaled to 64*64 to obtain the third negative sample image. The third negative sample image is used as the negative sample data for training the second-level network. Repeat the above operation until the number of negative samples generated for each level of the network is greater than 50.

[0257] A square with a side length of not less than 64 pixels can be randomly generated in the sample palm image, recorded as Rect2, and the IOU between Rect2 and each first positive sample image in the positive sample rectangle set PositiveRects is calculated;

[0258] The first ratio threshold is set to 0.7. When IOU ≥ 0.7, the area where Rect2 is located can be intercepted and scaled to 16*16 to obtain the second positive sample image. The second positive sample image is used as the positive sample data for training the first-level network; the area is scaled to 64*64 to obtain the third positive sample image. The third positive sample image is used as the positive sample data for training the second-level network, and the coordinate offset of the generated sample relative to each first positive sample image is saved. Repeat the above operation until the number of positive samples generated for each level of the network is greater than 150.

[0259] A square with a side length of not less than 64 pixels can be randomly generated in the sample palm image, recorded as Rect3, and the IOU between Rect3 and each first positive sample image in the positive sample rectangle set PositiveRects is calculated;

[0260] Set the first ratio threshold to 0.7 and the third ratio threshold to 0.64. When IOU ≥ 0.64 and IOU < 0.7, the area where Rect3 is located can be intercepted and scaled to 16*16 size, and used as the intermediate sample data for training the first-level network; scale the area to 64*64 size, and use it as the intermediate sample data for training the second-level network, and save the coordinate offset of the generated sample relative to each first positive sample image. Repeat the above operation until the number of intermediate samples generated for each level of the network is greater than 150.

[0261] A square with a side length of not less than 64 pixels can be randomly generated in the sample palm image, recorded as Rect4, and the IOU between Rect4 and each first negative sample image in the negative sample rectangle set NegativeRects is calculated;

[0262] The first ratio threshold is set to 0.7. When IOU ≥ 0.7, the area where Rect4 is located can be intercepted and scaled to 16*16 to obtain the second negative sample data, which is used as the negative sample data for training the first-level network; the area is scaled to 64*64 to obtain the third negative sample data, which is used as the negative sample data for training the second-level network. Repeat the above operation until the number of negative samples generated for each level of the network is greater than 100.

[0263] According to the palm area detection method provided in the embodiment of the present application, positive samples and negative samples are first manually marked, and then the area intersection and union ratio between the to-be-marked area in the sample palm image and the positive samples and negative samples is calculated. This can generate more positive sample data, intermediate sample data, and negative sample data near the positive samples or negative samples. The combination of manual marking and automatic generation provides a large number of valuable training samples for the network. By marking the negative samples, false detection will not occur even when the palm is closed, thereby improving the detection precision and accuracy of the network.

[0264] The palm area detection device provided in the present application is described below. The palm area detection device described below and the palm area detection method described above can be referenced to each other.

[0265] The palm area detection method provided in the embodiment of the present application can be executed by a palm area detection device. In the embodiment of the present application, the palm area detection device provided in the embodiment of the present application is described by taking the palm area detection device executing the palm area detection method as an example.

[0266] An embodiment of the present application also provides a palm area detection device.

[0267] As shown in FIG8 , the palm area detection device includes a first processing module 810 , a second processing module 820 and a third processing module 830 .

[0268] A first processing module 810 is configured to extract features from the acquired initial palm image based on a first-level network to obtain a plurality of first images; the first images include candidate corner features; and the first-level network is trained based on a first training set.

[0269] The second processing module 820 is configured to identify the type of each candidate corner feature based on the second-level network, and filter the plurality of first images based on the identified corner types to obtain a plurality of second images; the second images include target corner features and target position information corresponding to the target corner features, and the target corner features correspond to the corner types; the second-level network is trained based on the second training set;

[0270] The third processing module 830 is used to perform region positioning on the initial palm image based on the corner point type corresponding to each target corner point feature and the target position information to determine the target palm area; the target palm area is the area in the palm that includes at least part of the palm print features.

[0271] According to the palm area detection device provided in the embodiment of the present application, the initial palm image is feature extracted through the first-level network, and multiple potential corner point targets can be roughly detected in the entire palm image, and the multiple potential corner point targets are input into the second-level network. The second-level network can accurately judge and adjust the positions of the multiple potential corner point targets to obtain multiple target corner points, so that the target palm area can be extracted from the initial palm image based on the target corner points. The target palm area is detected by the first-level network and the second-level network, and the target palm area can be accurately detected even when the palm posture changes. The detection precision and accuracy are high, and while ensuring the accuracy, the network reasoning speed is also improved. The palm area can be extracted from the palm image more quickly, thereby meeting the speed requirements of palm recognition.

[0272] In some embodiments, the first processing module 810 may also be configured to:

[0273] Based on the initial palm image and the first-level network, position prediction is performed on multiple first candidate corner points in the initial palm image to obtain first predicted pixel positions and first predicted offset information corresponding to each first candidate corner point;

[0274] Based on the plurality of first prediction offset information and target parameters corresponding to the first-level network, the first predicted pixel positions are corrected to obtain a plurality of first images; the target parameters include: at least one of a downsampling multiple and an input size.

[0275] In some embodiments, the first processing module 810 may also be configured to:

[0276] Based on multiple scaling ratios, the initial palm image is scaled to obtain multiple first palm images of different sizes;

[0277] Inputting a target palm image from the plurality of first palm images into a first-stage network, obtaining a plurality of first candidate corner point images output by the first-stage network, first predicted offset information corresponding to the first candidate corner points in each of the first candidate corner point images, and first predicted pixel positions corresponding to each of the first candidate corner points; each of the first candidate corner point images has a first probability corresponding thereto;

[0278] When the first probability is greater than the first threshold, determining the first candidate corner point corresponding to the first probability as the first corner point;

[0279] Based on the downsampling multiple corresponding to the first-level network, the scaling ratio corresponding to the target palm image, the first prediction offset information corresponding to each first corner point, and the input size corresponding to the first-level network, the first predicted pixel position corresponding to each first corner point is corrected to obtain a second predicted pixel position of each first corner point on the initial palm image;

[0280] Based on the multiple second predicted pixel positions, deduplication processing is performed on images corresponding to the multiple first corner points to obtain multiple first images.

[0281] In some embodiments, the second processing module 820 may also be configured to:

[0282] Based on the input size of the second-stage network, scaling the plurality of first images to obtain a plurality of fourth images;

[0283] Inputting the plurality of fourth images into the second-level network, obtaining at least one second candidate corner point image corresponding to each corner point type output by the second-level network and second predicted offset information corresponding to the second candidate corner point in each second candidate corner point image; each second candidate corner point image corresponds to a second probability;

[0284] When the second probability is greater than the second threshold, determining the second candidate corner point corresponding to the second probability as the second corner point;

[0285] The second predicted pixel position corresponding to the second corner point is corrected based on the second predicted offset information corresponding to the second corner point, and multiple second images are obtained based on the image corresponding to the corrected second corner point; the second predicted pixel position is obtained by processing the initial palm image based on the first-level network.

[0286] In some embodiments, the first processing module 810 may also be configured to:

[0287] Get a sample palm image;

[0288] Manually labeling the sample palm images to obtain a plurality of first positive sample images and a plurality of first negative sample images; the first positive sample images include sample corner features and first labels corresponding to the sample corner features; the first negative sample images include background features and second labels corresponding to the background features;

[0289] Randomly determine the area to be marked from the sample palm image;

[0290] Calculating the area intersection-and-union ratios between the region to be marked and each first positive sample image respectively to obtain a plurality of first area intersection-and-union ratios; calculating the area intersection-and-union ratios between the region to be marked and each first negative sample image respectively to obtain a plurality of second area intersection-and-union ratios;

[0291] Based on multiple first area intersection-to-union ratios, multiple second area intersection-to-union ratios, and a target ratio threshold, each area to be marked is marked to obtain multiple second positive sample images, multiple second negative sample images, and multiple intermediate sample images; the second positive sample images include sample corner features and first labels corresponding to the sample corner features; the second negative sample images include background features and second labels corresponding to the background features; and the intermediate sample images include sample corner features and background features;

[0292] A first training set is obtained based on a plurality of first positive sample images, a plurality of first negative sample images, a plurality of second positive sample images, a plurality of second negative sample images, and a plurality of intermediate sample images; the image sizes of the plurality of first positive sample images, the plurality of first negative sample images, the plurality of second positive sample images, the plurality of second negative sample images, and the plurality of intermediate sample images are the same as the input size of the first-level network.

[0293] In some embodiments, the first processing module 810 may also be configured to:

[0294] When the multiple first area intersection-to-union ratios are not less than the first ratio threshold, determining the mark corresponding to the to-be-marked area as the first mark to obtain a second positive sample image;

[0295] When the first area intersection-to-union ratios are all less than the second ratio threshold, the mark corresponding to the area to be marked is determined as the second mark to obtain a second negative sample image; the second ratio threshold is less than the first ratio threshold;

[0296] When the plurality of first area intersection-to-union ratios are not less than the third ratio threshold and less than the first ratio threshold, the area to be marked is determined as an intermediate sample image; the third ratio threshold is greater than the second ratio threshold and less than the first ratio threshold;

[0297] When the multiple second area intersection-to-union ratios are not less than the first ratio threshold, the mark corresponding to the to-be-marked area is determined as the second mark to obtain a second negative sample image.

[0298] In some embodiments, the third processing module 830 may also be configured to:

[0299] Based on the corner point type corresponding to each target corner point feature and the target position information corresponding to each target corner point feature, a fitting process is performed on multiple first corner points of the target type to obtain a target straight line;

[0300] Based on a perpendicular line between a second corner point other than the first corner points in the plurality of target corner point features and the target straight line, a rotation angle between the perpendicular line and a longitudinal axis corresponding to the initial palm image is calculated;

[0301] Correcting the initial palm image based on the rotation angle to obtain a second palm image, corrected target corner features in the second palm image, and target position information corresponding to the corrected target corner features;

[0302] Based on the target position information corresponding to the corrected target corner point features, the initial palm image is region-localized to determine the target palm area.

[0303] The palm area detection device in the embodiment of the present application can be an electronic device or a component in an electronic device, such as an integrated circuit or chip. The electronic device can be a terminal or other device other than a terminal. For example, the electronic device can be a mobile phone, a tablet computer, a laptop computer, a PDA, a car-mounted electronic device, a mobile Internet device (MID), an augmented reality (AR) / virtual reality (VR) device, a robot, a wearable device, an ultra-mobile personal computer (UMPC), a netbook or a personal digital assistant (PDA), etc. It can also be a server, a network attached storage (NAS), a personal computer (PC), a television (TV), an ATM or a self-service machine, etc., and the embodiment of the present application does not specifically limit it. The palm area detection device in the embodiment of the present application can be a device with an operating system. The operating system can be an Android operating system, an iOS operating system, or other possible operating systems, and the embodiment of the present application does not specifically limit it.

[0304] The palm area detection device provided in the embodiment of the present application can implement each process implemented in the method embodiments of Figures 1 to 7. To avoid repetition, they are not described here.

[0305] In some embodiments, as shown in Figure 9, an embodiment of the present application also provides an electronic device 900, including a processor 901, a memory 902, and a computer program stored in the memory 902 and executable on the processor 901. When the program is executed by the processor 901, each process of the above-mentioned palm area detection method embodiment is implemented, and the same technical effect can be achieved. To avoid repetition, it will not be repeated here.

[0306] It should be noted that the electronic devices in the embodiments of the present application include the mobile electronic devices and non-mobile electronic devices mentioned above.

[0307] On the other hand, the present application also provides a computer program product, which includes a computer program stored on a non-transitory computer-readable storage medium, and the computer program includes program instructions. When the program instructions are executed by a computer, the computer can execute the various processes of the above-mentioned palm area detection method embodiment and can achieve the same technical effect. To avoid repetition, it will not be repeated here.

[0308] On the other hand, the present application also provides a non-transitory computer-readable storage medium on which a computer program is stored. When the computer program is executed by a processor, it is implemented to perform the various processes of the above-mentioned palm area detection method embodiment and can achieve the same technical effect. To avoid repetition, it will not be repeated here.

[0309] On the other hand, an embodiment of the present application further provides a chip, which includes a processor and a communication interface, the communication interface and the processor are coupled, and the processor is used to run programs or instructions to implement the various processes of the above-mentioned palm area detection method embodiment, and can achieve the same technical effect. To avoid repetition, it will not be repeated here.

[0310] It should be understood that the chip mentioned in the embodiments of the present application can also be referred to as a system-on-chip, a system-on-chip, a chip system, or a system-on-chip chip, etc. The device embodiments described above are merely illustrative, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed across multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the solution of this embodiment. A person of ordinary skill in the art can understand and implement the present invention without paying any creative effort.

[0311] Each module in the palm area detection device can be implemented in whole or in part by software, hardware, or a combination thereof. The network interface can be an Ethernet card or a network card. Each module can be embedded in or independent of a processor in a server in hardware form, or can be stored in a memory in a server in software form so that the processor can call and execute the operations corresponding to each of the above modules. As used in this application, the terms "component," "module," and "system" are intended to represent computer-related entities, which can be hardware, a combination of hardware and software, software, or software in execution. For example, a component can be, but is not limited to, a process running on a processor, a processor, an object, an executable code, an execution thread, a program, and / or a computer. As an illustration, both an application running on a server and a server can be components. One or more components can reside in a process and / or an execution thread, and a component can be located within a computer and / or distributed between two or more computers.

[0312] The technical features of the above-mentioned embodiments can be combined arbitrarily. In order to make the description concise, not all possible combinations of the technical features in the above-mentioned embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0313] The above-described embodiments merely represent several implementation methods of the present application. While the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the patent application. It should be noted that a person of ordinary skill in the art may make various modifications and improvements without departing from the spirit of the present application, and these modifications and improvements fall within the scope of protection of the present application. Therefore, the scope of protection of the present patent application shall be determined by the appended claims.

Claims

1. A palm region detection method, comprising: Performing feature extraction on an acquired initial palm image based on a first-level network to obtain multiple first images; The first images include candidate corner features; The first-level network is trained based on a first training set; Performing type recognition on each of the candidate corner features based on a second-level network, and performing screening processing on the multiple first images based on the recognized corner types to obtain multiple second images; the second images include target corner features and target position information corresponding to the target corner features, and the target corner features correspond to corner types; the second-level network is trained based on a second training set; Performing region positioning on the initial palm image based on the corner types corresponding to each of the target corner features and the target position information to determine a target palm region; the target palm region is a region in the palm that includes at least part of the palm print features.

2. The palm region detection method according to claim 1, wherein the performing feature extraction on the acquired initial palm image based on a first-level network to obtain multiple first images comprises: Based on the initial palm image and the first-level network, predicting the positions of multiple first candidate corners in the initial palm image to obtain first predicted pixel positions and first predicted offset information corresponding to each of the first candidate corners; Performing correction processing on the first predicted pixel positions based on multiple first predicted offset information and target parameters corresponding to the first-level network to obtain the multiple first images; The target parameters include at least one of a downsampling factor and an input size.

3. The palm region detection method according to claim 2, wherein the predicting the positions of multiple first candidate corners in the initial palm image based on the initial palm image and the first-level network to obtain first predicted pixel positions and first predicted offset information corresponding to each of the first candidate corners comprises: Performing scaling processing on the initial palm image based on multiple scaling ratios to obtain multiple first palm images with different sizes; Inputting a target palm image among the multiple first palm images into the first-level network to obtain multiple first candidate corner images output by the first-level network, first predicted offset information corresponding to the first candidate corners in each of the first candidate corner images, and first predicted pixel positions corresponding to each of the first candidate corners; Each of the first candidate corner images corresponds to a first probability; wherein the performing correction processing on the first predicted pixel positions based on multiple first predicted offset information and target parameters corresponding to the first-level network to obtain the multiple first images comprises: When the first probability is greater than a first threshold, determining the first candidate corner corresponding to the first probability as a first corner; Based on the downsampling multiple corresponding to the first-level network, the scaling ratio corresponding to the target palm image, the first predicted offset information corresponding to each of the first corner points, and the input size corresponding to the first-level network, correct the first predicted pixel positions corresponding to each of the first corner points to obtain the second predicted pixel positions of each of the first corner points on the initial palm image; Based on multiple second predicted pixel positions, perform duplicate removal processing on the images corresponding to multiple first corner points to obtain the multiple first images.

4. The palm region detection method according to any one of claims 1-3, wherein the type recognition of each candidate corner point feature is performed based on a second-level network, and the multiple first images are screened based on the recognized corner point types to obtain multiple second images, including: Based on the input size of the second-level network, perform scaling processing on the multiple first images to obtain multiple fourth images; Input the multiple fourth images into the second-level network to obtain at least one second candidate corner point image corresponding to each corner point type output by the second-level network and the second predicted offset information corresponding to the second candidate corner points in each second candidate corner point image; each second candidate corner point image corresponds to a second probability; In the case where the second probability is greater than a second threshold, determine the second candidate corner point corresponding to the second probability as a second corner point; Based on the second predicted offset information corresponding to the second corner point, correct the second predicted pixel position corresponding to the second corner point, and based on the image corresponding to the corrected second corner point, obtain multiple second images; the second predicted pixel position is obtained by processing the initial palm image based on the first-level network.

5. The palm region detection method according to any one of claims 1-3, wherein the first training set is obtained based on the following operations: Obtain a sample palm image; Manually mark the sample palm image to obtain multiple first positive sample images and multiple first negative sample images; the first positive sample image includes sample corner point features and a first mark corresponding to the sample corner point features; the first negative sample image includes background features and a second mark corresponding to the background features; Randomly determine a region to be marked from the sample palm image; Calculate the area intersection-over-union ratios between the region to be marked and each of the first positive sample images respectively to obtain multiple first area intersection-over-union ratios; Calculate the area intersection-over-union ratios between the region to be marked and each of the first negative sample images respectively to obtain multiple second area intersection-over-union ratios; Based on the multiple first area intersection-over-union ratios, the multiple second area intersection-over-union ratios, and a target ratio threshold, mark each region to be marked to obtain multiple second positive sample images, multiple second negative sample images, and multiple intermediate sample images; the second positive sample image includes sample corner point features and a first mark corresponding to the sample corner point features; the second negative sample image includes background features and a second mark corresponding to the background features; the intermediate sample image includes the sample corner point features and the background features; Obtain the first training set based on the multiple first positive sample images, the multiple first negative sample images, the multiple second positive sample images, the multiple second negative sample images, and the multiple intermediate sample images; The image sizes of the multiple first positive sample images, the multiple first negative sample images, the multiple second positive sample images, the multiple second negative sample images, and the multiple intermediate sample images are the same as the input size of the first-level network.

6. The palm region detection method according to claim 5, wherein the step of marking each of the regions to be marked based on the multiple first area intersection-over-union ratios, the multiple second area intersection-over-union ratios, and the target ratio threshold to obtain multiple second positive sample images, multiple second negative sample images, and multiple intermediate sample images includes: When all of the multiple first area intersection-over-union ratios are not less than the first ratio threshold, determine the mark corresponding to the region to be marked as the first mark to obtain the second positive sample image; When all of the multiple first area intersection-over-union ratios are less than the second ratio threshold, determine the mark corresponding to the region to be marked as the second mark to obtain the second negative sample image; The second ratio threshold is less than the first ratio threshold; When all of the multiple first area intersection-over-union ratios are not less than the third ratio threshold and less than the first ratio threshold, determine the region to be marked as the intermediate sample image; The third ratio threshold is greater than the second ratio threshold and less than the first ratio threshold; When all of the multiple second area intersection-over-union ratios are not less than the first ratio threshold, determine the mark corresponding to the region to be marked as the second mark to obtain the second negative sample image.

7. The palm region detection method according to any one of claims 1-3, wherein the step of performing region positioning on the initial palm image based on the corner types corresponding to the respective target corner features and the target position information to determine the target palm region includes: Perform fitting processing on multiple first corners of a target type based on the corner types corresponding to the respective target corner features and the target position information corresponding to the respective target corner features to obtain a target straight line; Calculate the rotation angle between the perpendicular line between the second corners other than the multiple first corners among the multiple target corner features and the target straight line and the vertical axis corresponding to the initial palm image; Correct the initial palm image based on the rotation angle to obtain a second palm image, the corrected target corner features in the second palm image, and the target position information corresponding to the corrected target corner features; Perform region positioning on the initial palm image based on the target position information corresponding to the corrected target corner features to determine the target palm region.

8. A palm region detection device, comprising: A first processing module, configured to perform feature extraction on an acquired initial palm image based on a first-level network to obtain multiple first images; The first images include candidate corner features; The first-level network is trained based on a first training set; A second processing module, configured to perform type recognition on each of the candidate corner features based on a second-level network, and perform screening processing on the multiple first images based on the recognized corner types to obtain multiple second images; the second images include target corner features and target position information corresponding to the target corner features, and the target corner features correspond to corner types; the second-level network is trained based on a second training set; A third processing module, configured to perform region positioning on the initial palm image based on the corner types corresponding to the target corner features and the target position information, and determine a target palm region; the target palm region is a region of the palm that includes at least part of the palmprint features.

9. The apparatus according to claim 8, wherein the first processing module is further configured to: Based on the initial palm image and the first-level network, perform position prediction on multiple first candidate corners in the initial palm image to obtain first predicted pixel positions and first predicted offset information corresponding to the first candidate corners; Based on a plurality of first prediction offset information and target parameters corresponding to the first-level network, perform correction processing on the first predicted pixel positions to obtain the plurality of first images; the target parameters include: At least one of the downsampling factor and the input size.

10. The apparatus according to claim 9, wherein the first processing module is further configured to: Perform scaling processing on the initial palm image based on multiple scaling ratios to obtain multiple first palm images with different sizes; Input the target palm image among multiple first palm images into the first-level network, and obtain multiple first candidate corner point images output by the first-level network, first prediction offset information corresponding to the first candidate corner points in each of the first candidate corner point images, and first prediction pixel positions corresponding to each of the first candidate corner points; Each of the first candidate corner images corresponds to a first probability; The correcting the first predicted pixel positions to obtain the multiple first images based on the multiple first predicted offset information and the target parameters corresponding to the first-level network includes: When the first probability is greater than a first threshold, determining the first candidate corner corresponding to the first probability as a first corner; Based on the downsampling factor corresponding to the first-level network, the scaling ratio corresponding to the target palm image, the first predicted offset information corresponding to each first corner, and the input size corresponding to the first-level network, correct the first predicted pixel positions corresponding to each first corner to obtain second predicted pixel positions of each first corner on the initial palm image; Perform duplicate removal processing on the images corresponding to the multiple first corners based on the multiple second predicted pixel positions to obtain the multiple first images.

11. An electronic device, comprising a memory, a processor; and A computer program stored on the memory and executable on the processor, and when the processor executes the program, the operations of the palm region detection method according to any one of claims 1-7 are implemented.

12. A computer program product, comprising a computer program, and when the computer program is executed by a processor, the operations of the palm region detection method according to any one of claims 1-7 are implemented.

Citation Information

Patent Citations

  • Palm and key point detection method, device and terminal device thereof

    CN109345553A

  • Identity recognition method and palmprint key point detection model training method and device

    CN112132099A

  • Palm vein image region-of-interest extraction method and device

    CN113963158A

  • Robust palm region-of-interest positioning method in natural scene

    CN115661872A

  • Palm area detection method and device

    CN117831082A