Object recognition method, electronic device, and storage medium
By employing a recognition model that performs quality inspection and adaptive feature compensation on palm print and palm vein images, the issues of recognition accuracy and stability in complex environments of the palm print and palm vein fusion recognition system are resolved, thereby improving the accuracy of object recognition.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- ZHEJIANG DAHUA TECH CO LTD
- Filing Date
- 2026-02-27
- Publication Date
- 2026-06-16
AI Technical Summary
Existing palmprint and palm vein fusion recognition systems suffer from decreased recognition accuracy and stability due to factors such as sensor malfunctions and environmental interference, resulting in low object recognition accuracy.
By acquiring palm print and palm vein images of the object to be identified, quality detection processing is performed on each image. Based on the detection results, the target image and prompt information are determined and input into the recognition model for recognition, thereby achieving adaptive feature compensation.
It improves the accuracy of object recognition, enhances the system's recognition stability and adaptability in complex environments, and avoids the decline in feature discrimination ability caused by poor image quality.
Smart Images

Figure CN122223757A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of image processing, and in particular to an object recognition method, electronic device, and storage medium. Background Technology
[0002] With the development of artificial intelligence technology, biometric features such as facial features, iris scans, and palm prints / palm veins have become digital identity cards for people entering the interconnected world. Palm print and palm vein biometric recognition technologies, with their uniqueness, security, and intelligent adaptability, have become leading solutions in the field of identity authentication. From a single-modal perspective, palm prints have advantages in feature richness and high compatibility. Palm prints contain complex textures such as main lines (lifeline, wisdom line, etc.), fine lines, wrinkles, and triangular areas, with a far greater number of feature points than fingerprints. Furthermore, palm prints can be captured using ordinary RGB cameras (such as mobile phone cameras), eliminating the need for dedicated infrared equipment and resulting in low hardware costs. However, palm prints are susceptible to surface contamination, making them less resistant to counterfeiting. Palm veins, on the other hand, offer significant advantages such as liveness detection and high anti-counterfeiting capabilities, high uniqueness and privacy, high environmental stability, and independence from skin condition. Their vascular patterns are hidden under the skin, making them extremely difficult to forge. However, palm vein imaging requires strict control of the acquisition posture and distance; low temperatures or low blood flow conditions can lead to unclear vein imaging. Palmprint and palm vein fusion recognition, by combining the complementary characteristics of the two modalities, significantly surpasses single-modal solutions in terms of security, user experience, and robustness. However, in real-world applications, factors such as sensor malfunctions, sensor defects in specific environments, lighting interference, and dirt occlusion can lead to low quality or missing modalities in one modality. This results in decreased accuracy and stability of the palmprint and palm vein fusion recognition system, leading to lower object recognition accuracy.
[0003] Therefore, there is an urgent need for an effective object recognition method. Summary of the Invention
[0004] This application provides at least one object recognition method, electronic device, and storage medium that can improve the accuracy of object recognition.
[0005] This application provides an object recognition method, which includes: acquiring an image of an object to be recognized, the image of which includes a palm print image and a palm vein image; performing quality detection processing on the palm print image and the palm vein image in the image to be recognized to obtain a quality detection result of the image to be recognized; determining a target image and target prompt information of the image to be recognized based on the quality detection result of the image to be recognized; and inputting the target image and the target prompt information of the image to be recognized into a recognition model to obtain the object recognition result output by the recognition model.
[0006] This application provides an object recognition device, comprising: an acquisition module, a quality detection module, a determination module, and a recognition module; the acquisition module is used to acquire an image of an object to be recognized, the image of which includes a palm print image and a palm vein image; the quality detection module is used to perform quality detection processing on the palm print image and the palm vein image in the image to be recognized respectively, to obtain a quality detection result of the image to be recognized; the determination module is used to determine a target image and target prompt information of the image to be recognized based on the quality detection result of the image to be recognized; the recognition module is used to input the target image and the target prompt information of the image to be recognized into a recognition model, to obtain an object recognition result output by the recognition model.
[0007] This application provides an electronic device, including a memory and a processor, wherein the processor is used to execute program instructions stored in the memory to implement the above-described object recognition method.
[0008] This application provides a computer-readable storage medium storing program instructions thereon, which, when executed by a processor, implement the above-described object recognition method.
[0009] The above scheme acquires the image of the object to be identified, which includes the palm print image and the palm vein image of the object to be identified. The palm print image and the palm vein image are respectively subjected to quality detection processing to obtain the quality detection result of the image to be identified, and the target image and the target prompt information of the image to be identified are determined. The target image and the target prompt information of the image to be identified are input into the recognition model to obtain the object recognition result output by the recognition model. In this way, the target prompt information corresponding to the quality detection result of the image to be identified can guide the recognition model to process the image information of the target image, thereby improving the accuracy of the recognition result output by the recognition model.
[0010] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit this application. Attached Figure Description
[0011] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with this application and, together with the specification, serve to explain the technical solutions of this application.
[0012] Figure 1 This is a flowchart illustrating an exemplary embodiment of the object identification method of this application; Figure 2 yes Figure 1 A schematic diagram of the sub-process of step S13; Figure 3 yes Figure 2 A schematic diagram of the sub-process of step S22; Figure 4 yes Figure 3 A schematic diagram of the sub-process of step S32; Figure 5 This is a schematic diagram of the framework of the identification model in one embodiment of the object identification method of this application; Figure 6 This is a schematic diagram of the structure of an embodiment of the object identification device of this application; Figure 7 This is a schematic diagram of the structure of an embodiment of the electronic device of this application; Figure 8 This is a schematic diagram of the structure of an embodiment of the computer-readable storage medium of this application. Detailed Implementation
[0013] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. It is understood that the specific embodiments described herein are only for explaining this application and not for limiting it. Furthermore, it should be noted that, for ease of description, only the parts related to this application are shown in the accompanying drawings, not all structures. All other embodiments obtained by those skilled in the art based on the embodiments of this application without creative effort are within the scope of protection of this application.
[0014] In the following description, specific details such as particular system architectures, interfaces, and technologies are presented for illustrative purposes rather than for limiting purposes, in order to provide a thorough understanding of this application.
[0015] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that includes a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus. The term "and / or" is merely a description of the association of related objects, indicating that three relationships can exist; for example, A and / or B can represent: A alone, A and B simultaneously, and B alone. Additionally, the character " / " in this document generally indicates that the preceding and following related objects are in an "or" relationship. Furthermore, "many" in this document means two or more. Furthermore, the term "at least one" in this document means any combination of at least two of any one or more of a plurality of elements, for example, including at least one of A, B, and C, which can mean including any one or more elements selected from the set consisting of A, B, and C. Additionally, the term "several" in this document means one or more.
[0016] This application provides several object recognition methods and devices. The application scenarios of these object recognition methods include, but are not limited to, image recognition scenarios or identity recognition scenarios for objects to be identified. The executing entity of the object recognition method can be an object recognition device, such as an object recognition model. For example, the object recognition device can be located in a terminal device, server, or other processing device. The terminal device can be a user equipment (UE), mobile device, user terminal, terminal, cellular phone, cordless phone, personal digital assistant (PDA), handheld device, computing device, vehicle-mounted device, etc. In some possible implementations, the object recognition method can be implemented by a processor calling computer-readable instructions stored in memory.
[0017] Please see Figure 1 , Figure 1 This is a flowchart illustrating an exemplary embodiment of the object identification method of this application. Specifically, the object identification method may include the following steps: Step S11: Obtain the image of the object to be identified.
[0018] The images to be identified include palm print images and palm vein images of the object to be identified. The object to be identified is an animal body that possesses palm print and palm vein features.
[0019] The images to be identified represent palm print images and palm vein images for the same object. There can be one or more palm print images. There can be one or more palm vein images.
[0020] Specifically, step S11 may involve acquiring an image of the object to be identified within a preset time period. For example, step S11 may involve acquiring a palmprint image of the palmprint modality of the object to be identified within a preset time period; and acquiring a palm vein image of the palm vein modality of the object to be identified within a preset time period.
[0021] In some application scenarios, step S11 can involve acquiring images of the palm print and palm veins of the object to be identified within a preset time period using an image acquisition device, and then preprocessing the acquired images to obtain the image to be identified. In other application scenarios, step S11 can involve retrieving images of the object to be identified from a preset database. In still other application scenarios, step S11 can involve acquiring images of the object to be identified within a preset time period, cropping the acquired images according to the region of interest to obtain a cropped image, and using the cropped image as the image to be identified. In yet another application scenario, step S11 can involve acquiring images of the object to be identified within a preset time period, including images of the palm print and palm veins, and aligning the images of the palm print and palm veins to obtain the corresponding palm print image and corresponding palm vein image.
[0022] Step S12: Perform quality detection processing on the palm print image and palm vein image in the image to be identified, and obtain the quality detection result of the image to be identified.
[0023] In some application scenarios, when the image to be identified is a palmprint image, the quality detection result of the image to be identified includes the quality detection result of the palmprint image. Step S12 above can be: performing quality detection processing on the palmprint image to obtain the quality detection result of the palmprint image. In other application scenarios, when the image to be identified is a palm vein image, the quality detection result of the image to be identified includes the quality detection result of the palm vein image. Step S12 above can be: performing quality detection processing on the palm vein image to obtain the quality detection result of the palm vein image. It is understood that the processing logic is the same or similar when the image to be identified is either a palmprint image or a palm vein image.
[0024] Quality inspection processing is a step that independently assesses the image quality of palm print and palm vein images within the image to be identified. The quality inspection result of the image to be identified indicates its overall image quality. This result includes a quality score, which indicates the image quality of the image on a preset evaluation metric. A higher quality score indicates higher image quality. The preset evaluation metric includes several sub-metrics.
[0025] In some application scenarios, step S12 above can involve performing a first quality detection on the palmprint image in the image to be recognized to obtain a quality detection result for the palmprint image; and performing a second quality detection on the palm vein image in the image to be recognized to obtain a quality detection result for the palm vein image. The preset evaluation index corresponding to the first quality detection and the preset evaluation index corresponding to the second quality detection share at least some evaluation sub-indicators. For example, the evaluation sub-indicators may include image clarity, image integrity, and the degree of image contamination in the image to be recognized.
[0026] In other application scenarios, an initial quality detection process is performed on the image to be recognized based on a first evaluation metric to obtain a first quality score; an initial quality detection process is then performed on the image to be recognized based on a second evaluation metric to obtain a second quality score; the first quality score and the second quality score are weighted and fused to obtain a total quality score, which is then used as the quality score of the image to be recognized. The evaluation sub-indicators in the first evaluation metric differ from those in the second evaluation metric. Specifically, the first evaluation metric includes at least one of the following: palm angle, palm exposure, palm brightness, palm and / or palm integrity, palm texture clarity, and palm texture wrinkling degree in the image to be recognized. The second evaluation metric includes at least one of the following: the distance between the palm of the object to be recognized and the acquisition device in the image to be recognized, the direction / continuity of the palm texture in the image to be recognized, and the percentage of the palm area in the image to be recognized. The palm angle characterizes the spatial relative relationship between the palm target in the image to be recognized and the acquisition device.
[0027] In other application scenarios, the image to be recognized undergoes quality detection processing corresponding to its acquisition quality to obtain a first quality score regarding acquisition quality. Then, the image to be recognized undergoes quality detection processing corresponding to its image content to obtain a second quality score regarding its image content. The first and second quality scores are then weighted and fused to obtain a total quality score, which is used as the overall quality score of the image to be recognized. The sub-indicators for acquisition quality include at least one of the following: sharpness, exposure, brightness, and completeness of the hand in the image to be recognized. The sub-indicators for image content include at least one of the following: hand angle, completeness of the hand and / or palm, clarity of the palm texture, degree of palm texture wrinkling, distance between the hand of the object being recognized and the acquisition device, direction / continuity of the palm texture in the image to be recognized, and percentage of the area of the palm region in the image to be recognized. The fusion weights for the first and second quality scores differ during the weighted fusion process.
[0028] In other application scenarios, step S12 can also involve inputting each palmprint image and each palm vein image in the image to be recognized into a quality assessment network, respectively, to obtain the quality detection results of each palmprint image and each palm vein image output by the quality assessment network. The quality assessment network is used to perform quality detection processing on the image to be recognized to obtain quality detection results. For example, the quality assessment network is pre-trained using training data, which includes sample palmprint images and sample palm vein images.
[0029] Step S13: Determine the target image and target prompt information for the image to be identified based on the quality detection results of the image to be identified.
[0030] The target image represents the image to be processed associated with the image to be identified. The image to be processed is at least one of a palm print image and a palm vein image. The target cue information of the image to be identified represents the cue information associated with the image to be identified and is used for feature generation of the image to be identified.
[0031] In some application scenarios, when the image to be identified is a palmprint image, the target image represents at least one of the palmprint image and the palm vein image, and the target prompt information of the image to be identified represents the prompt information associated with the palmprint image. Step S13 above can be: selecting the target image associated with the palmprint image from the palmprint image and the palm vein image based on the quality detection result of the palmprint image, including: selecting the target image associated with the palmprint image from the palmprint image and the palm vein image based on the quality score of the palmprint image. Obtaining the prompt information associated with the palmprint image based on the quality detection result of the palmprint image includes: selecting the target prompt information associated with the palmprint image from the preset palmprint generation prompt information or candidate palmprint generation prompt information corresponding to the palmprint image based on the quality score of the palmprint image. The preset palmprint generation prompt information is used to indicate prompt information for feature completion of the palmprint modality. The candidate palmprint generation prompt information represents prompt information for not performing feature completion of the palmprint modality; for example, the candidate palmprint generation prompt information can be empty. It is understood that step S13 above can only determine the target image and the target prompt information of the palmprint image when the image to be identified is a palmprint image.
[0032] In other application scenarios, when the image to be identified is a palm vein image, the target image represents at least one of a palm print image and a palm vein image, and the target prompt information of the image to be identified represents the prompt information associated with the palm vein image. Step S13 above may include: selecting the target image associated with the palm vein image from the palm vein image and the palm vein image based on the quality detection result of the palm vein image, including: selecting the target image associated with the palm vein image from the palm print image and the palm vein image based on the quality score of the palm vein image. Obtaining the prompt information associated with the palm vein image based on the quality detection result of the palm vein image includes: selecting the target prompt information associated with the palm vein image from the preset palm vein generation prompt information or candidate palm vein generation prompt information corresponding to the palm vein image based on the quality score of the palm vein image. The preset palm vein generation prompt information is used to indicate prompt information for feature completion of the palm vein modality. The candidate palm vein generation prompt information represents prompt information for not performing feature completion of the palm vein modality; for example, the candidate palm vein generation prompt information can be empty. It is understood that step S13 above may only determine the target image and the target prompt information of the palm vein image when the image to be identified is a palm vein image.
[0033] In other application scenarios, step S13 above may be, when the image to be identified is a palm print image, determining the target image of the palm print image and the target prompt information of the palm print image; and, when the image to be identified is a palm vein image, determining the target image of the palm vein image and the target prompt information of the palm vein image.
[0034] Step S14: Input the target image and the target prompt information of the image to be recognized into the recognition model to obtain the object recognition result output by the recognition model.
[0035] The recognition model is used to process palm print images and / or palm vein images in the image to be recognized based on target images and target cue information in the image to be recognized, thereby obtaining the object recognition result. The object recognition result represents the identity information of the object to be recognized.
[0036] The recognition model includes at least one functional module, each module responsible for processing palmprint and / or palm vein images in the image to be recognized to obtain object recognition results. The recognition model includes at least a feature extraction module, which extracts target features from the palmprint and / or palm vein images in the image to be recognized. Target information from the target image and the image to be recognized is connected to the recognition model via an input interface. Based on this target information, the recognition model performs adaptive feature compensation on the palmprint and / or palm vein images to obtain the target features of the palmprint and / or palm vein images.
[0037] In some application scenarios, the target image includes the target image of a palmprint image. Step S14 can be: inputting the target image of the palmprint image and the target prompt information of the palmprint image into the recognition model to obtain the object recognition result output by the recognition model. Specifically, based on the target image of the palmprint image and the target prompt information of the palmprint image, the extracted features corresponding to the target image are obtained, and these extracted features are used as the target palmprint features corresponding to the palmprint image; based on the target palmprint features corresponding to the palmprint image, the object recognition result is determined.
[0038] In other application scenarios, the target image includes the target image of the palm vein image. Step S14 can then be: inputting the target image of the palm vein image and the target prompt information of the palm vein image into the recognition model to obtain the object recognition result output by the recognition model. Specifically, based on the target image of the palm vein image and the target prompt information of the palm vein image, the extracted features corresponding to the target image are obtained, and these extracted features are used as the target palm vein features corresponding to the palm vein image; the object recognition result is determined based on the target palm vein features corresponding to the palm vein image.
[0039] In other application scenarios, the target image includes the target image of the palmprint image and the target image of the palm vein image. Step S14 can be: inputting the target image of the palmprint image, the target prompt information of the palmprint image, the target image of the palm vein image, and the target prompt information of the palm vein image into the recognition model to obtain the object recognition result output by the recognition model. This application takes this application scenario as an example. Specifically, based on the target image of the palmprint image and the target prompt information of the palmprint image, the extracted features corresponding to the target image are obtained, and these extracted features are used as the target palmprint features corresponding to the palmprint image; based on the target image of the palm vein image and the target prompt information of the palm vein image, the extracted features corresponding to the target image are obtained, and these extracted features are used as the target palm vein features corresponding to the palm vein image; based on the target palmprint features corresponding to the palmprint image and the target palm vein features corresponding to the palm vein image, the object recognition result is determined.
[0040] In other application scenarios, classification processing of target palmprint features yields a first recognition result for the object to be identified, representing the identity information of the object determined based on the target palmprint features. Classification processing of target palm vein features yields a second recognition result for the object to be identified, representing the identity information of the object determined based on the target palm vein features. The first and second recognition results are fused based on the quality score corresponding to the target palmprint image and the quality score of the target palm vein image to obtain the object recognition result for the object to be identified. This includes: in response to the first and second recognition results being the same, using either the first or second recognition result as the object recognition result for the object to be identified; in response to the first and second recognition results being different, determining a target quality score from the quality score corresponding to the target palmprint image and the quality score of the target palm vein image, including: using the maximum quality score between the quality score corresponding to the target palmprint image and the quality score of the target palm vein image as the target quality score; using the first recognition result of the target palmprint image corresponding to the target quality score as the object recognition result for the object to be identified, or using the second recognition result of the target palm vein image corresponding to the target quality score as the object recognition result for the object to be identified.
[0041] In other application scenarios, feature fusion processing is performed on the target palmprint features and target palm vein features to obtain the fused features corresponding to the object to be identified; or, based on the quality scores of the target palmprint image and the target palm vein image, the target palmprint features and palm vein features are fused to obtain the fused features corresponding to the object to be identified. The fused features corresponding to the object to be identified are then subjected to classification processing to obtain the object recognition result.
[0042] For example, the object recognition device first captures images of two modalities (palmprint modality and palm vein modality) of the object to be recognized using a binocular camera. Each captured image is preprocessed to obtain an image to be recognized (the image to be recognized contains both palmprint images of the palmprint modality and palm vein images of the palm vein modality). Then, quality detection processing is performed on the image to be recognized to obtain a quality detection result. Finally, the quality scores of the image to be recognized and the quality detection results are input into a recognition model for feature fusion and identity determination to obtain the object recognition result. For example, this application can use a convolutional neural network as a quality detection module to perform local texture analysis on the image to be recognized to obtain a quality detection result. In the recognition model, a quality level can be dynamically set based on a preset threshold. When the quality score is lower than the threshold, a feature compensation mechanism is triggered to compensate for the features of the image to be recognized to obtain the target features (target palmprint features and / or target palm vein features). Based on the target features of the image to be recognized, the object recognition result of the object to be recognized is determined.
[0043] It can be considered that the quality detection process enables real-time evaluation of the image quality of the palmprint and palm vein modal images to be identified. This allows for adaptive adjustment of the recognition strategy in cases of low-quality or missing modalities, avoiding a decline in feature discrimination capability due to poor image quality. The quality detection results guide the recognition model to effectively compensate for low-quality information, further improving the system's recognition stability and environmental adaptability in complex environments such as lighting interference and surface contamination, thereby enhancing the accuracy of object recognition results.
[0044] The above scheme acquires the image of the object to be identified, which includes the palm print image and the palm vein image of the object to be identified. The palm print image and the palm vein image are respectively subjected to quality detection processing to obtain the quality detection result of the image to be identified, and the target image and the target prompt information of the image to be identified are determined. The target image and the target prompt information of the image to be identified are input into the recognition model to obtain the object recognition result output by the recognition model. In this way, the target prompt information corresponding to the quality detection result of the image to be identified can guide the recognition model to process the image information of the target image, thereby improving the accuracy of the recognition result output by the recognition model.
[0045] In some embodiments, step S11 may include the following steps: acquiring a first image of the object to be identified by a palmprint acquisition device and a second image of the object to be identified by a palm vein acquisition device; extracting a region of interest (ROI) from the first image to obtain a first ROI; extracting a ROI from the second image to obtain a second ROI; and aligning the first and second images according to the first and second ROIs to obtain a palmprint image and a palm vein image, respectively.
[0046] The palmprint acquisition device represents an acquisition device used to acquire palmprint modalities of the object to be identified. The first acquired image represents an acquired image obtained by acquiring palmprint modalities of the object to be identified at a first preset time. The first region of interest represents the palmprint region obtained by extracting the region of interest from the palmprint target in the first acquired image. The palm vein acquisition device represents an acquisition device used to acquire palm vein modalities of the object to be identified. The second acquired image represents an acquired image obtained by acquiring palm vein modalities of the object to be identified at a second preset time. The second region of interest represents the palm vein region obtained by extracting the region of interest from the palm vein target in the second acquired image. The first preset time and the second preset time can be the same time or adjacent times within a preset time period. For example, the palmprint acquisition device and the palm vein acquisition device are respectively the white light path acquisition device and the infrared acquisition device in a binocular camera. The alignment process represents image registration of the first and second acquired images at the same spatial location. The palmprint image in the image to be identified represents the first acquired image after alignment. The palm vein image in the image to be identified represents the second acquired image after alignment. The palmprint target in the palmprint image and the palm vein target in the palm vein image are located at the same spatial location.
[0047] For example, step S11 above may specifically include: performing palm detection and key point localization processing on the first acquired image to obtain target key points of the palm in the first acquired image; performing palm detection and key point localization processing on the second acquired image to obtain target key points of the palm in the second acquired image; aligning the first acquired image and the second acquired image according to the target key points in the first acquired image and the target key points in the second acquired image to obtain a palm print image corresponding to the first acquired image and a palm vein image corresponding to the second acquired image.
[0048] For example, palmprint data and palm vein data of the object to be identified are collected to obtain a first acquired image and a second acquired image, respectively. A near-infrared binocular camera with white light and a fixed wavelength is constructed. The near-infrared binocular camera includes a near-infrared camera with a fixed wavelength (i.e., a palm vein acquisition device) and a white light acquisition device (i.e., a palmprint acquisition device). The second acquired image containing palm features is acquired using the fixed wavelength near-infrared camera, and the first acquired image containing palmprint features is acquired using the white light path. During the process of palm detection and hand key point localization on the first or second acquired image, a preset network (such as YOLOv8) is used to perform palm detection and hand key point localization on the two acquired infrared and white light images, respectively. The specific palm detection and hand key point localization results are as follows: the coordinates of the palm and the key points between the index finger and middle finger, the key points between the middle finger and ring finger, and the key points between the ring finger and little finger are located. If a palm target is detected in the image to be identified (i.e., the palm detection result indicates the presence of a palm target), the quality of the palm in the two modalities (palmprint modality and palm vein modality) is evaluated separately. The image to be identified is then input into the quality evaluation network, which outputs the quality score of the palm target in each modality.
[0049] For example, taking the palm vein modality as an example, the palm vein quality assessment process is described, specifically including the following: The palm target is matted based on the palm coordinate box located in the second acquired image, and simultaneously scaled to a fixed size to obtain a palm vein image; the scaled-down, fixed-size palm target (i.e., the palm vein image) is input into the constructed quality assessment network, and the last layer of the network outputs the quality score of the palm vein target assessment. Similarly, the palm target is matted based on the palm coordinate box located in the first acquired image, and simultaneously scaled to a fixed size to obtain a palmprint image; the scaled-down, fixed-size palm target (i.e., the palmprint image) is input into the constructed quality assessment network, and the last layer of the network outputs the quality score of the palm vein target assessment. The quality assessment network can be pre-trained. The sample quality score of the sample image output by the quality assessment network is compared with the label quality score of the sample image using a loss function, which is an L2 regression loss function. Through gradient backpropagation, the quality assessment network is continuously trained until a fully trained quality assessment network is obtained. The construction of the label quality scores for sample images requires manual annotation. The overall sample scoring principle is as follows: three coarse scoring levels, with each level further refined into fine scoring levels. The three coarse scoring levels are set as follows: Level 1, low score: palm angle exceeds 20 degrees, palm is overexposed, palm is too dark, palm is incomplete, etc.; Level 2, middle score: palm angle is within 20 degrees, palm veins and palm print texture are relatively visible; Level 3, high score: palm is slightly open facing the camera, palm texture is wrinkle-free, palm veins and texture are clear. After the three coarse scoring levels, the scores for Level 2 and Level 3 are further refined within each interval according to the principle that the clearer the palm veins and texture, the higher the score. Region of Interest (ROI) detection: Target key points from the first and second acquired images are used to align and extract the ROIs of the two modalities of the palm, obtaining the corresponding palm print region and palm vein region to construct palm print images and palm vein images respectively.
[0050] It can be assumed that by aligning the palm print modality with the palm vein modality through palm detection and key point localization, the accuracy of subsequent feature extraction processing of the palm print and palm vein regions can be improved, and the effectiveness of feature completion can be enhanced when using feature completion between different modalities.
[0051] In some application scenarios, step S12 above can be: inputting the image to be recognized into a quality assessment network to obtain the quality detection result of the image to be recognized output by the quality assessment network. Specifically, inputting a palmprint image into the quality assessment network to obtain the quality detection result of the palmprint image output by the quality assessment network; inputting a palm vein image into the quality assessment network to obtain the quality detection result of the palm vein image output by the quality assessment network.
[0052] In other application scenarios, step S12 above can be: performing palm detection and key point localization processing on the image to be recognized to obtain the palm detection result of the image to be recognized; obtaining the quality detection result of the image to be recognized based on the palm detection result, specifically including: performing palm detection on the image to be recognized to obtain the palm detection result of the image to be recognized. If the palm detection result of the image to be recognized is empty, a preset quality detection result is set as the quality detection result of the image to be recognized, specifically including: if the palm detection result for a palmprint is empty, a preset score is set as the quality score in the quality detection result of the palmprint image; if the palm detection result for a palm vein image is empty, a preset score is set as the quality score in the quality detection result of the palm vein image. If the palm detection result of the image to be recognized is not empty, the image to be recognized is input into the quality assessment network to obtain the quality detection result of the image to be recognized output by the quality assessment network.
[0053] In some embodiments, the image to be identified is a palmprint image. Step S13 may include the following steps: If the quality detection result of the palmprint image indicates that the quality score of the palmprint image is greater than a first score threshold, then the target image is determined to be a palmprint image, and the target prompt information is empty. If the quality detection result of the palmprint image indicates that the quality score of the palmprint image is less than or equal to the first score threshold and greater than a second score threshold, then the target image is determined to be both a palmprint image and a palm vein image, and a preset palmprint generation prompt information is used as the target prompt information. If the quality detection result of the palmprint image indicates that the quality score of the palmprint image is less than or equal to the second score threshold, then the target image is determined to be a palm vein image, and the preset palmprint generation prompt information is used as the target prompt information.
[0054] When the image to be identified is a palmprint image, the target image and corresponding target prompt information are determined based on the quality score in the palmprint image quality detection result. The palmprint image quality score represents the quality score obtained by evaluating the palmprint image according to a preset quality assessment index. The higher the palmprint image quality score, the higher the image quality. A first score threshold is greater than a second score threshold. The preset palmprint prompt information is used to indicate the prompt information for feature completion of the palmprint modality.
[0055] If the quality score of the palmprint image is greater than the first score threshold, it indicates that a hand is present in the palmprint image and the image quality is high. If the quality score of the palmprint image is less than or equal to the first score threshold but greater than the second score threshold, it indicates that a hand is present in the palmprint image and the image quality is low. If the quality score of the palmprint image is less than or equal to the second score threshold, it indicates that a hand is not present in the palmprint image.
[0056] In some application scenarios, step S14 above can be: inputting the target image of the palm print image and the corresponding target prompt information into the recognition model to obtain the object recognition result output by the recognition model.
[0057] In some embodiments, the image to be identified is a palm vein image. Step S13 may include the following steps: If the quality detection result of the palm vein image indicates that the quality score of the palmprint image is greater than a first score threshold, then the target image is determined to be a palm vein image, and the target prompt information is empty. If the quality detection result of the palm vein image indicates that the quality score of the palmprint image is less than or equal to the first score threshold and greater than a second score threshold, then the target image is determined to be both a palmprint image and a palm vein image, and a preset palm vein generation prompt information is used as the target prompt information. If the quality detection result of the palm vein image indicates that the quality score of the palmprint image is less than or equal to the second score threshold, then the target image is determined to be a palmprint image, and the preset palm vein generation prompt information is used as the target prompt information.
[0058] When the image to be identified is a palm vein image, the target image and corresponding target prompt information are determined based on the quality score in the palm vein image quality detection results. The quality score of the palm vein image represents the quality score obtained by evaluating the palm vein image according to a preset quality assessment index. The higher the quality score of the palm vein image, the higher the image quality. A first score threshold is greater than a second score threshold. The preset palm vein prompt information is used to indicate the prompt information for feature completion of the palm vein modality.
[0059] If the quality score of the palm vein image is greater than the first score threshold, it indicates that a palm target is present in the palm vein image and the image quality is high. If the quality score of the palm vein image is less than or equal to the first score threshold but greater than the second score threshold, it indicates that a palm target is present in the palm vein image and the image quality is low. If the quality score of the palm vein image is less than or equal to the second score threshold, it indicates that a palm target is not present in the palm vein image.
[0060] In some application scenarios, step S14 above can be: inputting the target image of the palm vein image and the corresponding target prompt information into the recognition model to obtain the object recognition result output by the recognition model.
[0061] In other application scenarios, step S14 above can be: inputting the target image of the palm print image, the target prompt information corresponding to the palm print image, the target image of the palm vein image, and the target prompt information corresponding to the palm vein image into the recognition model to obtain the object recognition result output by the recognition model.
[0062] Please see Figure 2 , Figure 2 yes Figure 1 A schematic diagram of the sub-process of step S14.
[0063] In some embodiments, the recognition model includes a first feature extraction module and a second feature extraction module. Step S14 above may include the following steps: Step S21: Perform feature extraction processing on the target image through the first feature extraction module to obtain the initial features of the target image.
[0064] The first feature extraction module is used to extract low-dimensional features from the target image to obtain the initial features of the target image. The initial features of the target image represent the low-dimensional features of the hand target in the target image. For example, the first feature extraction module is equipped with an image segmentation sub-module, or an image segmentation sub-module and a preset convolutional layer (e.g., 1×1 convolution).
[0065] Specifically, step S21 above can be performed by the first feature extraction module as follows: the first feature extraction module can perform image block processing and flattening processing on the target image to obtain the features to be extracted from the target image; the features to be extracted from the target image can be directly used as the initial features of the target image; or, the features to be extracted from the target image can be subjected to preset convolution processing to obtain the initial features of the target image, wherein the preset convolution processing can be 1×1 convolution processing.
[0066] In some application scenarios, regardless of whether the image to be identified is a palm print image or a palm vein image, if the target image of the image to be identified only includes a palm print image, the first feature extraction module performs the following steps: The first feature extraction module may perform image block processing and flattening processing on the palm print image to obtain the extracted features of the palm print image; the extracted features of the palm print image are directly used as the initial features of the palm print image; or, the extracted features of the palm print image are subjected to a preset convolution processing to obtain the initial features of the palm print image, wherein the preset convolution processing may be a 1×1 convolution processing.
[0067] In some application scenarios, regardless of whether the image to be identified is a palm print image or a palm vein image, when the target image only includes a palm vein image, the first feature extraction module performs the following steps: The first feature extraction module may perform image block processing and flattening processing on the palm vein image to obtain the extracted features of the palm vein image; the extracted features of the palm vein image are directly used as the initial features of the palm vein image; or, the extracted features of the palm vein image are subjected to a preset convolution processing to obtain the initial features of the palm vein image, wherein the preset convolution processing may be a 1×1 convolution processing.
[0068] It is understandable that, regardless of whether the image to be identified is a palmprint image or a palm vein image, if the target image of the image to be identified can include both palmprint and palm vein images, the palmprint and palm vein images in the target image are respectively input into the first feature extraction module to obtain the initial features of the target image. The initial features of the target image include the initial features of the palmprint image and the initial features of the palm vein image. In some application scenarios, the first feature extraction module can be equipped with a shared feature extraction network corresponding to the palmprint modality and the palm vein modality, or a feature extraction network corresponding to each modality. The palmprint image and the palm vein image are respectively input into the first feature extraction module to obtain the initial features of the palmprint image and the initial features of the palm vein image output by the first feature extraction module. It is understandable that, in the case of a shared feature extraction network, if there are identical images between the target images of the corresponding images to be identified when the images to be identified are different, in order to save computing resources, the initial features output by the target image of one image to be identified after passing through the shared feature extraction network can be used as the initial features of the target image of the other image to be identified. For example, when the image to be identified is a palm print image, the target image of the palm print image includes the palm print image. When the image to be identified is a palm vein image, the target image of the palm vein image includes both the palm print image and the palm vein image. The initial features of the palm vein image output by the shared feature extraction network can be used as the shared features for the two cases respectively, without repeated calculation.
[0069] Step S22: The initial features and target prompt information of the target image are processed by the second feature extraction module to obtain the object recognition result output by the recognition model.
[0070] The second feature extraction module is used to process the initial features of the target image and the target prompt information of the corresponding image to be recognized to obtain the object recognition result.
[0071] Specifically, step S22 may involve fusing the initial features of the target image of the image to be recognized with the target prompt information of the corresponding image to be recognized to obtain the target features of the image to be recognized; classifying the target features of the image to be recognized to obtain the candidate recognition results of the object to be recognized; and determining the object recognition result based on the candidate recognition results of the object to be recognized.
[0072] In some application scenarios, the candidate recognition result includes the first recognition result obtained based on the target features of the palmprint image. The above step S22 may be: when the image to be recognized is a palmprint image, the initial features of the target image of the palmprint image and the target prompt information of the palmprint image are fused to obtain the target features of the palmprint image (i.e., target palmprint features); the target features of the palmprint image are classified to obtain the first recognition result of the object to be recognized; and the object recognition result is determined based on the first recognition result.
[0073] In other application scenarios, the candidate recognition result includes a second recognition result obtained based on the target features of the palm vein image. Step S22 above may be: when the image to be recognized is a palm vein image, the initial features of the target image of the palm vein image and the target prompt information of the palm vein image are fused to obtain the target features of the palm vein image (i.e., target palm vein features); the target features of the palm vein image are classified to obtain the second recognition result of the object to be recognized; and the object recognition result is determined based on the second recognition result.
[0074] In other application scenarios, the candidate recognition results include a first recognition result obtained from the target features of the palmprint image and a second recognition result obtained from the target features of the palm vein image. Step S22 may be as follows: when the image to be recognized is a palmprint image, the initial features of the target image of the palmprint image and the target prompt information of the palmprint image are fused to obtain the target features of the palmprint image; the target features of the palmprint image are classified to obtain the first recognition result of the object to be recognized; and when the image to be recognized is a palm vein image, the initial features of the target image of the palm vein image and the target prompt information of the palm vein image are fused to obtain the target features of the palm vein image; the target features of the palm vein image are classified to obtain the second recognition result of the object to be recognized; and the object recognition result is determined based on the first recognition result and the second recognition result.
[0075] Specifically, the step of determining the object recognition result based on the first recognition result and the second recognition result may be: using the first recognition result as the object recognition result; or, using the second recognition result as the object recognition result; or, determining the object recognition result based on the quality score of the palmprint image, the quality score of the palm vein image, the first recognition result, and the second recognition result, including: in response to the quality score of the palmprint image being greater than the quality score of the palm vein image, using the first recognition result as the object recognition result; in response to the quality score of the palmprint image being less than or equal to the quality score of the palm vein image, using the second recognition result as the object recognition result.
[0076] The step of classifying the target features of the image to be identified to obtain candidate recognition results can involve inputting the target features of the image to be identified into a large language model / preset knowledge graph for similarity comparison, and using the preset image corresponding to the preset feature with the highest similarity to the target feature as the candidate recognition result. In the process of fusing the first and second recognition results to obtain the object recognition result, the fusion weight of the corresponding candidate recognition result can be determined by combining the quality score of the image to be identified with the similarity weight between the preset database objects and the object to be identified in the corresponding candidate recognition results. Taking a palm vein image as an example, the target palm vein features are classified to obtain the second recognition result of the object to be identified. The fusion weight of the second recognition result is determined by weighting the quality score of the palm vein image with the similarity weight between the preset database objects and the object to be identified in the second recognition result. It is understood that the method for determining the fusion weight of the first recognition result is similar to the method for determining the fusion weight of the second recognition result, and will not be elaborated here. The above-mentioned method of fusing the first recognition result and the second recognition result based on the quality score of the palm print image and the quality score of the palm vein image to obtain the object recognition result of the object to be recognized may further include: responding to the first recognition result and the second recognition result being the same, using either the first recognition result or the second recognition result as the object recognition result of the object to be recognized; responding to the first recognition result and the second recognition result being different, selecting a target fusion weight from the fusion weight of the first recognition result and the fusion weight of the second recognition result, and using the recognition result corresponding to the target fusion weight as the object recognition result of the object to be recognized. Specifically, if the target fusion weight is the fusion weight of the first recognition result, the first recognition result is used as the object recognition result. If the target fusion weight is the fusion weight of the second recognition result, the second recognition result is used as the object recognition result.
[0077] Please see Figure 3 , Figure 3 yes Figure 2 A schematic diagram of the sub-process of step S22.
[0078] In some embodiments, the second feature extraction module includes a feature completion module and a third feature extraction module. Step S22 above may include the following steps: Step S31: The initial features and target prompt information of the target image are processed by the feature completion module to obtain advanced features.
[0079] The feature completion module is used to complete the initial features of the image to be recognized based on the target prompt information, thereby obtaining the advanced features of the image to be recognized. The advanced features of the image to be recognized represent the corresponding completed initial features. It can be understood that, depending on the specific content of the target prompt information, the advanced features of the image to be recognized are the result of processing the initial features of the image to be recognized using different completion methods.
[0080] In some application scenarios, when the image to be identified is a palmprint image and the quality score of the palmprint image is greater than a first score threshold, the initial features of the target image of the palmprint image include the initial features of the palmprint image, and the target prompt information of the image to be identified is empty. In this case, step S31 above can be: processing the initial features of the palmprint image through the feature completion module to obtain the advanced features of the palmprint image; or, when the image to be identified is a palmprint image and the quality score of the palmprint image is less than or equal to the first score threshold and greater than the second score threshold, the initial features of the target image of the palmprint image include the initial features of the palmprint image and the initial features of the palm vein image, and the target prompt information of the image to be identified is a preset palm vein image. In the case of palm print generation prompt information, step S31 above can be: processing the initial features of the palm print image, the initial features of the palm vein image, and the preset palm print generation prompt information through the feature completion module to obtain the advanced features of the palm print image; or, when the image to be identified is a palm print image and the quality score of the palm print image is less than or equal to the second score threshold, and the initial features of the target image of the palm print image include the initial features of the palm vein image, and the target prompt information of the image to be identified is the preset palm print generation prompt information, step S31 above can be: processing the initial features of the palm vein image and the preset palm print generation prompt information through the feature completion module to obtain the advanced features of the palm print image.
[0081] In other application scenarios, when the image to be identified is a palm vein image and the quality score of the palm vein image is greater than a first score threshold, the initial features of the target image of the palm vein image include the initial features of the palm vein image, and the target prompt information of the image to be identified is empty. In this case, step S31 can be: processing the initial features of the palm vein image through the feature completion module to obtain the advanced features of the palm vein image; or, when the image to be identified is a palm vein image and the quality score of the palm vein image is less than or equal to the first score threshold and greater than the second score threshold, the initial features of the target image of the palm vein image include the initial features of the palm vein image and the initial features of the palm print image, and the target prompt information of the image to be identified is a preset... In the case of palm vein generation prompt information, step S31 above may be: processing the initial features of the palm vein image, the initial features of the palm print image, and the preset palm vein generation prompt information through the feature completion module to obtain the advanced features of the palm vein image; or, when the image to be identified is a palm vein image and the quality score of the palm vein image is less than or equal to the second score threshold, and the initial features of the target image of the palm vein image include the initial features of the palm print image, and the target prompt information of the image to be identified is the preset palm vein generation prompt information, step S31 above may be: processing the initial features of the palm print image and the preset palm vein generation prompt information through the feature completion module to obtain the advanced features of the palm vein image.
[0082] In other application scenarios, step S31 above can be used when the images to be identified are palmprint images and palm vein images, respectively, to determine the advanced features of the palmprint image and the advanced features of the palm vein image. The process of determining the advanced features of the palmprint image and the palm vein image separately can refer to the above content, and will not be repeated here. It can be understood that in the feature completion module, when the palmprint image and the palm vein image are used as the images to be identified respectively, the construction of the advanced features of the palmprint image and the palm vein image can be carried out in a separate serial manner or in parallel.
[0083] The feature completion module may be equipped with a preset feature extraction network. Specifically, the preset feature extraction network may be a convolutional neural network, a recurrent neural network, a long short-term memory network, a gated recurrent unit, a BiGRU neural network, an attention mechanism, or other feature extraction networks. In other application scenarios, the preset feature extraction network may be a preset perceptron network. The preset perceptron network may be obtained from at least one perceptron module. Each of the at least one perceptron module may be a multilayer perceptron (MLP).
[0084] In some embodiments, step S31 may be any one of steps S41 to S43, and / or any one of steps S44 to S46; wherein, step S41: the initial features of the palmprint image are processed by the feature completion module to obtain the advanced features of the palmprint image. Step S42: the initial features of the palmprint image, the initial features of the palm vein image, and the preset palmprint generation prompt information are processed by the feature completion module to obtain the advanced features of the palmprint image. Step S43: the initial features of the palm vein image and the preset palmprint generation prompt information are processed by the feature completion module to obtain the advanced features of the palmprint image. Step S44: the initial features of the palm vein image are processed by the feature completion module to obtain the advanced features of the palm vein image. Step S45: the initial features of the palm vein image, the initial features of the palmprint image, and the preset palm vein generation prompt information are processed by the feature completion module to obtain the advanced features of the palm vein image. Step S46: The initial features of the palm print image and the preset palm vein generation prompt information are processed by the feature completion module to obtain the advanced features of the palm vein image.
[0085] In other application scenarios, a quality score greater than a first score threshold indicates high image quality and high reliability of the initial features of the image to be identified. When the quality score of the image to be identified is greater than the first score threshold, using the feature completion module to take the initial features of the image to be identified as the advanced features of the image to be identified can be as follows: Step S41 above is: taking the initial features of the palmprint image as the advanced features of the palmprint image; and / or, Step S44 above is: taking the initial features of the palm vein image as the advanced features of the palm vein image.
[0086] In other application scenarios, a quality score less than or equal to a first score threshold indicates low image quality and low reliability of the initial features. The convolution result obtained by the feature completion module through convolution processing of the target image and the preset generated prompt information of the modality to which the image belongs is used as an advanced feature of the image to be recognized; or, the convolution result is mapped to obtain a mapping result, which is then used as an advanced feature of the image to be recognized. It is understood that steps S42, S43, S45, and S46 can respectively involve first performing convolution processing on the target image and the preset generated prompt information of the modality to which the image belongs, and directly using the result of the convolution processing as the corresponding advanced feature, or mapping the result of the convolution processing to obtain the corresponding advanced feature. Further details are omitted here.
[0087] The second score threshold is less than the first score threshold. In some application scenarios, a quality score of the image to be identified that is less than or equal to the second score threshold indicates that the modality to which the image belongs is missing, that is, there is no hand target in the image to be identified, the initial features of the image to be identified are absent, or it indicates a high degree of missing hand target, and the initial features of the image to be identified are unreliable. In this case, step S31 may only be executed as step S43 and / or only as step S46. In other application scenarios, a quality score of the image to be identified that is greater than the second score threshold and less than or equal to the first score threshold indicates that the image quality of the image to be identified is low, and the reliability of the initial features of the image to be identified is low. When the quality score of the image to be identified is greater than the second score threshold and less than or equal to the first score threshold, step S31 may only be executed as step S42 and / or only as step S45.
[0088] It is understood that step S31 above can be achieved by using the feature completion module to take the palm print image and the palm vein image as images to be identified respectively, and performing any one of steps S41 to S43 on the palm print image to obtain advanced features of the palm print image and / or performing any one of steps S44 to S46 on the palm vein image to obtain advanced features of the palm vein image.
[0089] Step S32: The advanced features are processed by the third feature extraction module to obtain the object recognition result output by the recognition model.
[0090] The third feature extraction module is used to extract high-dimensional features from the image to be recognized, and then determine the object recognition result based on these high-dimensional features. This third feature extraction module may be equipped with a preset feature extraction network. Specifically, the preset feature extraction network may be a convolutional neural network, a recurrent neural network, a long short-term memory network, a gated recurrent unit, a BiGRU neural network, an attention mechanism, or other feature extraction networks. In some application scenarios, the preset feature extraction network may be a preset perceptron network. This preset perceptron network can be obtained from at least one perceptron module. Each of the at least one perceptron module may be a multilayer perceptron (MLP).
[0091] Please see Figure 4 , Figure 4 yes Figure 3 A schematic diagram of the sub-process of step S32.
[0092] In some embodiments, the third feature extraction module includes a preset feature extraction module and an object determination module. Step S32 may include the following steps: Step S51: The preset feature extraction module performs feature extraction processing on the advanced features to obtain the target features. Step S52: The object determination module performs feature matching processing on the target features and the preset base database features to obtain the object recognition result.
[0093] When the image to be identified is a palmprint image and / or a palm vein image, the target features of the image to be identified include the target features of the palmprint image and / or the target features of the palm vein image. The preset base database features include preset palmprint features and / or preset palm vein features of several preset base database objects. The target features of the image to be identified represent the advanced features of the image to be identified after feature extraction processing.
[0094] Specifically, step S51 can be: when the image to be identified is a palmprint image, performing feature extraction processing on the advanced features of the palmprint image to obtain the target features of the palmprint image; or, when the image to be identified is a palm vein image, performing feature extraction processing on the advanced features of the palm vein image to obtain the target features of the palm vein image; or, when the image to be identified is a palmprint image, performing feature extraction processing on the advanced features of the palmprint image to obtain the target features of the palmprint image, and when the image to be identified is a palm vein image, performing feature extraction processing on the advanced features of the palm vein image to obtain the target features of the palm vein image.
[0095] In some embodiments, the step of performing feature matching processing on target features and preset database features by the object determination module to obtain object recognition results may include the following steps: using the similarity between the target features of the palmprint image and the preset palmprint features of each preset database object as the first similarity between the object to be identified and each preset database object; and / or using the similarity between the target features of the palm vein image and the preset palm vein features of each preset database object as the second similarity between the object to be identified and each preset database object. For each preset database object, obtaining the first target similarity between the object to be identified and the preset database object based on the first similarity and / or the second similarity between the object to be identified and the preset database object. Determining the object recognition result based on the first target similarity between the object to be identified and each preset database object.
[0096] The first similarity score characterizes the similarity between the target features of the palmprint image of the object to be identified and the preset palmprint features of objects in the preset database. The second similarity score characterizes the similarity between the target features of the palm vein image of the object to be identified and the preset palm vein features of objects in the preset database. The first target similarity score characterizes the similarity determined based on the first similarity and / or the second similarity between the object to be identified and objects in the preset database.
[0097] In some embodiments, the preset database features further include preset fusion features between preset palmprint features and preset palm vein features of each preset database object. The step of determining the object recognition result based on the first target similarity between the object to be identified and each preset database object may include the following steps: fusing the target features of the palmprint image and the target features of the palm vein image to obtain the fusion features of the object to be identified; using the similarity between the fusion features of the image to be identified and the preset fusion features corresponding to each preset database object as the third similarity between the object to be identified and each preset database object; for each preset database object, updating the first target similarity between the object to be identified and the preset database object based on the third similarity between the object to be identified and the preset database object to obtain the second target similarity between the object to be identified and the preset database object; and determining the object recognition result based on the second target similarity between the object to be identified and each preset database object.
[0098] The third similarity represents the similarity between the fused features of the object to be identified and the preset fused features of the preset base database objects. The second target similarity represents the similarity determined based on the first target similarity and the third similarity between the object to be identified and the preset base database objects.
[0099] It is understandable that the methods for determining the preset palmprint features, preset palm vein features, and preset fusion features for each preset base object are the same as the methods for determining the target features of the palmprint image, the target features of the palm vein image, and the fusion features of the object to be identified, and will not be elaborated here.
[0100] In some embodiments, the step of obtaining the first target similarity between the object to be identified and the preset database object based on the first similarity and / or second similarity between the object to be identified and the preset database object for each preset database object may include the following steps: determining the palmprint fusion weight of the palmprint image based on the quality detection result of the acquired palmprint image; determining the palm vein fusion weight of the palm vein image based on the quality detection result of the acquired palm vein image; and for each preset database object, performing weighted fusion of the first similarity, second similarity, palmprint fusion weight, and palm vein fusion weight between the object to be identified and the preset database object to obtain the first target similarity between the object to be identified and the preset database object.
[0101] In some embodiments, the third feature extraction module includes a preset feature extraction module. The process of determining the object recognition result through the third feature extraction module can be as follows: The preset feature extraction module performs feature extraction processing on the advanced features of the image to be recognized, obtaining the target features of the image to be recognized. The similarity between the target features of the palmprint image and the preset palmprint features of several preset database objects is used as the first similarity between the object to be recognized and each preset database object. The similarity between the target features of the palm vein image and the preset palm vein features of several preset database objects is used as the second similarity between the object to be recognized and each preset database object. Based on each first similarity and each second similarity, the recognition result of the object to be recognized is determined.
[0102] The target features of the image to be identified include target features of the palmprint image and target features of the palm vein image. In some application scenarios, the preset feature extraction module is a shared feature extraction layer that performs high-dimensional feature extraction on the advanced features of the palmprint image and the advanced features of the palm vein image, respectively. In other application scenarios, the preset feature extraction module includes a first preset extraction module for performing high-dimensional feature extraction on the advanced features of the palmprint image and a second preset extraction module for performing high-dimensional feature extraction on the advanced features of the palm vein image. This application takes as an example a preset feature extraction module that includes a first preset extraction module and a second preset extraction module that independently perform high-dimensional feature extraction on the advanced features of the palmprint image and the advanced features of the palm vein image. The first preset extraction module and the second preset extraction module can be two modules with the same structure but different parameters.
[0103] The advanced features of the image to be recognized are extracted by a preset feature extraction module to obtain the target features of the image to be recognized. This includes: extracting advanced features from the palmprint image to obtain the target features of the palmprint image; specifically, the advanced features of the palmprint image are input into a first preset extraction module to obtain the target features of the palmprint image output by the first preset extraction module. Similarly, the advanced features of the palm vein image are extracted from the palm vein image to obtain the target features of the palm vein image; specifically, the advanced features of the palm vein image are input into a second preset extraction module to obtain the target features of the palm vein image output by the second preset extraction module.
[0104] The preset base object can be the object to be compared. If the similarity between the object to be identified and the preset base object meets the requirements, the preset base object is used as the object identification result of the object to be identified. At this time, the object identification result of the object to be identified indicates that the object to be identified is the preset base object.
[0105] The preset database objects can be one or more preset database objects. The identity information, preset palm print features, and preset palm vein features of these preset database objects are stored in the preset database. As needed, the identity information, preset palm print features, and preset palm vein features of these preset database objects can be retrieved from the preset database.
[0106] It is understandable that during user registration, the collected images of each preset database object are processed by the recognition model to obtain the target features of the palm print image and the target features of the palm vein image of the preset database object. The target features of the palm print image of the preset database object are registered as the preset palm print features of the preset database object and stored in the preset database, and the target features of the palm vein image of the preset database object are registered as the preset palm vein features of the preset database object and stored in the preset database.
[0107] It is understandable that the similarity calculation method for determining the similarity between the target features of a palm print image and the preset palm print features of several preset database objects, and the similarity calculation method for determining the similarity between the target features of a vein image and the preset palm vein features of several preset database objects, can be by calculating Euclidean distance, cosine similarity, Manhattan distance, etc. The same similarity calculation method can be used for different modalities, or different similarity calculation methods can be used. There is no specific limitation here.
[0108] In some application scenarios, the steps described above for determining the recognition result of the object to be identified based on each first similarity and each second similarity may include the following: selecting a preset similarity from each first similarity and each second similarity, and using the preset database object corresponding to the preset similarity as the object recognition result of the object to be identified. The preset similarity may be the maximum similarity among each first similarity and each second similarity.
[0109] In other application scenarios, the steps of determining the recognition result of the object to be identified based on each first similarity and each second similarity may include the following: selecting a preset similarity from each first similarity as a first candidate similarity, and using the palm print image of the object to be identified corresponding to the first candidate similarity as a candidate palm print image; selecting a preset similarity from each second similarity as a second candidate similarity; and using the palm vein image of the object to be identified corresponding to the second candidate similarity as a candidate palm vein image; in response to the quality score in the quality detection result of the candidate palm print image being greater than the quality score in the quality detection result of the candidate palm vein image, using the preset database object corresponding to the first candidate similarity as the object recognition result of the object to be identified; in response to the quality score in the quality detection result of the candidate palm print image being less than the quality score in the quality detection result of the candidate palm vein image, using the preset database object corresponding to the second candidate similarity as the object recognition result of the object to be identified.
[0110] In some application scenarios, the step of obtaining the first target similarity between the object to be identified and the preset database object based on the first similarity and / or second similarity between the object to be identified and the preset database object for each preset database object can be performed as follows for each preset database object: using the first similarity between the object to be identified and the preset database object as the first target similarity between the object to be identified and the preset database object; or using the second similarity between the object to be identified and the preset database object as the first target similarity between the object to be identified and the preset database object; or performing weighted fusion of the first similarity and the second similarity between the object to be identified and the preset database object to obtain the first target similarity between the object to be identified and the preset database object. During the weighted fusion process, the weights of the first similarity and the second similarity can be dynamically set according to the user's attention to palmprint images and palm vein images or other indicators.
[0111] In other application scenarios, the step of weightedly fusing the first and second similarities between the object to be identified and objects in the preset database to obtain the first target similarity between the object to be identified and objects in the preset database may include the following steps: determining the palmprint fusion weight based on the quality detection results of the acquired palmprint image; determining the palm vein fusion weight based on the quality detection results of the acquired palm vein image; and for each object in the preset database, weightedly fusing the first similarity, palmprint fusion weight, second similarity, and palm vein fusion weight of the preset database object to obtain the first target similarity between the object to be identified and objects in the preset database.
[0112] In some application scenarios, palmprint fusion weights are determined based on the quality score in the quality inspection results of palmprint images. Specifically, the quality score in the palmprint image quality inspection results is normalized and used as the palmprint fusion weight; or, the quality score in the palmprint image quality inspection results is directly used as the palmprint fusion weight. Similarly, in some application scenarios, palm vein fusion weights are determined based on the quality score in the quality inspection results of palm vein images. Specifically, the quality score in the palm vein image quality inspection results is normalized and used as the palm vein fusion weight; or, the quality score in the palm vein image quality inspection results is directly used as the palm vein fusion weight.
[0113] It is understandable that similarity represents the degree of similarity between the object to be identified and the preset database objects corresponding to the similarity. The higher the similarity, the higher the degree of similarity between the object to be identified and the preset database objects corresponding to the similarity.
[0114] In some application scenarios, similarity includes a first target similarity in palmprint modality between the object to be identified and a preset database object corresponding to a similarity score. For each preset database object, the first similarity score of the preset database object is adjusted according to the palmprint fusion weight to obtain a first target similarity score. Specifically, this includes multiplying the palmprint fusion weight and the first similarity score of the preset database object as the first target similarity score. In some application scenarios, similarity includes a first target similarity in palm vein modality between the object to be identified and a preset database object corresponding to a similarity score. For each preset database object, a second similarity score of the preset database object is adjusted according to the palm vein fusion weight to obtain a first target similarity score. Specifically, this includes multiplying the palm vein fusion weight and the second similarity score of the preset database object as the first target similarity score. In some application scenarios, the above steps of determining the object recognition result based on the first target similarity score between the object to be identified and each preset database object include: selecting a preset similarity score from each first target similarity score, and using the preset database object corresponding to the preset similarity score as the object recognition result of the object to be identified. The preset similarity can be the maximum similarity among the similarities of each first target.
[0115] In other application scenarios, similarity includes characterizing the similarity between the object to be identified and a preset database object corresponding to the similarity score in palmprint and palm vein modalities. For each preset database object, the weighted fusion result of the first similarity score of the preset database object, the palmprint fusion weight, the second similarity score of the preset database object, and the palm vein fusion weight is used as the first target similarity between the preset database object and the object to be identified. Specifically, this includes: multiplying the palmprint fusion weight and the first similarity score of the preset database object as the first product; multiplying the palm vein fusion weight and the second similarity score of the preset database object as the second product; and summing the first and second products as the first target similarity between the preset database object and the object to be identified. In some application scenarios, the above steps of determining the object recognition result based on the first target similarity between the object to be identified and each preset database object include: selecting a preset similarity score from each first target similarity score, and using the preset database object corresponding to the preset similarity score as the object recognition result of the object to be identified. The preset similarity score can be the maximum similarity score among the first target similarities.
[0116] In some embodiments, the third feature extraction module includes a preset feature extraction module and a fusion module. Step S32 above may include the following steps: The preset feature extraction module performs feature extraction processing on the advanced features of the image to be identified to obtain target features of the image to be identified. The target features of the image to be identified include target features of the palmprint image and target features of the palm vein image. The fusion module performs fusion processing on the target features of the palmprint image and the target features of the palm vein image to obtain fused features of the object to be identified. The similarity between the target features of the palmprint image and the preset palmprint features of several preset database objects is used as the first similarity between the object to be identified and each preset database object. And / or, the similarity between the target features of the palm vein image and the preset palm vein features of several preset database objects is used as the second similarity between the object to be identified and each preset database object. The similarity between the fused features of the image to be identified and the preset fused features of several preset database objects is used as the third similarity between the image to be identified and each preset database object. Based on each third similarity and each first similarity and / or each second similarity, the object recognition result of the object to be identified is determined.
[0117] Please refer to the above content for the preset feature extraction module; it will not be repeated here. The fusion module is used to fuse the target features of the palm print image and the target features of the palm vein image of the object to be identified, thereby obtaining the fused features of the object to be identified.
[0118] The identity information, preset palmprint features, preset palm vein features, and preset fusion features of several preset base objects are stored in a preset database. As needed, the identity information, preset palmprint features, preset palm vein features, etc., of several preset base objects can be retrieved from the preset database. It can be understood that during user registration, the captured images of each preset base object are processed by a recognition model to obtain the target features of the palmprint image and the target features of the palm vein image of the preset base object. These target features are then fused to obtain the fused feature of the preset base object, which is then registered as the preset fused feature of the preset base object and stored in the preset database.
[0119] In some application scenarios, the similarity between the target features of the palmprint image and the preset palmprint features of several preset database objects is used as the first similarity between the object to be identified and each preset database object. The similarity between the fused features of the image to be identified and the preset fused features of several preset database objects is used as the third similarity between the image to be identified and each preset database object. Based on each third similarity and each first similarity, the object recognition result of the object to be identified is determined.
[0120] In some application scenarios, the similarity between the target features of the palm vein image and the preset palm vein features of several preset database objects is used as the second similarity between the object to be identified and each preset database object. The similarity between the fused features of the image to be identified and the preset fused features of several preset database objects is used as the third similarity between the image to be identified and each preset database object. Based on each third similarity and each second similarity, the object recognition result of the object to be identified is determined.
[0121] In some application scenarios, the similarity between the target features of a palmprint image and the preset palmprint features of several preset database objects is used as the first similarity between the object to be identified and each preset database object. The similarity between the target features of a palm vein image and the preset palm vein features of several preset database objects is used as the second similarity between the object to be identified and each preset database object. The similarity between the fused features of the image to be identified and the preset fused features of several preset database objects is used as the third similarity between the image to be identified and each preset database object. The first target similarity is updated based on each third similarity to obtain the corresponding second target similarity, and the object recognition result of the object to be identified is determined based on each second target similarity.
[0122] In some application scenarios, the step of determining the object recognition result of the object to be identified based on each third similarity, each first similarity, and / or each second similarity may include: taking the largest similarity among each third similarity and each first similarity as the selected similarity; or, taking the largest similarity among each third similarity and each second similarity as the selected similarity; or, taking the largest similarity among each third similarity, each first similarity, and each second similarity as the selected similarity; and taking the preset database object corresponding to the selected similarity as the object recognition result of the object to be identified.
[0123] In other application scenarios, the step of determining the object recognition result of the object to be identified based on each third similarity, each first similarity, and / or each second similarity may include: selecting a preset similarity from each third similarity as a third candidate similarity; selecting a preset similarity from each first similarity as a first candidate similarity, and using the palmprint image of the object to be identified corresponding to the first candidate similarity as a candidate palmprint image; and / or, selecting a preset similarity from each second similarity as a second candidate similarity, and using the palm vein image of the object to be identified corresponding to the second candidate similarity as a candidate palm vein image. The preset similarity is the maximum similarity. In response to the third candidate similarity being the largest among the first, second, and third candidate similarities, the preset database object corresponding to the third candidate similarity is used as the object recognition result of the object to be identified. In response to the third candidate similarity not being the largest among the first, second, and third candidate similarities, the object recognition result of the object to be identified is determined based on the preset database object corresponding to the first and second candidate similarities. The steps described above for determining the object recognition result of the object to be recognized based on the preset database objects corresponding to the first candidate similarity and the preset database objects corresponding to the second candidate similarity include: in response to the quality score in the quality detection result of the candidate palmprint image being greater than the quality score in the quality detection result of the candidate palm vein image, using the preset database object corresponding to the first candidate similarity as the object recognition result of the object to be recognized; and in response to the quality score in the quality detection result of the candidate palmprint image being less than the quality score in the quality detection result of the candidate palm vein image, using the preset database object corresponding to the second candidate similarity as the object recognition result of the object to be recognized.
[0124] In other application scenarios, the step of determining the object recognition result of the object to be identified based on each third similarity, each first similarity, and / or each second similarity may include: for each preset database object, performing a weighted fusion of the first similarity, palmprint fusion weight, second similarity, palm vein fusion weight, and third similarity of the preset database object to obtain the target similarity between the object to be identified and the preset database object. The object recognition result of the object to be identified is then determined based on each target similarity. Specifically, for each preset database object, the following steps are performed: adjusting the third similarity of the preset database object according to a preset fusion weight, or a combination of palmprint fusion weight and palm vein weight, to obtain the second target similarity corresponding to the third similarity. This specifically includes: using the product of the preset fusion weight, or the combination of palmprint fusion weight and palm vein weight, and the third similarity of the preset database object as the second target similarity corresponding to the third similarity. In some application scenarios, the above steps for determining the object recognition result of the object to be identified based on the similarity of each target include: taking the maximum similarity among each second target similarity and each first similarity as the selected target similarity; or, taking the maximum similarity among each second target similarity and each second similarity as the selected target similarity; or, taking the maximum similarity among each second target similarity and each first similarity and each second similarity as the selected target similarity; and taking the preset base database object corresponding to the selected target similarity as the object recognition result of the object to be identified. The calculation methods for the first and second similarities are described above and will not be repeated here.
[0125] In some application scenarios, the step described above, for each preset database object, updating the first target similarity between the object to be identified and the preset database object based on the third similarity between the object to be identified and the preset database object to obtain the second target similarity between the object to be identified and the preset database object, can be performed on each preset database object by performing the following steps: using the third similarity between the object to be identified and the preset database object as the second target similarity between the object to be identified and the preset database object; or, using the sum of the third similarity and the first target similarity between the object to be identified and the preset database object as the second target similarity between the object to be identified and the preset database object; or, performing a weighted fusion of the third similarity and the first target similarity between the object to be identified and the preset database object to obtain the second target similarity between the object to be identified and the preset database object. During the weighted fusion process, the weights of the first target similarity and the third similarity can be dynamically set based on the user's attention to the first target similarity and the third similarity or other indicators. The steps described above for determining the object recognition result based on the second target similarity between the object to be identified and each preset base database object may include the following: selecting a preset similarity from the second target similarity as the selected similarity; and using the preset base database object corresponding to the selected similarity as the object recognition result of the object to be identified.
[0126] It is understandable that the similarity calculation methods for determining the similarity between the target features of the palmprint image and the preset palmprint features of several preset database objects, the similarity between the target features of the vein image and the preset palm vein features of several preset database objects, and the similarity between the fusion features of the object to be identified and the preset fusion features of several preset database objects can be achieved by calculating Euclidean distance, cosine similarity, Manhattan distance, etc. The same similarity calculation method can be used for the three methods, or different similarity calculation methods can be used. There is no specific limitation here.
[0127] To facilitate understanding of this application, this application takes the following example: obtaining each third similarity, each first similarity, and each second similarity; obtaining a second target similarity based on each third similarity, each first similarity, and each second similarity; and determining the recognition result of the object to be identified based on each second target similarity. The implementation of determining the recognition result of the object to be identified based solely on each third similarity and each first similarity, and determining the recognition result of the object to be identified based solely on each third similarity and each second similarity, can be found in the relevant content on determining the recognition result of the object to be identified based on each third similarity, each first similarity, and / or each second similarity, which will not be elaborated here.
[0128] For example, this application can achieve the fusion of missing / low-quality modality perception and palmprint / palm vein features by using the quality detection results of palmprint images and palm vein images, thereby improving the accuracy of object recognition results. Figure 5 As shown, the recognition model includes a feature completion module, a feature extraction module, and a fusion module for generating missing / low-quality modalities. In the process of generating advanced features for missing / low-quality modalities, to avoid the previous problem of needing to design a generation model for each missing modality, resulting in a large overall model and high deployment difficulty, this method combines the quality scores of the images to be recognized corresponding to each modality (palmprint modality and palm vein modality) to construct generation prompts for low-quality / missing modalities, thus recovering the missing information in a simple way. Specifically, the process of generating features for missing / low-quality modalities, i.e., advanced features, is described in the following formulas (1) and (2): Formula (1); Formula (2); The preset prompt information is set as follows: ,in These represent the generation prompts for palm print mode and palm vein mode, respectively (i.e., the preset palm print generation prompt information and the preset palm vein generation prompt information, respectively). ,in These represent the dimension and length of the prompt message, respectively. This represents the specific text content that generates the prompt message. Taking palmprint modality loss (i.e., when the image to be recognized is a palmprint image, and the quality score of the palmprint image is less than or equal to the second score threshold) as an example, if the input is... . This represents a convolutional block consisting of a 1D convolutional layer and an activation function. This indicates the generation of one mode from another. To provide learnable prompts, target prompt information corresponding to the quality score in the quality detection result is added as guidance information when constructing the guidance prompts. It can be understood that when the quality score in the quality detection result is lower than a preset threshold and the modality exists, the modality generation feature is generated by the combined action of low-quality data from this modality and data from another modality; while when the modality is missing, the modality generation feature is generated solely by data from another modality. It can be understood that formula (1) represents the process of completing the initial features of the palmprint image when the image to be identified is a palmprint image, provided that the quality score in the quality detection result of the palmprint image of the palmprint modality is greater than the second score threshold and less than or equal to the first score threshold. The initial features representing the palm print image; The initial features representing the candidate image, i.e., the initial features of the palm vein image; The preset information prompts corresponding to the palm print modality are the preset palm print generation prompts. The generated palmprint features are the advanced features corresponding to the aforementioned palmprint image. It can be understood that formula (2) represents the process of completing the initial features of the palm vein image when the image to be identified is a palm vein image, provided that the quality score in the quality detection result of the palm vein image in the palm vein modality is greater than the second score threshold and less than or equal to the first score threshold. Initial features representing the palmar vein image; The initial features representing the candidate image, i.e., the initial features of the palm print image; The preset information prompts corresponding to the palm vein modality are the preset palm vein generation prompts. The generated palm vein features are the advanced features corresponding to the aforementioned palm vein image.
[0129] like Figure 5 The diagram shows a framework for a palmprint and palm vein feature fusion network based on cue-based learning to perceive modal quality. It is assumed that... Represents the initial features of the image to be identified, where, Represents the initial features of the palm print image. This represents the initial features of the palm vein image. (Using...) and These represent the missing or low-quality initial features of the palmprint and palm vein modalities, respectively. First, a quality score is obtained from the quality detection results of the image to be identified through a quality assessment network. For modalities with scores below a first threshold or those that are missing, a cue message for that modality is obtained. This cue message, along with the initial features of the low-quality modality and the initial features from another normal modality, are fed into a missing / low-quality modality generation module consisting of a 1D convolutional layer and an activation function to obtain the advanced features of the image to be identified. Second, the normal modality data features and the generated modality features (i.e., the advanced features of the palmprint image and the advanced features of the palm vein image) are extracted separately through a transformer feature extraction branch to obtain the target features of the palmprint image and the palm vein image. It is understood that the above-mentioned preset feature extraction module can include a transformer feature extraction algorithm. Then, the extracted target features of the palmprint image and the palm vein image are concatenated to obtain the fused palmprint and palm vein features, i.e., the fused features of the object to be identified, which can be represented as... .
[0130] For example, the object recognition method of this application relies on palmprint and palm vein recognition testing. For any input binocular video stream of the palmprint and palm vein recognition system, dual-path palm detection is performed, and then the quality scores of the palmprint image and the palm vein image are obtained through the single-modal quality assessment modality. The quality scores of the palmprint image and the palm vein image, along with the initial features corresponding to the two modalities, are input into the missing / low-quality modality generation module (i.e., the feature completion module). Based on the quality scores of the palmprint image and the palm vein image and the obtained prompt information matrix, the prompt information matrix is combined with the two inputs to generate compensation features (i.e., advanced features) for the low-quality / missing modality. Then, the feature extraction network in the preset extraction module and the feature fusion network in the fusion module are used to obtain the final palmprint and palm vein fusion features (i.e., the fusion features of the object to be recognized).
[0131] During registration in the database (i.e., the preset database storing several preset database objects), only images of the palm print and palm vein that both meet the threshold quality scores can be registered. The database features of each preset database object's identifier (ID) are composed of concatenated palm print and palm vein features, with the palm print and palm vein features having the same length. During database registration, the same ID is used to extract and register single palm print features (i.e., the preset palm print features of the preset database object), single palm vein features (i.e., the preset palm vein features of the preset database object), and palm print and palm vein fusion features (i.e., the preset fusion features of the preset database object). During recognition, the palm print and palm vein fusion features extracted from any input binocular video stream of the palm print and palm vein recognition system are also composed of concatenated palm print and palm vein features. During comparison, the first half of the fused features (i.e., the target features corresponding to the palmprint image) is compared with the palmprint features of each ID in the database to obtain a first similarity score; the second half of the fused features (i.e., the target features corresponding to the palm vein image) is compared with the palm vein features of each ID in the database to obtain a second similarity score; and the fused features of the object to be identified are compared with the palmprint and palm vein spliced features of each ID to obtain a third similarity score. For example, each first similarity score and each second similarity score are multiplied by the normalized quality score of the corresponding modality, and then a voting decision is made among the three scores (first similarity score and second similarity score). The ID with the highest score is selected as the identified identity ID and used as the object recognition result for the object to be identified.
[0132] In some embodiments, the recognition model includes a preset feature extraction module and a fusion module; the method further includes a training step for the recognition model, the training step including: acquiring sample images of sample objects and sample annotation information of the sample images, the sample annotation information including a first annotation recognition result determined based on the annotation target features of the sample images and / or a second annotation recognition result determined based on the annotation fusion features of the sample images. Inputting the sample images into the recognition model, obtaining the sample target features of the sample images output by the preset feature extraction module in the recognition model and the sample fusion features of the sample objects output by the fusion module in the recognition model. Classifying the sample target features of the sample images to obtain a first sample recognition result of the sample objects; and / or, classifying the sample fusion features of the sample images to obtain a second sample recognition result of the sample objects. Determining a first loss based on the difference between the first annotation recognition result and the first sample recognition result of the sample objects; and / or, determining a second loss based on the difference between the second annotation recognition result and the second sample recognition result of the sample objects. Adjusting the parameters in the recognition model based on the first loss and / or the second loss to obtain a trained recognition model.
[0133] The sample images include sample palm print images and sample palm vein images of the same sample object.
[0134] In some application scenarios, the above-mentioned classification processing of sample target features of sample images to obtain the first sample recognition result of sample objects includes: classifying sample target features of sample palmprint images to obtain the first sample recognition result of palmprint modality. Specifically, the classification processing may involve inputting the sample target features of sample palmprint images into a first preset classifier to obtain the first sample recognition result of palmprint modality output by the first preset classifier; classifying sample target features of sample palm vein images to obtain the first sample recognition result of palm vein modality. Specifically, the classification processing may involve inputting the sample target features of sample palm vein images into a second preset classifier to obtain the first sample recognition result of palm vein modality output by the second preset classifier.
[0135] In some application scenarios, the first annotation recognition result includes the first annotation recognition result corresponding to the sample palm print image and the first annotation recognition result corresponding to the sample palm vein image; the step of determining the first loss based on the difference between the first annotation recognition result and the first sample recognition result of the sample object includes: determining the first sub-loss based on the difference between the first annotation recognition result corresponding to the sample palm print image and the first sample recognition result of the palm print modality; determining the second sub-loss based on the difference between the first annotation recognition result corresponding to the sample palm vein image and the first sample recognition result of the palm vein modality; and determining the first loss based on the first sub-loss and / or the second sub-loss.
[0136] It is understandable that the steps of inputting a sample image into the recognition model to obtain the sample target features of the sample image output by the preset feature extraction module in the recognition model and the sample fusion features of the sample object output by the fusion module in the recognition model are the same as the process of inputting the image to be recognized into the recognition model to obtain the target features of the image to be recognized and the fusion features of the object to be recognized in the intermediate result of the recognition model, and will not be elaborated here.
[0137] In some application scenarios, the parameters in the recognition model are adjusted based on the first loss and the second loss to obtain the trained recognition model; the parameters in the recognition model are adjusted based on the second loss to obtain the trained recognition model; the parameters in the recognition model are adjusted based on the first loss and the second loss to obtain the trained recognition model.
[0138] For example, the parameters in the recognition model are adjusted based on the first loss and / or the second loss to obtain the trained recognition model, including the following: feeding the fused features into a classifier related to identity information, where the classification loss function is denoted using Arcface. For example Figure 5 The recognition model shown divides training into two parts: pre-training of the complete modality and fine-tuning of the fusion features generated from missing / low-quality modalities. First, pre-training related to identity information is performed using a one-to-one complete modality of palm print and palm vein. The training loss function uses Arcface, that is, the first preset extraction module is trained using the first sub-loss, and the second preset extraction module is trained using the second sub-loss. Then, the parameters of each modality transformer branch network (i.e., the first preset extraction module and the second preset extraction module) are frozen, and the parameters of the missing / low-quality modality generation module (feature completion module) and the final fusion module are fine-tuned using the second loss.
[0139] The palmprint and palm vein feature fusion recognition method proposed in this application, based on prompt learning to perceive modal quality, fuses palmprint and palm vein features to design a non-contact biometric method that combines the high anti-counterfeiting capabilities of palm vein recognition with the high universality of palmprint recognition, thereby achieving a more stable and secure identity recognition effect. This application addresses the practical problem of low-quality or missing modalities in palmprint and palm vein feature fusion recognition due to sensor defects, lighting interference, or dirt occlusion under specific environmental conditions. It introduces the theory of prompt learning in large model domains and combines it with single-modal quality assessment scores to construct learnable modal quality perception prompts. This provides a simpler way to recover information from missing / low-quality modalities, alleviating the problem of decreased fusion feature discrimination ability caused by the absence or low quality of a single modality. Simultaneously, a base database registration strategy incorporating modal quality perception prompts is designed to improve the recognition accuracy and system stability of the palmprint and palm vein recognition system under low-quality or missing modalities.
[0140] The above scheme acquires the image of the object to be identified, which includes the palm print image and the palm vein image of the object to be identified. The palm print image and the palm vein image are respectively subjected to quality detection processing to obtain the quality detection result of the image to be identified, and the target image and the target prompt information of the image to be identified are determined. The target image and the target prompt information of the image to be identified are input into the recognition model to obtain the object recognition result output by the recognition model. In this way, the target prompt information corresponding to the quality detection result of the image to be identified can guide the recognition model to process the image information of the target image, thereby improving the accuracy of the recognition result output by the recognition model.
[0141] Please see Figure 6 , Figure 6 This is a schematic diagram of an embodiment of the object recognition device of this application. The object recognition device 60 includes an acquisition module 61, a quality detection module 62, a determination module 63, and a recognition module 64. The acquisition module 61 is used to acquire an image of the object to be recognized, which includes a palm print image and a palm vein image of the object to be recognized. The quality detection module 62 is used to perform quality detection processing on the palm print image and the palm vein image in the image to be recognized, respectively, to obtain a quality detection result of the image to be recognized. The determination module 63 is used to determine a target image and target prompt information of the image to be recognized based on the quality detection result of the image to be recognized. The recognition module 64 is used to input the target image and the target prompt information of the image to be recognized into a recognition model to obtain the object recognition result output by the recognition model.
[0142] Please refer to the object identification method for the functions performed by each module; they will not be repeated here.
[0143] The above scheme acquires the image of the object to be identified, which includes the palm print image and the palm vein image of the object to be identified. The palm print image and the palm vein image are respectively subjected to quality detection processing to obtain the quality detection result of the image to be identified, and the target image and the target prompt information of the image to be identified are determined. The target image and the target prompt information of the image to be identified are input into the recognition model to obtain the object recognition result output by the recognition model. In this way, the target prompt information corresponding to the quality detection result of the image to be identified can guide the recognition model to process the image information of the target image, thereby improving the accuracy of the recognition result output by the recognition model.
[0144] Please see Figure 7 , Figure 7 This is a schematic diagram of the structure of an embodiment of the electronic device of this application. The electronic device 70 includes a memory 71 and a processor 72. The processor 72 is used to execute program instructions stored in the memory 71 to implement the steps in the above-described object recognition method embodiment. In a specific implementation scenario, the electronic device 70 may include, but is not limited to, a microcomputer or a server. In addition, the electronic device 70 may also include mobile devices such as laptops and tablets, which are not limited here.
[0145] Specifically, processor 72 controls itself and memory 71 to implement the steps in the above-described object recognition method embodiments. Processor 72 can also be referred to as a CPU (Central Processing Unit). Processor 72 may be an integrated circuit chip with signal processing capabilities. Processor 72 can also be a general-purpose processor, digital signal processor (DSP), application-specific integrated circuit (ASIC), field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. A general-purpose processor can be a microprocessor or any conventional processor. Furthermore, processor 72 can be implemented using integrated circuit chips.
[0146] The above scheme acquires the image of the object to be identified, which includes the palm print image and the palm vein image of the object to be identified. The palm print image and the palm vein image are respectively subjected to quality detection processing to obtain the quality detection result of the image to be identified, and the target image and the target prompt information of the image to be identified are determined. The target image and the target prompt information of the image to be identified are input into the recognition model to obtain the object recognition result output by the recognition model. In this way, the target prompt information corresponding to the quality detection result of the image to be identified can guide the recognition model to process the image information of the target image, thereby improving the accuracy of the recognition result output by the recognition model.
[0147] Please see Figure 8 , Figure 8 This is a schematic diagram of a computer-readable storage medium according to an embodiment of the present application. The computer-readable storage medium 80 stores program instructions 801 thereon, which, when executed by a processor, implement the steps in any of the above-described object recognition method embodiments.
[0148] The above scheme acquires the image of the object to be identified, which includes the palm print image and the palm vein image of the object to be identified. The palm print image and the palm vein image are respectively subjected to quality detection processing to obtain the quality detection result of the image to be identified, and the target image and the target prompt information of the image to be identified are determined. The target image and the target prompt information of the image to be identified are input into the recognition model to obtain the object recognition result output by the recognition model. In this way, the target prompt information corresponding to the quality detection result of the image to be identified can guide the recognition model to process the image information of the target image, thereby improving the accuracy of the recognition result output by the recognition model.
[0149] In some embodiments, the functions or modules of the apparatus provided in this disclosure can be used to perform the methods described in the above method embodiments. The specific implementation can be referred to the description of the above method embodiments, and for the sake of brevity, it will not be repeated here.
[0150] The description of the various embodiments above tends to emphasize the differences between the various embodiments. The similarities or similarities between them can be referred to, and for the sake of brevity, they will not be repeated here.
[0151] In the several embodiments provided in this application, it should be understood that the disclosed methods and apparatus can be implemented in other ways. For example, the apparatus implementations described above are merely illustrative. For instance, the division of modules or units is only a logical functional division, and in actual implementation, there may be other division methods. For example, units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the mutual coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection of devices or units may be electrical, mechanical, or other forms.
[0152] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0153] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) or processor to execute all or part of the steps of the methods of various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0154] If the technical solution of this application involves personal information, the product using this technical solution has clearly informed the user of the personal information processing rules and obtained the user's voluntary consent before processing the personal information. If the technical solution of this application involves sensitive personal information, the product using this technical solution has obtained the user's separate consent before processing the sensitive personal information, and also meets the requirement of "express consent". For example, at personal information collection devices such as cameras, clear and prominent signs are set up to inform users that they have entered the scope of personal information collection and that personal information will be collected. If an individual voluntarily enters the collection scope, it is deemed that they have agreed to the collection of their personal information; or on the personal information processing device, with clear signs / information informing users of the personal information processing rules, authorization is obtained from the individual through pop-up information or by asking the individual to upload their personal information; wherein, the personal information processing rules may include information such as the personal information processor, the purpose of personal information processing, the processing method, and the types of personal information processed.
Claims
1. An object recognition method, characterized in that, The method includes: Acquire an image of the object to be identified, the image of the object to be identified including a palm print image and a palm vein image of the object to be identified; The palm print image and palm vein image in the image to be identified are subjected to quality detection processing respectively to obtain the quality detection result of the image to be identified; The target image and the target prompt information of the image to be identified are determined based on the quality detection results of the image to be identified. The target image and the target prompt information of the image to be identified are input into the recognition model to obtain the object recognition result output by the recognition model.
2. The method according to claim 1, characterized in that, The recognition model includes a first feature extraction module and a second feature extraction module. The step of inputting the target image and the target prompt information of the image to be recognized into the recognition model to obtain the object recognition result output by the recognition model includes: The target image is processed by the first feature extraction module to obtain the initial features of the target image; The initial features of the target image and the target prompt information are processed by the second feature extraction module to obtain the object recognition result output by the recognition model.
3. The method according to claim 2, characterized in that, The second feature extraction module includes a feature completion module and a third feature extraction module. The step of processing the initial features of the target image and the target prompt information through the second feature extraction module to obtain the object recognition result output by the recognition model includes: The feature completion module processes the initial features and target prompt information of the target image to obtain advanced features; The advanced features are processed by the third feature extraction module to obtain the object recognition result output by the recognition model.
4. The method according to claim 3, characterized in that, The third feature extraction module includes a preset feature extraction module and an object determination module. The step of processing the advanced features through the third feature extraction module to obtain the object recognition result output by the recognition model includes: The advanced features are extracted using the preset feature extraction module to obtain the target features. The object identification module performs feature matching processing on the target features and the preset base database features to obtain the object recognition result.
5. The method according to claim 4, characterized in that, The target features include the target features of the palm print image and / or the palm vein image, and the preset database features include preset palm print features and / or preset palm vein features of a number of preset database objects; The step of performing feature matching processing between the target features and the preset base database features through the object determination module to obtain the object recognition result includes: The similarity between the target features of the palmprint image and the preset palmprint features of each preset database object is respectively used as the first similarity between the object to be identified and each preset database object; and / or The similarity between the target features of the palm vein image and the preset palm vein features of each preset database object is used as the second similarity between the object to be identified and each preset database object. For each of the preset base objects, a first target similarity between the object to be identified and the preset base objects is obtained based on the first similarity and / or the second similarity between the object to be identified and the preset base objects; The object recognition result is determined based on the first target similarity between the object to be identified and each preset base database object.
6. The method according to claim 5, characterized in that, The preset base database features also include preset fusion features between preset palm print features and preset palm vein features of each preset base database object; The step of determining the object recognition result based on the first target similarity between the object to be identified and each preset base database object includes: The target features of the palm print image and the target features of the palm vein image are fused to obtain the fused features of the object to be identified. The similarity between the fusion features of the image to be identified and the preset fusion features corresponding to each preset base object is used as the third similarity between the image to be identified and each preset base object. For each of the preset base objects, the first target similarity between the object to be identified and the preset base objects is updated based on the third similarity between the object to be identified and the preset base objects to obtain the second target similarity between the object to be identified and the preset base objects. The object recognition result is determined based on the second target similarity between the object to be identified and each preset base database object.
7. The method according to claim 5, characterized in that, The step of obtaining the first target similarity between the object to be identified and the preset database object for each of the preset database objects, based on the first similarity and / or second similarity between the object to be identified and the preset database object, includes: Based on the quality detection results of the acquired palmprint image, the palmprint fusion weight of the palmprint image is determined. Based on the quality detection results of the acquired palm vein images, the palm vein fusion weights of the palm vein images are determined. For each of the preset database objects, the first similarity, the second similarity, the palm print fusion weight, and the palm vein fusion weight between the object to be identified and the preset database objects are weighted and fused to obtain the first target similarity between the object to be identified and the preset database objects.
8. The method according to claim 1, characterized in that, The image to be identified is a palm print image; The step of determining the target image and the target prompt information of the target image based on the quality detection result of the image to be identified includes: If the quality detection result of the palmprint image indicates that the quality score of the palmprint image is greater than a first score threshold, then the target image is determined to be a palmprint image, and the target prompt information is empty; In response to the quality detection result of the palmprint image indicating that the quality score of the palmprint image is less than or equal to the first score threshold and greater than the second score threshold, the target image is determined to be a palmprint image and a palm vein image, and the preset palmprint generation prompt information is used as the target prompt information; In response to the quality detection result of the palmprint image indicating that the quality score of the palmprint image is less than or equal to the second score threshold, the target image is determined to be a palm vein image, and the preset palmprint generation prompt information is used as the target prompt information.
9. The method according to claim 1, characterized in that, The image to be identified is a palm vein image; The step of determining the target image and the target prompt information of the target image based on the quality detection result of the image to be identified includes: If the quality detection result of the palm vein image indicates that the quality score of the palm print image is greater than a first score threshold, then the target image is determined to be a palm vein image, and the target prompt information is empty; In response to the quality detection result of the palm vein image indicating that the quality score of the palm print image is less than or equal to the first score threshold and greater than the second score threshold, the target image is determined to be a palm print image and a palm vein image, and the preset palm vein generation prompt information is used as the target prompt information; In response to the quality detection result of the palm vein image indicating that the quality score of the palm print image is less than or equal to the second score threshold, the target image is determined to be a palm print image, and the preset palm vein generation prompt information is used as the target prompt information.
10. The method according to claim 1, characterized in that, The step of obtaining the image of the object to be identified includes: Acquire a first image of the object to be identified by a palm print acquisition device and a second image of the object to be identified by a palm vein acquisition device; The region of interest is extracted from the first acquired image to obtain the first region of interest of the first acquired image; The region of interest (ROI) of the second acquired image is extracted to obtain the second ROI of the second acquired image; The first acquired image and the second acquired image are aligned according to the first region of interest and the second region of interest, respectively, to obtain the palm print image and the palm vein image.
11. An electronic device, characterized in that, include: A memory and a processor, wherein the memory stores program instructions, and the processor retrieves the program instructions from the memory to perform the method as claimed in any one of claims 1-10.
12. A computer-readable storage medium having program instructions stored thereon, characterized in that, When the program instructions are executed by the processor, they are used to implement the method as described in any one of claims 1-10.