Identity verification method and device based on multi-modal feature fusion, equipment and medium

CN122597930APending Publication Date: 2026-08-18CHONGQING TECH & BUSINESS UNIV +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610734881.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-05-26
Publication Date
2026-08-18

AI Technical Summary

Technical Problem

[0005]本发明的目的是提供一种基于多模态特征融合的身份核验方法、装置、设备及介质,用以解决现有技术中因模态单一所存在的身份核验可靠性差的问题

Benefits of technology

(1)本发明融合现场人脸图像、身份证件图像、身份文本信息及人脸数据库等多维特征,来形成了多重验证闭环,即先基于人脸特征,来进行现场人脸与证件图像上人脸的比对识别,而后,在对比通过后,再利用身份文本信息在数据库中搜索出匹配人脸,以及对现场人脸进行人脸识别,得到身份识别信息;最后,基于前述匹配人脸、现场人脸、身份文本信息和身份识别信息,来进行身份的交叉验证,从而在交叉验证后,得到核验结果;通过上述设计,不法分子即使利用高精度Deepfake或面具成功欺骗现场人脸与证件人脸的单一比对,本发明仍会依据身份文本信息在数据库中搜索出匹配人脸,并与现场人脸、身份识别信息进行交叉验证,任一维度异常即会导致核验失败,如此,本发明可有效抵御人脸伪造、证件篡改及冒用等复合型攻击,从而提高了身份核验的可靠性。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122597930A_ABST
    Figure CN122597930A_ABST
Patent Text Reader

Abstract

This invention discloses an identity verification method, device, equipment, and medium based on multimodal feature fusion. By fusing multidimensional features such as on-site facial images, ID card images, document text information, and facial databases, this invention forms a closed loop capable of multiple verifications. Therefore, even if criminals successfully deceive a single comparison between the on-site face and the ID card face using high-precision deepfakes or masks, this invention will still search for matching faces in the database based on the identity text information and cross-verify them with the on-site face and identity recognition information. Any anomaly in any dimension will lead to verification failure. Thus, this invention can effectively resist complex attacks such as face forgery, document tampering, and impersonation, thereby significantly improving the security and reliability of verification.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of identity recognition technology, specifically relating to an identity verification method, device, equipment, and medium based on multimodal feature fusion. Background Technology

[0002] Identity verification is a crucial step in many fields, including finance, government affairs, security, travel, and internet services. Its core objective is to confirm the consistency between the "person" and the "claimed identity." With increasing digitalization, the demand for remote and contactless identity authentication is growing, placing higher demands on the accuracy, convenience, and security of the verification process.

[0003] Currently, the mainstream identity verification methods are mainly based on facial recognition comparison methods. This involves capturing facial images of people on-site through a camera and then comparing them with facial images stored in a database or uploaded by users. This method utilizes deep learning technology and has achieved high accuracy in controlled environments.

[0004] However, the aforementioned existing technologies have the following shortcomings: they are single-modal and vulnerable to attack. That is, there are obvious security vulnerabilities in single-modal face recognition. For example, high-precision forged face images or videos can successfully deceive pure visual comparison systems. Although there are existing technologies that use OCR to recognize text information on ID documents, text recognition and face comparison are usually processed as independent modules in sequence. They lack deep fusion and cross-verification mechanisms of multi-source information and are difficult to resist complex attacks such as document forgery and face replacement. Therefore, based on the aforementioned shortcomings, how to overcome the limitations of single-modal information in existing identity verification methods has become an urgent problem to be solved. Summary of the Invention

[0005] The purpose of this invention is to provide an identity verification method, apparatus, device, and medium based on multimodal feature fusion, in order to solve the problem of poor reliability of identity verification due to single modality in the prior art.

[0006] To achieve the above objectives, the present invention adopts the following technical solution: Firstly, an identity verification method based on multimodal feature fusion is provided, including: Obtain facial images and ID card images of the target individuals; The image of the identity document is processed by text recognition to obtain the identity text information of the target person; The face image is extracted from the ID card image, and features are extracted from the face capture image and the face image to obtain the actual captured face features and the ID card face features, respectively. Based on the actual facial features captured and the facial features on the ID card, determine whether the facial image captured in the photo and the facial image on the ID card belong to the same face; If so, then in the face database, a matching face corresponding to the identity text information is searched, and the face image is subjected to face recognition to obtain the identity information of the target person; Based on the matched face, the captured face image, the identity recognition information, and the identity text information, the target person's identity is cross-verified to obtain the cross-verification result; Based on the cross-validation results, the identity verification results of the target person are generated.

[0007] Based on the aforementioned disclosures, this invention integrates multi-dimensional features such as on-site facial images, ID card images, identity text information, and a facial database to form a multi-layered verification closed loop. First, based on facial features, it compares and identifies the face on the on-site image with the face on the ID card image. Then, after a successful comparison, it uses the identity text information to search the database for a matching face and performs facial recognition on the on-site face to obtain identity verification information. Finally, based on the aforementioned matching face, on-site face, identity text information, and identity verification information, it performs cross-verification of the identity, thus obtaining the verification result. Through this design, even if criminals successfully deceive the single comparison of the on-site face and the ID card face using high-precision Deepfakes or masks, this invention will still search the database for a matching face based on the identity text information and perform cross-verification with the on-site face and identity verification information. Any anomaly in any dimension will lead to verification failure. Therefore, this invention can effectively resist complex attacks such as face forgery, document tampering, and impersonation, thereby improving the reliability of identity verification.

[0008] In one possible design, feature extraction is performed on the captured facial image to obtain the actual captured facial features, including: The face image is subjected to a two-dimensional discrete cosine transform to obtain a DCT coefficient matrix, and low-frequency coefficient matrix blocks are extracted from the DCT coefficient matrix to form face frequency domain features. Texture features are extracted from the captured face image to obtain a face texture histogram; Gradient intensity feature extraction is performed on the face texture histogram to obtain texture gradient features; Based on the texture gradient features and the face texture histogram, facial spatial features are generated. The frequency domain features and spatial domain features of the face are fused to generate the actual captured face features.

[0009] In one possible design, texture features are extracted from the captured facial image to obtain a facial texture histogram, including: Get texture neighborhood parameters; For any pixel in the face image, based on the texture neighborhood parameters, determine the corresponding texture neighborhood pixels. Based on the gray values ​​of each texture neighboring pixel, the texture feature value of any pixel is calculated, and after traversing all pixels in the face image, the texture feature value of each pixel is obtained. The texture feature values ​​of each pixel are converted into binary numbers, and binary numbers that satisfy the condition that the number of circular transitions is less than or equal to a preset value are selected from the binary numbers corresponding to each pixel as regular texture features. The number of circular transitions is the number of transitions from 0 to 1 and from 1 to 0 in the binary number. The occurrence frequency of each regular texture feature is counted, and based on the occurrence frequency of each regular texture feature, a histogram of the corresponding pixel points for each regular texture feature is obtained; The face texture histogram is constructed by using the histograms of the pixels corresponding to each regular texture feature.

[0010] In one possible design, the texture neighborhood parameters include: the neighborhood major axis, the neighborhood minor axis, and the total number of neighborhood pixels; Specifically, based on texture neighborhood parameters, each texture neighborhood pixel corresponding to any given pixel is determined, including: The polar angle of each texture neighboring pixel is calculated based on the total number of neighboring pixels. Based on the polar angle of each texture neighboring pixel, the major axis of the neighborhood, and the minor axis of the neighborhood, the radial radius between each texture neighboring pixel and any pixel is calculated. The coordinates of each texture neighboring pixel are calculated using the polar angle and radial radius of each texture neighboring pixel. Based on the coordinates of each texture neighboring pixel, the corresponding texture neighboring pixels are determined for any given pixel.

[0011] In one possible design, gradient intensity feature extraction is performed on the face texture histogram to obtain texture gradient features, including: For any regular texture feature in the face texture histogram, the positions of the pixels corresponding to the regular texture feature in the face image are statistically determined. For the pixel at position k, obtain the texture neighborhood pixels of the pixel at position k; Based on the coordinates and grayscale value of the pixel at the k-th position, and the coordinates and grayscale values ​​of each texture neighboring pixel, calculate the gradient magnitude of each texture neighboring pixel relative to the pixel at the k-th position. Take the largest gradient value among all gradient magnitudes as the maximum gradient value of the pixel at the k-th position; Increment k by 1 and reacquire the texture neighborhood pixels of the pixel at the k-th position until k equals K. This yields the maximum gradient value of each pixel corresponding to any regular texture feature, where the initial value of k is 1 and K is the total number of occurrences of pixels corresponding to any regular texture feature. The mean of each maximum gradient value is used as the gradient feature of all pixels corresponding to any regular texture feature. After all regular texture features in the face texture histogram have been circumvented, several gradient intensity features are obtained. The texture gradient features are generated using several gradient intensity features.

[0012] In one possible design, before performing feature extraction on the captured facial image, the method further includes: Obtain a set of sample images of faces, and calculate the standard illumination range of faces based on the set of sample images of faces; The average brightness of the captured face image is calculated, and based on the standard illumination range of the face and the average brightness, it is determined whether the captured face image is a low-light image. If so, then extract the high-frequency detail layer from the captured face image; The illumination component estimation process is performed on the face image to obtain the illumination component in the face image; Based on the facial image and the illumination components, the reflection components in the facial image are determined; The high-frequency detail layer and the reflection component are used to generate an initially enhanced face image; Adaptive color compensation processing is performed on the initial enhanced face image to obtain an enhanced face image, so that feature extraction can be performed on the enhanced face image to obtain the actual captured face features.

[0013] In one possible design, the face image is subjected to illumination component estimation processing to obtain the illumination components in the face image, including: Obtain several window parameters at different scales, where any window parameter at any scale includes the window radius, window orientation, and the position of the pixel to be processed on the window boundary; For any pixel in the face image, multiple filtering windows for that pixel at the p-th scale are constructed based on several window parameters at the p-th scale. By using the Gaussian kernel function and multiple filtering windows at the p-th scale, any pixel is filtered to obtain the filtered output value output by the multiple filtering windows at the p-th scale. Calculate the difference between the filtered output value of each filter window at the p-th scale and the gray value of any pixel. The filtered output value with the smallest difference is taken as the actual output value at the p-th scale, and the illumination component at the p-th scale is obtained after all pixels in the face image have been traversed. Increment p by 1, and reconstruct multiple filtering windows for any pixel at the p-th scale based on several window parameters at the p-th scale, until p equals P, to obtain the illumination components at each scale, where P is the total number of scales; The illumination components of the captured face image are composed using the illumination components at each scale.

[0014] Secondly, an identity verification device based on multimodal feature fusion is provided, including: The acquisition unit is used to acquire facial images and ID card images of the target personnel; The recognition unit is used to perform text recognition processing on the image of the ID card to obtain the identity text information of the target person; The feature extraction unit is used to extract the face image from the ID card image and to extract features from the face capture image and the face image to obtain the actual captured face features and the ID card face features, respectively. A face comparison unit is used to determine, based on the actual captured face features and the ID card face features, whether the face image captured in the face image and the face image in the ID card image belong to the same face; The cross-verification unit is used to search the face database for a matching face corresponding to the identity text information when the face comparison unit determines that the face image and the face image in the ID card image belong to the same face, and to perform face recognition on the face image to obtain the identity information of the target person. The cross-verification unit is used to perform cross-verification of the target person's identity based on the matched face, the captured face image, the identity recognition information, and the identity text information, and obtain the cross-verification result; The cross-verification unit is also used to generate the identity verification result of the target person based on the cross-verification result.

[0015] Thirdly, another identity verification device based on multimodal feature fusion is provided. Taking the device as an electronic device as an example, it includes a memory, a processor, and a transceiver that are connected in sequence. The memory is used to store a computer program, the transceiver is used to send and receive messages, and the processor is used to read the computer program and execute the identity verification method based on multimodal feature fusion as described in the first aspect or any possible design in the first aspect.

[0016] Fourthly, a storage medium is provided, on which instructions are stored, which, when executed on a computer, perform the identity verification method based on multimodal feature fusion as described in the first aspect or any possible design of the first aspect.

[0017] Fifthly, a computer program product containing instructions is provided, which, when executed on a computer, causes the computer to perform the identity verification method based on multimodal feature fusion as described in the first aspect or any possible design of the first aspect.

[0018] Beneficial effects: (1) This invention integrates multi-dimensional features such as on-site facial images, ID card images, identity text information, and facial databases to form a multi-verification closed loop. First, based on facial features, the on-site face is compared and identified with the face on the ID card image. Then, after the comparison is successful, the identity text information is used to search for matching faces in the database, and the on-site face is also recognized to obtain identity recognition information. Finally, based on the aforementioned matching faces, on-site faces, identity text information, and identity recognition information, cross-verification of identity is performed, and the verification result is obtained after cross-verification. Through the above design, even if criminals successfully deceive the single comparison of on-site face and ID card face using high-precision Deepfake or masks, this invention will still search for matching faces in the database based on identity text information and cross-verify them with on-site faces and identity recognition information. Any abnormality in any dimension will lead to verification failure. In this way, this invention can effectively resist complex attacks such as face forgery, ID card tampering, and impersonation, thereby improving the reliability of identity verification. Attached Figure Description

[0019] Figure 1 This is a flowchart illustrating the steps of the identity verification method based on multimodal feature fusion provided in an embodiment of the present invention. Figure 2 This is a structural diagram of the identity verification device based on multimodal feature fusion provided in an embodiment of the present invention; Figure 3 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present invention. Detailed Implementation

[0020] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the present invention will be briefly introduced below in conjunction with the accompanying drawings and descriptions of the embodiments or the prior art. Obviously, the following description of the structure of the accompanying drawings is only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort. It should be noted that the description of these embodiments is for the purpose of helping to understand the present invention, but does not constitute a limitation of the present invention.

[0021] It should be understood that although the terms first, second, etc., may be used herein to describe various units, these units should not be limited by these terms. These terms are only used to distinguish one unit from another. For example, a first unit may be referred to as a second unit, and similarly, a second unit may be referred to as a first unit, without departing from the scope of the exemplary embodiments of the invention.

[0022] It should be understood that the term "and / or" that may appear in this document is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can mean: A exists alone, B exists alone, and A and B exist simultaneously. The term " / and" that may appear in this document describes another relationship between related objects, indicating that two relationships can exist. For example, A / and B can mean: A exists alone, and A and B exist alone. In addition, the character " / " that may appear in this document generally indicates that the related objects before and after it are in an "or" relationship.

[0023] Example: See Figure 1 As shown, the identity verification method based on multimodal feature fusion provided in this embodiment can be executed by a computer device with certain computing resources, but is not limited to. For example, the computer device can be, but is not limited to, a server, an edge computer, or a personal computer (PC, which refers to a multi-purpose computer of a size, price, and performance suitable for personal use; desktop computers, laptops, mini-laptops, tablets, and ultrabooks are all personal computers), a smartphone, or a personal digital assistant (PDA). It is understood that the aforementioned executing entity does not constitute a limitation on the embodiments of this application. Accordingly, the operation steps of this method can be, but are not limited to, the steps S1 to S7 below.

[0024] S1. Acquire the facial image and ID card image of the target person; in this embodiment, for example, but not limited to, capturing the real-time facial image of the target person through a camera, thereby using it as the facial image; at the same time, acquire the ID card image (such as ID card image, passport image, etc.) through a high-speed scanner or camera; thus, after completing the image acquisition, the identity text information can be extracted, the process of which is shown in step S2 below.

[0025] S2. Perform text recognition processing on the ID card image to obtain the target person's identity text information; in specific implementation, for example, perform OCR (Optical Character Recognition) recognition on the ID card image to obtain the target person's identity text information, wherein the identity text information may include, but is not limited to: name, age, gender, registered address and ID number, etc.; of course, OCR recognition is a commonly used text recognition technology, and its principle will not be elaborated here.

[0026] After extracting the identity text information, the first level of verification can be carried out, which is to compare the face on site with the face on the document. The process is as shown in step S3 below.

[0027] S3. Extract the face image from the ID card image, and perform feature extraction on the face capture image and the face image to obtain the actual captured face features and the ID card face features, respectively. In this embodiment, for example, but not limited to, a face detection algorithm can be used to extract the face image from the ID card image; such as using a YOLOv5 neural network for face segmentation. Of course, face detection algorithms are commonly used techniques in face recognition, and their principles will not be elaborated here.

[0028] Furthermore, when identifying specific faces, lighting conditions can be quite complex, such as images captured in low-light environments like at night, dusk, or indoors, or images captured in high-light conditions. Images captured in low-light conditions often suffer from insufficient brightness and blurred details, while images captured in high-light conditions suffer from excessive illumination. Therefore, before feature extraction, it is necessary to perform quality enhancement processing on the captured face images and the face images themselves.

[0029] The following uses a face image as an example to illustrate the quality enhancement process, which may be, but is not limited to, the steps S031 to S037 below.

[0030] S031. Obtain a set of face capture sample images and calculate the standard illumination range of the face based on the set of face capture sample images. In this embodiment, a set of face capture sample images can be formed by selecting several face capture samples (usually, the samples in the face image library are face images captured under normal lighting conditions) from an existing face image library. In this way, after obtaining the set of face capture sample images, the standard illumination range of the face can be calculated based on it.

[0031] Optionally, but not limited to, the mean brightness and standard deviation of brightness of all face image samples in the face image sample set can be calculated first; then, the aforementioned standard brightness range of the face can be calculated based on the mean brightness and standard deviation of brightness.

[0032] For example, the following formula can be used, but is not limited to, to calculate the standard illumination range for a human face.

[0033] ; In the formula, Indicates the standard illumination range for a human face. These represent the mean and standard deviation of luminance, respectively. This represents the adjustable coefficient (which can take a value of 1.4). The brightness threshold is 0.05.

[0034] Thus, based on the aforementioned formula, after calculating the standard illumination range of the face, it is possible to determine whether the face image needs quality enhancement, as shown in step S032 below.

[0035] S032. Calculate the average brightness of the captured face image, and based on the standard illumination range of the face and the average brightness, determine whether the captured face image is a low-light image.

[0036] In practice, if the average brightness of a face image is less than the left endpoint of the standard illumination range for a face, it is determined to be a low-light image, and quality enhancement is required. At the same time, if the average brightness is greater than the right endpoint of the standard illumination range for a face, it is determined to be a high-light image, and quality enhancement is also required. If the average brightness is within the standard illumination range for a face, no processing is required, and feature extraction can be performed directly.

[0037] The traditional Retinex algorithm is prone to causing edge blurring when processing low-light images and has a weak ability to preserve image details. Therefore, this embodiment provides an improved low-light image quality enhancement algorithm. It first extracts the high-frequency detail layer to prevent the loss of details during subsequent quality enhancement processing. Then, it extracts the reflection component. Next, it fuses the reflection component with the high-frequency detail layer to complete the initial image enhancement. Finally, it performs adaptive color compensation to obtain the enhanced face image (color compensation can avoid color distortion during the enhancement process).

[0038] Optionally, the aforementioned quality enhancement process may be, but is not limited to, the steps S033 to S037 below.

[0039] S033. If so, the high-frequency detail layer in the face image is extracted. In a specific application, this embodiment first performs dual gamma correction on the face image to obtain a corrected image; then, it performs Gaussian filtering on the corrected image to obtain a filtered image (the filtered image is a low-frequency image); finally, it calculates the difference image between the corrected image and the filtered image, and uses the difference image as the high-frequency detail layer.

[0040] After extracting the high-frequency detail layer, illumination component estimation can be performed, as shown in step S034 below. S034. Perform illumination component estimation processing on the face image to obtain the illumination components in the face image. In this embodiment, the traditional technique uses a Gaussian kernel function for illumination component estimation, which can lead to edge blurring. This is because the center of the local window of the Gaussian filter function coincides with the pixel to be processed. When the pixel to be processed is an edge region of the image, the filtering algorithm will cause the edge to blur. Therefore, this embodiment provides a new illumination component estimation method that does not use the pixel to be processed as the center of the local window, thereby avoiding the edge blurring problem.

[0041] Optionally, the process for estimating the illumination components is as shown in steps S034a to S034g below.

[0042] S034a. Obtain several window parameters at different scales, wherein any window parameter at any scale includes the window radius, window direction, and the position of the pixel to be processed on the window boundary; in specific implementation, the window radius is half the side length of a square window, the window direction is the angle between the window and the horizontal line, and the position of the pixel to be processed on the window boundary is represented by u (its value is [0, d], where d is an integer); further, when u is 0, it means that the pixel to be processed is at the top corner of the window, and when u is d, it means that the pixel to be processed is on the side of the window; at the same time, the window direction can be q×(π / 2), where q takes the value 0, 1, 2, 3; thus, there are eight window directions.

[0043] Specifically, when u is d, the window direction is 0, π / 2, π, 3π / 2. Similarly, when u is 0, there are also four window directions. Based on this, the eight windows are windows in the up, down, left, right, northeast, southeast, southwest, and northwest directions. For example, when u is d and the window direction is 0, it corresponds to the left window (i.e., the pixel to be processed is located at the right center of the window). When u is d and the window direction is π / 2, it corresponds to the top window, and so on. Of course, the other windows will not be described in detail.

[0044] Furthermore, the window radius is the same at the same scale, but different at different scales.

[0045] Thus, after obtaining several window parameters at different scales, multiple filter windows at different scales can be constructed based on these parameters, as shown in step S034b below.

[0046] S034b. For any pixel in the face image, multiple filtering windows are constructed at the p-th scale based on several window parameters at the p-th scale. In this embodiment, the multiple filtering windows at the p-th scale have the same radius but different directions. Thus, after constructing the multiple filtering windows at the p-th scale, the multiple filtering windows can be used to perform filtering processing on any pixel. The process is shown in step S034c below.

[0047] S034c. Using a Gaussian kernel function and multiple filtering windows at the p-th scale, filter the pixel to obtain the filtered output value output by the multiple filtering windows at the p-th scale.

[0048] In practical implementation, for any filtering window, its filtered output value is: ; In the formula, This represents the filtered output value corresponding to any of the filtering windows. Represents any of the filtering windows, This represents the pixel points within any of the filtering windows. Represents the Gaussian kernel function. This represents the image corresponding to any of the filtering windows. Represents the convolution symbol.

[0049] Thus, based on the aforementioned formula, after calculating the filtered output values ​​of multiple filter windows at the p-th scale, the actual output value of any pixel at the p-th scale can be calculated, as shown in steps S034d and S034e below.

[0050] S034d. Calculate the difference between the filtered output value of each filter window at the p-th scale and the gray value of any pixel.

[0051] S034e. The filtered output value with the smallest difference is taken as the actual output value at the p-th scale, and the illumination component at the p-th scale is obtained after all pixels in the face image have been traversed.

[0052] In practical applications, the smallest difference between the output value corresponding to each filtering window and the gray value of any pixel is taken as the actual filtered output of that pixel. In this way, after traversing all pixels in the face image as described above, the actual output value of each pixel at the p-th scale can be obtained. Then, the actual output value of each pixel at the p-th scale can be used to generate the illumination component at the p-th scale.

[0053] Thus, based on the aforementioned steps, after estimating the illumination component at the p-th scale, the remaining scales can be processed in the same way, as shown in step S034f below.

[0054] S034f. Increment p by 1, and reconstruct multiple filtering windows for any pixel at the p-th scale based on several window parameters at the p-th scale, until p equals P, to obtain the illumination components at each scale; in this embodiment, P can be set to 3 (i.e., the total number of scales is set to 3), that is, the illumination components at three scales are extracted; then, the illumination components at the three scales can be used to form the illumination components of the face image, as shown in step S034g below.

[0055] S034g. The illumination components of the captured face image are composed using the illumination components at each scale.

[0056] Based on the aforementioned steps S034a to S034g, after extracting the illumination components at different scales, the reflection components are determined by directly combining them with the face image. The process is shown in step S035 below.

[0057] S035. Based on the face image and the illumination component, determine the reflection component in the face image; in specific implementations, for example, but not limited to, the following formula can be used to determine the aforementioned reflection component.

[0058] ; In the formula, This represents the reflection component. This refers to the captured image of the face. This represents the illumination component at the p-th scale. Represents the convolution symbol, This represents the weight of the p-th scale. This represents the total number of scales.

[0059] As can be seen from the aforementioned formula, this embodiment no longer uses the logarithmic function in the traditional Retinex algorithm. The reason is that while the logarithmic function has the advantage of compressing the dynamic range of an image and effectively balancing the overall brightness distribution, thus enhancing the detail in dark areas, it is extremely sensitive to low grayscale values ​​close to 0, easily amplifying weak pixel values, leading to a significant increase in noise in dark areas. Simultaneously, the monotonic response of the logarithmic function in bright areas can easily cause excessive brightness enhancement, producing unnatural artifacts and losing details. Therefore, this embodiment uses a non-linear activation function (its form is...). Instead of the logarithmic function, the nonlinear activation function is used to achieve a linear response near 0, which avoids excessive amplification of noise, ensures bounded output (approaching ±1), and suppresses over-enhancement. This helps to prevent noise amplification in dark areas and reduce artifacts and brightness overflow.

[0060] Thus, based on the aforementioned formula, after calculating the reflection component, the aforementioned high-frequency detail layer can be combined to generate the initially enhanced face image, as shown in step S036 below.

[0061] S036. Using the high-frequency detail layer and the reflection component, an initially enhanced face image is generated; in this embodiment, the initially enhanced face image can be obtained by using the high-frequency detail layer plus the reflection component.

[0062] After the initial enhancement of the face image is completed, a second enhancement, namely color compensation, can be performed, as shown in step S037 below.

[0063] S037. Adaptive color compensation processing is performed on the initially enhanced face image to obtain an enhanced face image, so that feature extraction can be performed on the enhanced face image to obtain the actual captured face features. In specific applications, for example, but not limited to, calculating a color threshold based on the initially enhanced face image; then, selecting pixels above the color threshold from the initially enhanced face image as highlights; next, calculating a first compensation parameter and a second compensation parameter for each color channel based on each highlight; finally, using the first compensation parameter and the second compensation parameter for each color channel, the pixel value of each pixel in the initially enhanced face image in each color channel can be compensated to obtain the enhanced face image after compensation processing.

[0064] The color threshold is calculated as follows: the brightness of each pixel in the initially enhanced face image is calculated (i.e., the sum of the R channel value, G channel value and B channel value of each pixel is used as the brightness), then the average brightness of the image is obtained; finally, the average brightness of the image is multiplied by a scaling factor (the value of which can be set according to actual use, and is not specifically limited in this embodiment) to obtain the color threshold.

[0065] After calculating the color threshold, pixels with color values ​​greater than the color threshold can be designated as highlights. The color value of any pixel is essentially the sum of its R channel value, G channel value, and B channel value. Simultaneously, for any color channel, this embodiment first calculates the mean and standard deviation of the pixel values ​​of each highlight in that color channel. Then, the mean and standard deviation are used as the first compensation parameter and the second compensation parameter for that color channel.

[0066] Based on this, after calculating the first compensation parameter and the second compensation parameter, color compensation can be performed. For example, but not limited to, the R channel value of any pixel can be compensated according to the following formula.

[0067] ; In the formula, This represents the pixel value of the R channel after compensation for any given pixel. This represents the pixel value of the R channel of any given pixel. This represents the maximum pixel value in the initially enhanced face image (i.e., the maximum sum of the R, G, and B channels of each pixel in the image). These represent the first and second compensation parameters of channel R, respectively; of course, the compensation formulas for the other channels are the same, and will not be repeated in this embodiment.

[0068] Thus, based on the aforementioned formula, after completing the adaptive color compensation process, an enhanced face image can be obtained; of course, the enhancement process for face images in ID card images is also the same, and will not be elaborated here.

[0069] Meanwhile, when the face image is a high-light image, it is converted into a grayscale image, and the grayscale values ​​in the grayscale image are inverted (i.e., using 1-grayscale image) to obtain the inverted grayscale image; then, the image is enhanced using the same method as the aforementioned steps S031 to S037; finally, the enhanced image is subjected to another grayscale value inversion operation to obtain the enhanced face image.

[0070] Based on the aforementioned steps S031 to S037, after completing the enhancement processing of the captured face image and the face image, feature extraction can be performed; in this embodiment, the captured face image is used as an example to illustrate the feature extraction process.

[0071] In practical applications, traditional technologies typically use only a single algorithm for feature extraction, which is insufficient to capture multifaceted information about a face. Therefore, in order to better obtain facial features, this embodiment provides a face feature extraction algorithm that extracts and fuses multiple features. The process can be, but is not limited to, the steps S31 to S35 below.

[0072] S31. Perform a two-dimensional discrete cosine transform on the face image to obtain a DCT (discrete cosine transform) coefficient matrix, and extract low-frequency coefficient matrix blocks from the DCT coefficient matrix to form face frequency domain features using the low-frequency coefficient matrix blocks; in this embodiment, the 10×10 matrix block in the upper left corner of the DCT coefficient matrix (whose starting point is the upper left corner endpoint of the DCT coefficient matrix) can be used as the low-frequency coefficient matrix block, but is not limited to.

[0073] Of course, the two-dimensional discrete cosine transform is a commonly used technique for facial feature extraction, and its principle will not be elaborated here.

[0074] Since the discrete cosine transform can obtain global face image information but cannot obtain the texture features of the face, after extracting the face frequency domain features, it is also necessary to extract texture features to complete the features. The process can be, but is not limited to, the steps shown in step S32 below.

[0075] S32. Extract texture features from the captured face image to obtain a face texture histogram; in specific implementations, for example, but not limited to, the following steps S32a to S32f can be used to generate the face texture histogram.

[0076] S32a. Obtain texture neighborhood parameters; in this embodiment, the texture neighborhood parameters may include, but are not limited to: neighborhood major axis, neighborhood minor axis and total number of neighborhood pixels.

[0077] Thus, after obtaining the texture neighborhood parameters, the texture neighborhood pixels of each pixel in the face image (here referring to the enhanced image) can be determined based on these parameters, as shown in step S32b below.

[0078] S32b. For any pixel in the captured face image, based on the texture neighborhood parameters, determine the corresponding texture neighborhood pixels. In specific implementation, for example, but not limited to, first calculate the polar angle of each texture neighborhood pixel based on the total number of neighborhood pixels; then, calculate the radial radius between each texture neighborhood pixel and the any pixel based on the polar angle, the major axis, and the minor axis of the neighborhood pixels; finally, use the polar angle and radial radius of each texture neighborhood pixel to calculate the coordinates of each texture neighborhood pixel; thus, the corresponding texture neighborhood pixels can be determined based on the coordinates of each texture neighborhood pixel.

[0079] Optionally, one of the formulas for calculating the polar angle is disclosed below: ; In the formula, represents the polar angle of the i-th texture neighboring pixel, and n represents the total number of neighboring pixels.

[0080] Thus, based on the aforementioned formula, after calculating the polar angle of the i-th texture neighboring pixel, the radial radius between the i-th texture neighboring pixel and any other pixel can be calculated by combining the major and minor axes of the neighborhood. The calculation formula is as follows: ; In the formula, This represents the radial radius between the i-th texture neighbor pixel and any other pixel. These represent the major axis and the minor axis of the neighborhood, respectively.

[0081] After calculating the radial radius of the i-th pixel based on the aforementioned formula, its corresponding x-coordinate can be determined by combining it with its corresponding polar angle. The calculation formula is as follows: ; In the formula, These represent the x and y coordinates of the i-th texture neighbor pixel, respectively.

[0082] Therefore, based on the aforementioned formula, the coordinates of each texture neighboring pixel corresponding to any pixel can be calculated, and the texture neighboring pixel of any pixel can be determined based on the coordinates; then, the texture feature value of any pixel can be calculated based on the gray value of each texture neighboring pixel, as shown in step S32c below.

[0083] S32c. Based on the grayscale values ​​of each texture neighboring pixel, calculate the texture feature value of any pixel, and obtain the texture feature value of each pixel after traversing all pixels in the face image.

[0084] Optionally, for example, but not limited to, the texture feature value of any pixel can be calculated according to the following formula.

[0085] ; In the formula, This represents the texture feature value of any given pixel. This represents the grayscale value of any given pixel. This represents the grayscale value of the i-th texture neighbor pixel. Let be the texture feature mapping function, where, when When greater than or equal to 0, The value is 1, when When less than 0, The value is 0, and n is the total number of neighboring pixels.

[0086] Therefore, the texture feature value of each pixel in the face image can be calculated using the aforementioned formula. Then, feature statistics can be performed on the texture feature values ​​to generate the corresponding face histogram. The process is shown in steps S32d to S32f below.

[0087] S32d. Convert the texture feature value of each pixel into a binary number, and select the binary number that satisfies the condition that the number of circular transitions is less than or equal to a preset value from the binary numbers corresponding to each pixel, so as to use it as a regular texture feature. The number of circular transitions is the number of transitions from 0 to 1 and from 1 to 0 in the binary number. In this embodiment, the preset value is 2.

[0088] The following example illustrates the selection process for regular texture features: Assume pixel 1 corresponds to the binary number 00000000, pixel 2 to the binary number 00001111, pixel 3 to the binary number 00111100, and pixel 4 to the binary number 01010101. Pixel 1's binary number has 0 transitions between 0 and 1, and between 1 and 0, which conforms to the aforementioned selection rule, so it is classified as a regular texture feature (Category 1). Similarly, pixel 2's binary number has 1 transition between 0 and 1, and between 1 and 0, also conforming to the aforementioned selection rule, so it is classified as the second regular texture feature (Category 2). Pixel 3's binary number has 2 transitions between 0 and 1, and between 1 and 0, so it is classified as the third regular texture feature (Category 3). Finally, pixel 4's binary number has 7 transitions between 0 and 1, and between 1 and 0, which is greater than 2, therefore it is not selected.

[0089] Thus, based on the aforementioned method, regular texture features can be selected; then, statistics on regular texture features can be performed, as shown in step S32e below.

[0090] S32e. Count the occurrence frequency of each regular texture feature, and obtain the histogram of the corresponding pixel points for each regular texture feature based on the occurrence frequency of each regular texture feature. In this embodiment, based on the previous example, for category 1, i.e., 00000000, all pixels with a texture feature value of 00000000 are grouped into the same category. For example, assuming there are 4 pixels with a texture feature value of 00000000, then their occurrence frequency is 4. At this time, the horizontal axis of their corresponding histogram is the texture feature value, and the vertical axis is the occurrence frequency. Thus, after obtaining the histogram of the corresponding pixel points for each regular texture feature based on the above method, a face texture histogram can be generated based on it, as shown in step S32f below.

[0091] S32f. The face texture histogram is formed by using the histograms of the pixels corresponding to each regular texture feature; in this embodiment, the face texture histogram is obtained by splicing the various histograms.

[0092] Therefore, after extracting the facial texture features through the aforementioned steps S32a to S32f, in order to compensate for the shortcomings of traditional histograms that ignore gradient magnitude, this embodiment also sets up gradient feature extraction to enhance the robustness of face recognition to changes in illumination and contrast; wherein, the gradient feature extraction process is as shown in step S33 below.

[0093] S33. Perform gradient intensity feature extraction processing on the face texture histogram to obtain texture gradient features; in this embodiment, for example, but not limited to, the following steps S33a to S33g can be used to extract texture gradient features.

[0094] S33a. For any regular texture feature in the face texture histogram, count the positions of the pixels corresponding to the regular texture feature in the face image. In this embodiment, as previously explained, a regular texture feature corresponds to at least one pixel. For example, if the binary number corresponding to any regular texture feature is 00000000, then the pixel with the texture feature value of 00000000 is taken as the pixel corresponding to the regular texture feature. Assuming there are 4, then it is necessary to obtain the positions of these 4 pixels in the face image.

[0095] After calculating the positions of the pixels corresponding to any regular texture feature in the face image, the gradient magnitude of the texture neighborhood pixels at each position relative to the corresponding pixel at each position can be calculated so that the maximum gradient value of the pixels at each position can be calculated based on the gradient magnitude. The gradient magnitude calculation process is as shown in steps S33b and S33c below.

[0096] S33b. For the pixel at the k-th position, obtain the texture neighbor pixels of the pixel at the k-th position; In this embodiment, when calculating the texture feature value, the texture neighbor pixels of each pixel in the face image are already known, so this step can be used directly.

[0097] After obtaining the texture neighborhood pixels of the pixel at the k-th position, the gradient magnitude of the corresponding texture neighborhood region relative to the pixel at the k-th position can be calculated. The calculation process is shown in step S33c below.

[0098] S33c. Based on the coordinates and grayscale value of the pixel at the k-th position, and the coordinates and grayscale values ​​of each texture neighboring pixel, calculate the gradient magnitude of each texture neighboring pixel relative to the pixel at the k-th position. In specific implementation, for any texture neighboring pixel, for example, but not limited to, first calculate the square of the Euclidean distance between the any texture neighboring pixel and the pixel at the k-th position based on the coordinates of the any texture neighboring pixel and the coordinates of the pixel at the k-th position; then, calculate the difference between the grayscale value of the any texture neighboring pixel and the grayscale value of the pixel at the k-th position; finally, the ratio between the absolute value of the difference and the square of the Euclidean distance can be used as the gradient magnitude of the any texture neighboring pixel relative to the pixel at the k-th position.

[0099] Thus, after calculating the gradient magnitude of each texture neighborhood pixel relative to the pixel at the k-th position, the maximum gradient value of the pixel at the k-th position can be calculated based on this, as shown in step S33d below.

[0100] S33d. Take the largest gradient value among all gradient magnitudes as the maximum gradient value of the pixel at the k-th position.

[0101] After calculating the maximum gradient value of the pixel at the k-th position, the maximum gradient values ​​of the pixels at the remaining positions can be calculated in the same way, as shown in step S33e below.

[0102] S33e. Increment k by 1 and reacquire the texture neighborhood pixels of the pixel at the k-th position until k equals K, to obtain the maximum gradient value of each pixel corresponding to any regular texture feature, where the initial value of k is 1 and K is the total number of occurrences of the pixels corresponding to any regular texture feature.

[0103] After calculating the maximum gradient value of each pixel corresponding to any regular texture feature, the gradient feature corresponding to any regular texture feature can be calculated based on each maximum gradient value, as shown in step S33f below.

[0104] S33f. The mean of each maximum gradient value is used as the gradient feature of all pixels corresponding to any regular texture feature, and after all regular texture features in the face texture histogram have been circumvented, several gradient intensity features are obtained.

[0105] After generating the gradient features corresponding to all regular texture features, they are concatenated to generate texture gradient features, as shown in step S33g below.

[0106] S33g. The texture gradient features are generated using several gradient intensity features.

[0107] Therefore, after extracting the texture gradient features of the face image through the aforementioned steps S33a to S33g, the face texture histogram can be combined to generate the face spatial features, as shown in step S34 below.

[0108] S34. Based on the texture gradient features and the face texture histogram, generate face spatial features; in this embodiment, the texture gradient features and the face texture histogram (which can also be represented as a vector) are concatenated to obtain face spatial features; then, face frequency domain features can be combined to generate actual captured face features, as shown in step S35 below.

[0109] S35. Perform feature fusion on the face frequency domain features and the face spatial domain features to generate the actual captured face features; in specific implementation, for example, but not limited to, first normalize the face frequency domain features and face spatial domain features to obtain normalized frequency domain features and normalized spatial domain features; then, concatenate the normalized frequency domain features and normalized spatial domain features to obtain fused features; finally, perform PCA dimensionality reduction processing on the fused features to obtain the actual captured face features.

[0110] Therefore, after extracting the features of the face image from the face capture image and the ID card image through the aforementioned steps S31 to S35, the face captured on-site can be compared with the face on the ID card, as shown in step S4 below.

[0111] S4. Based on the actual captured facial features and the facial features on the ID card, determine whether the facial images in the captured image and the ID card image belong to the same face. In specific applications, for example, but not limited to, calculating the cosine similarity between the actual captured facial features and the ID card facial features, if the cosine similarity is greater than or equal to the similarity threshold, it can be determined that the facial images in the captured image and the ID card image belong to the same face, and in this case, the subsequent cross-validation process can be initiated; otherwise, the verification result of failure can be directly output.

[0112] The cross-validation process is shown in steps S5 and S6 below.

[0113] S5. If so, then in the face database, search for the matching face corresponding to the identity text information, and perform face recognition on the face image to obtain the identity information of the target person; in this embodiment, cross-validation uses the identity text information identified in the aforementioned step S2 to perform face matching in the known face database, and then performs face recognition on the face captured on site (i.e., face image) (existing face recognition models can be used to perform face recognition, such as R-CNN model, YOLO model, etc.) to obtain the identity information of the face on site; finally, cross-validation is performed based on the matched face and face image, as well as the identity text information on the ID card and the identity information obtained based on face recognition, as shown in step S6 below.

[0114] S6. Based on the matched face, the captured face image, the identity recognition information, and the identity text information, perform cross-verification on the target person to obtain a cross-verification result. In specific implementation, this involves determining whether the matched face and the captured face image belong to the same face, and whether the identity recognition information and the identity text information are the same, thereby obtaining a cross-verification result after the determination. Optionally, the process of comparing and recognizing the matched face and the captured face image can be referred to the aforementioned steps S3 and S4, and will not be repeated here. At the same time, the Jaro-Winkler similarity between the identity recognition information and the identity text information can be calculated to determine whether the two are the same.

[0115] S7. Generate the identity verification result of the target person based on the cross-validation result; in this embodiment, the identity verification result will be output only if the above two conditions are met (i.e., the matched person and the face image belong to the same face, and the identity recognition information and the identity text information are the same). Otherwise, if either condition is not met, the identity verification result will be output.

[0116] Therefore, through the multimodal feature fusion-based identity verification method described in detail in steps S1 to S7 above, this invention forms a multi-verification closed loop by integrating multi-dimensional features such as on-site facial images, ID card images, document text information, and facial databases. Thus, even if criminals successfully deceive the single comparison between the on-site face and the ID card face using high-precision deepfakes or masks, this invention will still search for matching faces in the database based on the identity text information and cross-verify them with the on-site face and identity recognition information. Any abnormality in any dimension will lead to verification failure. Thus, this invention can effectively resist complex attacks such as face forgery, document tampering, and impersonation, thereby significantly improving the security and reliability of verification.

[0117] like Figure 2 As shown, the second aspect of this embodiment provides a hardware device for implementing the identity verification method based on multimodal feature fusion described in the first aspect of the embodiment, comprising: The acquisition unit is used to acquire facial images and ID card images of the target person.

[0118] The recognition unit is used to perform text recognition processing on the image of the ID card to obtain the identity text information of the target person.

[0119] The feature extraction unit is used to extract the face image from the ID card image and to perform feature extraction on the face capture image and the face image to obtain the actual captured face features and the ID card face features, respectively.

[0120] The face comparison unit is used to determine whether the face image captured in the actual photograph and the face image captured in the ID card belong to the same face based on the actual photographed face features and the face features captured in the ID card image.

[0121] The cross-verification unit is used to search the face database for a matching face corresponding to the identity text information when the face comparison unit determines that the face image and the face image in the ID card image belong to the same face, and to perform face recognition on the face image to obtain the identity information of the target person.

[0122] The cross-verification unit is used to perform cross-verification of the target person's identity based on the matched face, the captured face image, the identity recognition information, and the identity text information, and obtain the cross-verification result.

[0123] The cross-verification unit is also used to generate the identity verification result of the target person based on the cross-verification result.

[0124] The working process, working details and technical effects of the device provided in this embodiment can be found in the first aspect of the embodiment, and will not be repeated here.

[0125] like Figure 3 As shown, the third aspect of this embodiment provides another identity verification device based on multimodal feature fusion. Taking the device as an electronic device as an example, it includes: a memory, a processor, and a transceiver that are connected in sequence. The memory is used to store a computer program, the transceiver is used to send and receive messages, and the processor is used to read the computer program and execute the identity verification method based on multimodal feature fusion as described in the first aspect of the embodiment.

[0126] For specific examples, the memory may include, but is not limited to, random access memory (RAM), read-only memory (ROM), flash memory, first-in-first-out (FIFO) memory, and / or first-in-last-out (FILO) memory, etc.; specifically, the processor may include one or more processing cores, such as a 4-core processor, an 8-core processor, etc. The processor may be implemented using at least one hardware form of DSP (Digital Signal Processing), FPGA (Field-Programmable Gate Array), PLA (Programmable Logic Array). The processor may also include a main processor and a coprocessor. The main processor, also known as the CPU (Central Processing Unit), is used to process data in the wake-up state; the coprocessor is a low-power processor used to process data in the standby state.

[0127] In some embodiments, the processor may integrate a GPU (Graphics Processing Unit), which is responsible for rendering and drawing the content to be displayed on the screen. For example, the processor may not be limited to microprocessors of the STM32F105 series, reduced instruction set computer (RISC) microprocessors, x86 architecture processors, or processors with integrated neural network processing units (NPUs). The transceiver may be, but is not limited to, a Wi-Fi transceiver, a Bluetooth transceiver, a General Packet Radio Service (GPRS) transceiver, a ZigBee (a low-power LAN protocol based on the IEEE 802.15.4 standard) transceiver, a 3G transceiver, a 4G transceiver, and / or a 5G transceiver. Furthermore, the device may also include, but is not limited to, a power module, a display screen, and other necessary components.

[0128] The working process, working details and technical effects of the electronic device provided in this embodiment can be found in the first aspect of the embodiment, and will not be repeated here.

[0129] The fourth aspect of this embodiment provides a storage medium that stores instructions containing the identity verification method based on multimodal feature fusion as described in the first aspect of the embodiment. That is, the storage medium stores instructions that, when the instructions are run on a computer, execute the identity verification method based on multimodal feature fusion as described in the first aspect of the embodiment.

[0130] The storage medium refers to a carrier for storing data, which may include, but is not limited to, floppy disks, optical disks, hard disks, flash memory, USB flash drives, and / or Memory Sticks. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices.

[0131] The working process, working details and technical effects of the storage medium provided in this embodiment can be found in the first aspect of the embodiment, and will not be repeated here.

[0132] The fifth aspect of this embodiment provides a computer program product containing instructions that, when executed on a computer, cause the computer to perform the identity verification method based on multimodal feature fusion as described in the first aspect of the embodiment, wherein the computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device.

[0133] Finally, it should be noted that the above description is merely a preferred embodiment of the present invention and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.

Claims

1. An identity verification method based on multimodal feature fusion, characterized in that, include: Obtain facial images and ID card images of the target individuals; The image of the identity document is processed by text recognition to obtain the identity text information of the target person; The face image is extracted from the ID card image, and features are extracted from the face capture image and the face image to obtain the actual captured face features and the ID card face features, respectively. Based on the actual facial features captured and the facial features on the ID card, determine whether the facial image captured in the photo and the facial image on the ID card belong to the same face; If so, then in the face database, a matching face corresponding to the identity text information is searched, and the face image is subjected to face recognition to obtain the identity information of the target person; Based on the matched face, the captured face image, the identity recognition information, and the identity text information, the target person's identity is cross-verified to obtain the cross-verification result; Based on the cross-validation results, the identity verification results of the target person are generated.

2. The method according to claim 1, characterized in that, Feature extraction is performed on the captured facial image to obtain the actual captured facial features, including: The face image is subjected to a two-dimensional discrete cosine transform to obtain a DCT coefficient matrix, and low-frequency coefficient matrix blocks are extracted from the DCT coefficient matrix to form face frequency domain features. Texture features are extracted from the captured face image to obtain a face texture histogram; Gradient intensity feature extraction is performed on the face texture histogram to obtain texture gradient features; Based on the texture gradient features and the face texture histogram, facial spatial features are generated. The frequency domain features and spatial domain features of the face are fused to generate the actual captured face features.

3. The method according to claim 2, characterized in that, Texture feature extraction is performed on the captured face image to obtain a face texture histogram, including: Get texture neighborhood parameters; For any pixel in the face image, based on the texture neighborhood parameters, determine the corresponding texture neighborhood pixels. Based on the gray values ​​of each texture neighboring pixel, the texture feature value of any pixel is calculated, and after traversing all pixels in the face image, the texture feature value of each pixel is obtained. The texture feature values ​​of each pixel are converted into binary numbers, and binary numbers that satisfy the condition that the number of circular transitions is less than or equal to a preset value are selected from the binary numbers corresponding to each pixel as regular texture features. The number of circular transitions is the number of transitions from 0 to 1 and from 1 to 0 in the binary number. The occurrence frequency of each regular texture feature is counted, and based on the occurrence frequency of each regular texture feature, a histogram of the corresponding pixel points for each regular texture feature is obtained; The face texture histogram is constructed by using the histograms of the pixels corresponding to each regular texture feature.

4. The method according to claim 3, characterized in that, The texture neighborhood parameters include: the neighborhood major axis, the neighborhood minor axis, and the total number of neighborhood pixels; Specifically, based on texture neighborhood parameters, each texture neighborhood pixel corresponding to any given pixel is determined, including: The polar angle of each texture neighboring pixel is calculated based on the total number of neighboring pixels. Based on the polar angle of each texture neighboring pixel, the major axis of the neighborhood, and the minor axis of the neighborhood, the radial radius between each texture neighboring pixel and any pixel is calculated. The coordinates of each texture neighboring pixel are calculated using the polar angle and radial radius of each texture neighboring pixel. Based on the coordinates of each texture neighboring pixel, the corresponding texture neighboring pixels are determined for any given pixel.

5. The method according to claim 3, characterized in that, Gradient intensity feature extraction is performed on the facial texture histogram to obtain texture gradient features, including: For any regular texture feature in the face texture histogram, the positions of the pixels corresponding to the regular texture feature in the face image are statistically determined. For the pixel at position k, obtain the texture neighborhood pixels of the pixel at position k; Based on the coordinates and grayscale value of the pixel at the k-th position, and the coordinates and grayscale values ​​of each texture neighboring pixel, calculate the gradient magnitude of each texture neighboring pixel relative to the pixel at the k-th position. Take the largest gradient value among all gradient magnitudes as the maximum gradient value of the pixel at the k-th position; Increment k by 1 and reacquire the texture neighborhood pixels of the pixel at the k-th position until k equals K. This yields the maximum gradient value of each pixel corresponding to any regular texture feature, where the initial value of k is 1 and K is the total number of occurrences of pixels corresponding to any regular texture feature. The mean of each maximum gradient value is used as the gradient feature of all pixels corresponding to any regular texture feature. After all regular texture features in the face texture histogram have been circumvented, several gradient intensity features are obtained. The texture gradient features are generated using several gradient intensity features.

6. The method according to claim 1, characterized in that, Before performing feature extraction on the captured facial image, the method further includes: Obtain a set of sample images of faces, and calculate the standard illumination range of faces based on the set of sample images of faces; The average brightness of the captured face image is calculated, and based on the standard illumination range of the face and the average brightness, it is determined whether the captured face image is a low-light image. If so, then extract the high-frequency detail layer from the captured face image; The illumination component estimation process is performed on the face image to obtain the illumination component in the face image; Based on the facial image and the illumination components, the reflection components in the facial image are determined; The high-frequency detail layer and the reflection component are used to generate an initially enhanced face image; Adaptive color compensation processing is performed on the initial enhanced face image to obtain an enhanced face image, so that feature extraction can be performed on the enhanced face image to obtain the actual captured face features.

7. The method according to claim 6, characterized in that, The illumination component estimation process is performed on the captured face image to obtain the illumination components in the captured face image, including: Obtain several window parameters at different scales, where any window parameter at any scale includes the window radius, window orientation, and the position of the pixel to be processed on the window boundary; For any pixel in the face image, multiple filtering windows for that pixel at the p-th scale are constructed based on several window parameters at the p-th scale. By using the Gaussian kernel function and multiple filtering windows at the p-th scale, any pixel is filtered to obtain the filtered output value output by the multiple filtering windows at the p-th scale. Calculate the difference between the filtered output value of each filter window at the p-th scale and the gray value of any pixel. The filtered output value with the smallest difference is taken as the actual output value at the p-th scale, and the illumination component at the p-th scale is obtained after all pixels in the face image have been traversed. Increment p by 1, and reconstruct multiple filtering windows for any pixel at the p-th scale based on several window parameters at the p-th scale, until p equals P, to obtain the illumination components at each scale, where P is the total number of scales; The illumination components of the captured face image are composed using the illumination components at each scale.

8. An identity verification device based on multimodal feature fusion, characterized in that, include: The acquisition unit is used to acquire facial images and ID card images of the target personnel; The recognition unit is used to perform text recognition processing on the image of the ID card to obtain the identity text information of the target person; The feature extraction unit is used to extract the face image from the ID card image and to extract features from the face capture image and the face image to obtain the actual captured face features and the ID card face features, respectively. A face comparison unit is used to determine, based on the actual captured face features and the ID card face features, whether the face image captured in the face image and the face image in the ID card image belong to the same face; The cross-verification unit is used to search the face database for a matching face corresponding to the identity text information when the face comparison unit determines that the face image and the face image in the ID card image belong to the same face, and to perform face recognition on the face image to obtain the identity information of the target person. The cross-verification unit is used to perform cross-verification of the target person's identity based on the matched face, the captured face image, the identity recognition information, and the identity text information, and obtain the cross-verification result; The cross-verification unit is also used to generate the identity verification result of the target person based on the cross-verification result.

9. An electronic device, characterized in that, include: A memory, a processor, and a transceiver are sequentially connected in communication, wherein the memory is used to store computer programs, the transceiver is used to send and receive messages, and the processor is used to read the computer programs and execute the identity verification method based on multimodal feature fusion as described in any one of claims 1 to 7.

10. A storage medium, characterized in that, The storage medium stores instructions that, when executed on a computer, perform the identity verification method based on multimodal feature fusion as described in any one of claims 1 to 7.