Image processing method, descriptor extraction method and device thereof, and electronic device

Through local and overall brightness alignment and shape alignment processing, the problem of background texture interference in electronic images is solved, and clear target image acquisition is achieved in different environments. In particular, the clarity of fingerprint images is significantly improved in under-screen optical fingerprint recognition.

CN114648548BActive Publication Date: 2025-09-19ARCSOFT CORP LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202011502579.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-12-17
Publication Date
2025-09-19
Estimated Expiration
2040-12-17

AI Technical Summary

Technical Problem

The existing technology cannot effectively remove the background texture in electronic images under different environments, resulting in unclear target images.

Method used

By collecting the first and second original images, performing local brightness alignment processing and overall brightness alignment processing, combined with shape alignment processing, background texture is removed to obtain a clear target image.

Benefits of technology

It effectively removes background textures in various environments and obtains clear target images. In particular, it can clearly display fingerprint images in under-screen optical fingerprint recognition, eliminating the texture interference of the display itself.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114648548B_ABST
    Figure CN114648548B_ABST
Patent Text Reader

Abstract

The present invention discloses an image processing method, a descriptor extraction method, a device therefor, and an electronic device. The image processing method comprises: acquiring a first original image and a second original image, wherein the first original image is one of an original target image and a background image, and the second original image is the other of the original target image and the background image; performing local brightness alignment processing on the first original image to obtain a first processed image; and obtaining a target image with the background removed based on the first processed image and the second original image. The present invention solves the technical problem in the prior art of being unable to obtain a clear target image by removing background texture in different environments.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to image processing technology, and in particular to an image processing method, a descriptor extraction method and a device thereof, and an electronic device. Background Art

[0002] With the development of electronic and communication technologies, the use of electronic images is becoming increasingly widespread. However, electronic images often contain background content that users do not want to appear. For example, in the field of under-screen optical fingerprint recognition, the sensor is located below a specific area of ​​the screen. When the user's finger presses on this area, the phone screen lights up, and the sensor receives the reflection from the surface of the user's finger, forming a fingerprint image. Because the display screen is between the sensor and the finger, the imaging result contains not only fingerprint information, but also the texture information of the display screen itself (i.e., background texture), and the background texture is usually more obvious than the fingerprint information. In order to obtain accurate fingerprint recognition results, we need to remove the background texture to obtain a clear fingerprint image.

[0003] Therefore, image processing technology that can remove background textures contained in electronic images in various environments and obtain clear background-removed images is becoming increasingly important. Summary of the Invention

[0004] The embodiments of the present invention provide an image processing method, a descriptor extraction method and a device thereof, and an electronic device, which at least solve the technical problem in the prior art that a clear target image cannot be obtained by removing background texture under different environments.

[0005] According to one aspect of an embodiment of the present invention, there is provided an image processing method, comprising: acquiring a first original image and a second original image, wherein the first original image is one of an original target image and a background image, and the second original image is the other of the original target image and the background image; performing local brightness alignment processing on the first original image to obtain a first processed image; and obtaining a target image with the background removed based on the first processed image and the second original image.

[0006] Optionally, the original target image is an image of the target object captured when the target object is pressed on the display screen surface of the electronic device, and the background image is a texture image of the display screen itself captured when a simulated object with a reflectivity close to that of the target object and a smooth surface is pressed on the display screen surface of the electronic device.

[0007] Optionally, the target object is a finger, and the simulation object is a skin-colored rubber block.

[0008] Optionally, when the target object fully presses the screen, the image processing method continuously captures multiple frames of target object images, performs overall or local weighted fusion on the multiple frames of target object images according to their overall quality or local quality, and obtains the fused target image as the original target image.

[0009] Optionally, performing local brightness alignment processing on the first original image to obtain a first processed image includes: calculating the pixel value of each pixel in the first original image and the pixel average of all pixels in the neighborhood window of each pixel, as well as the pixel value of each pixel in the second original image and the pixel average of all pixels in the neighborhood window of each pixel; and obtaining the first processed image that is locally brightness aligned with the second original image based on the pixel value of each pixel in the first original image and the pixel average of all pixels in the neighborhood window of each pixel, as well as the pixel value of each pixel in the second original image and the pixel average of all pixels in the neighborhood window of each pixel.

[0010] Optionally, obtaining the first result image based on the first processed image and the second original image includes: performing a subtraction operation on the first processed image and the second original image to obtain the first result image.

[0011] Optionally, the image processing method further includes performing an overall brightness alignment process before or after the local brightness alignment process.

[0012] Optionally, after performing local brightness alignment processing on the first original image to obtain the first processed image, performing overall brightness alignment processing on the first processed image to obtain the second processed image; wherein, performing overall brightness alignment processing on the first processed image to obtain the second processed image includes: respectively calculating the maximum pixel value and the minimum pixel value of the first processed image, and the maximum pixel value and the minimum pixel value of the second original image; obtaining the overall brightness scale coefficient and the overall brightness offset coefficient of the first processed image relative to the second original image based on the maximum pixel value and the minimum pixel value of the first processed image, and the maximum pixel value and the minimum pixel value of the second original image; performing a linear transformation on the first processed image based on the overall brightness scale coefficient and the overall brightness offset coefficient to obtain a second processed image whose overall brightness is aligned with the second original image.

[0013] Optionally, the second original image is subjected to overall brightness alignment processing to obtain a third processed image, including: respectively calculating the maximum pixel value and the minimum pixel value of the first processed image, and the maximum pixel value and the minimum pixel value of the second original image; obtaining the overall brightness scale coefficient and the overall brightness offset coefficient of the second original image relative to the first processed image based on the maximum pixel value and the minimum pixel value of the first processed image, and the maximum pixel value and the minimum pixel value of the second original image; and performing a linear transformation on the second original image based on the overall brightness scale coefficient and the overall brightness offset coefficient to obtain a third processed image whose overall brightness is aligned with that of the first processed image.

[0014] Optionally, the image processing method further includes performing a smoothing process before the overall brightness alignment.

[0015] Optionally, the smoothing process includes at least one of the following: mean filtering and Gaussian filtering.

[0016] Optionally, the image processing method further includes: using an image of the target object captured when the target object just contacts the screen but does not fully press the display screen as a background image.

[0017] Optionally, the image processing method further includes: performing shape alignment processing before or after the local brightness alignment processing.

[0018] Optionally, after performing local brightness alignment processing on the first original image to obtain the first processed image, shape alignment processing is performed on the first processed image to obtain a fourth processed image; wherein, performing shape alignment processing on the first processed image to obtain the fourth processed image includes: calculating the position offset of each pixel point in the first processed image to the corresponding target pixel point in the second original image; fitting the displacement parameters and scale parameters of the first processed image relative to the second original image based on the position offset; performing shape alignment processing on the first processed image based on the displacement parameters and scale parameters to obtain a fourth processed image that is shape-aligned with the second original image.

[0019] Optionally, shape alignment processing is performed on the second original image to obtain a fifth processed image, including: calculating the position offset of each pixel point in the first processed image to the corresponding target pixel point in the second original image; fitting the displacement parameters and scale parameters of the second original image relative to the first processed image based on the position offset; and performing shape alignment processing on the second original image based on the displacement parameters and scale parameters to obtain a fifth processed image that is shape-aligned with the first processed image.

[0020] Optionally, the image processing method further includes: performing at least one of the following processing on the target image with the background removed: local contrast enhancement, fast non-local mean denoising, and three-dimensional block matching filtering.

[0021] Optionally, the image processing method further includes: before performing the shape alignment process, determining whether the background image has deformation.

[0022] Optionally, the image processing method further includes: acquiring a candidate background image, wherein the candidate background image is a plurality of frames of target object images sampled from the time when the target object just contacts the screen to the time when the target object completely presses the screen.

[0023] Optionally, the image processing method further includes: obtaining a result image after removing the candidate background based on the candidate background image and the first original image or the second original image that has undergone local brightness alignment processing.

[0024] Optionally, the image processing method further includes: taking a better one of the result image after background removal and the result image after candidate background removal as a target image for background removal.

[0025] According to another aspect of an embodiment of the present invention, there is provided an image processing method, comprising: acquiring a third original image and a fourth original image, wherein the third original image is one of an original target image and a background image, and the fourth original image is the other of the original target image and the background image; obtaining an input image based on the third original image and the fourth original image; training an initial image generation network to construct a trained image generation network; wherein the image generation network is trained using a template image as a reference object, and the template image is a target image with the background removed obtained using any of the above-mentioned image processing methods; and inputting the input image into the trained image generation network to obtain a target image with the background removed.

[0026] Optionally, obtaining the input image according to the third original image and the fourth original image includes: performing local brightness alignment processing on the third original image and the fourth original image, and using the processing result as the input image.

[0027] Optionally, the initial image generation network is trained, and constructing a trained image generation network includes: obtaining a first sample image and a second sample image; inputting the first sample image and the second sample image into the initial image generation network to obtain a first training image; inputting the first training image and the template image into a first loss function module, training the initial image generation network according to a first loss value output by the first loss function module, and constructing a trained image generation network.

[0028] Optionally, the initial image generation network is trained, and constructing a trained image generation network also includes: inputting the first training image and the template image into the feature extraction network to obtain high-level semantic features of the first training image and the high-level semantic features of the template image; inputting the high-level semantic features of the first training image and the high-level semantic features of the template image into the second loss function module, and training the image generation network and / or the feature extraction network according to the second loss value output by the second loss function module.

[0029] According to another aspect of an embodiment of the present invention, a descriptor extraction method is provided, comprising: obtaining key points of a target image, wherein the target image is a target image with the background removed obtained using any of the above-mentioned image processing methods; classifying multiple target images with the same key points into the same category, and marking the category information as image labels; centering on the key points, intercepting image blocks from multiple target images with image labels representing the same category; training an initial descriptor extraction network to construct a trained descriptor extraction network; and inputting the image blocks into the trained descriptor extraction network to obtain feature descriptors.

[0030] Optionally, before inputting the image block into a trained descriptor extraction network to obtain a feature descriptor, the descriptor extraction method further includes: obtaining the direction and angle of the key point; and rotating and flattening the image block according to the direction and angle of the key point to obtain an aligned image block.

[0031] Optionally, the descriptor extraction method uses an end-to-end training method to train the image generation network, the feature extraction network, and the descriptor extraction network.

[0032] According to another aspect of an embodiment of the present invention, an image processing device is provided, including: an image acquisition unit, configured to acquire a first original image and a second original image, wherein the first original image is one of an original target image and a background image, and the second original image is the other of the original target image and the background image; a local brightness alignment unit, configured to perform local brightness alignment processing on the first original image to obtain a first processed image; and a result acquisition unit, configured to obtain a first result image based on the first processed image and the second original image.

[0033] Optionally, the local brightness alignment unit includes: a first pixel value calculation unit, used to respectively calculate the pixel value of each pixel point in the first original image and the pixel average of all pixels points in the neighborhood window of each pixel point, as well as the pixel value of each pixel point in the second original image and the pixel average of all pixels points in the neighborhood window of each pixel point; a first processed image acquisition unit, used to obtain a first processed image that is locally brightness aligned with the second original image based on the pixel value of each pixel point in the first original image and the pixel average of all pixels points in the neighborhood window of each pixel point, as well as the pixel value of each pixel point in the second original image and the pixel average of all pixels points in the neighborhood window of each pixel point.

[0034] Optionally, the result acquisition unit obtains the first result image by performing a subtraction operation on the first processed image and the second original image.

[0035] Optionally, the image processing apparatus further includes an overall brightness alignment unit configured to perform overall brightness alignment processing before or after the local brightness alignment processing.

[0036] Optionally, the overall brightness alignment unit includes: a second pixel value calculation unit, used to calculate the maximum pixel value and the minimum pixel value of the first processed image, and the maximum pixel value and the minimum pixel value of the second original image respectively; an overall brightness coefficient calculation unit, which obtains the overall brightness scale coefficient and the overall brightness offset coefficient of the first processed image relative to the second original image based on the maximum pixel value and the minimum pixel value of the first processed image, and the maximum pixel value and the minimum pixel value of the second original image; a linear transformation unit, which performs a linear transformation on the first processed image based on the overall brightness scale coefficient and the overall brightness offset coefficient to obtain a second processed image aligned with the overall brightness of the second original image.

[0037] Optionally, the overall brightness alignment unit includes: a second pixel value calculation unit, used to calculate the maximum pixel value and the minimum pixel value of the first processed image, and the maximum pixel value and the minimum pixel value of the second original image respectively; an overall brightness coefficient calculation unit, which obtains the overall brightness scale coefficient and the overall brightness offset coefficient of the second original image relative to the first processed image based on the maximum pixel value and the minimum pixel value of the first processed image, and the maximum pixel value and the minimum pixel value of the second original image; a linear transformation unit, which performs a linear transformation on the second original image based on the overall brightness scale coefficient and the overall brightness offset coefficient to obtain a third processed image whose overall brightness is aligned with the first processed image.

[0038] Optionally, the image processing device further includes a smoothing processing unit configured to perform smoothing processing before overall brightness alignment.

[0039] Optionally, the image processing apparatus further includes a shape alignment unit configured to perform shape alignment processing before or after the local brightness alignment processing.

[0040] Optionally, the shape alignment unit includes: a position calculation unit, used to calculate the position offset of each pixel point in the first processed image to the corresponding target pixel point in the second original image; a fitting unit, used to fit the displacement parameters and scale parameters of the first processed image relative to the second original image based on the position offset; and an alignment unit, used to perform shape alignment processing on the first processed image based on the displacement parameters and scale parameters to obtain a fourth processed image that is shape-aligned with the second original image.

[0041] Optionally, the shape alignment unit includes: a position calculation unit, used to calculate the position offset of each pixel point in the first processed image to the corresponding target pixel point in the second original image; a fitting unit, used to fit the displacement parameters and scale parameters of the second original image relative to the first processed image based on the position offset; and an alignment unit, used to perform shape alignment processing on the second original image based on the displacement parameters and scale parameters to obtain a fifth processed image that is shape-aligned with the first processed image.

[0042] According to another aspect of an embodiment of the present invention, an image processing device is provided, including: an image acquisition unit, which acquires a third original image and a fourth original image, wherein the third original image is one of the original target image and the background image, and the fourth original image is the other of the original target image and the background image; an input image acquisition unit, which obtains an input image based on the third original image and the fourth original image; a network construction unit, which trains an initial image generation network to construct a trained image generation network; wherein the image generation network is trained with a template image as a reference object, and the template image is a target image with the background removed obtained using any of the above-mentioned image processing methods; and a result acquisition unit, which inputs the input image into the trained image generation network to obtain the target image with the background removed.

[0043] Optionally, the input image acquisition unit is configured to perform local brightness alignment processing on the third original image and the fourth original image, and use the processing result as the input image.

[0044] Optionally, the network construction unit includes: a sample acquisition unit for acquiring a first sample image and a second sample image; a training image acquisition unit for inputting the first sample image and the second sample image into an initial image generation network to obtain a first training image; a first training unit for inputting the first training image and the template image into a first loss function module, training the initial image generation network according to the first loss value output by the first loss function module, and constructing the trained image generation network.

[0045] Optionally, the network construction unit also includes: a feature extraction unit, used to input the first training image and the template image into the feature extraction network to obtain high-level semantic features of the first training image and the high-level semantic features of the template image; a second training unit, used to input the high-level semantic features of the first training image and the high-level semantic features of the template image into the second loss function module, and train the image generation network and / or feature extraction network according to the second loss value output by the second loss function module.

[0046] According to another aspect of an embodiment of the present invention, a descriptor extraction device is provided, comprising: a key point acquisition unit for acquiring key points of a target image, wherein the target image is a target image obtained by removing the background using any of the above-mentioned image processing methods; a classification unit for classifying multiple target images with the same key points into the same category, and marking the category information as an image label; a screenshot unit for capturing image blocks centered on the key points from multiple target images with image labels representing the same category; a descriptor extraction network construction unit for training an initial descriptor extraction network and constructing a trained descriptor extraction network; and a feature descriptor acquisition unit for inputting image blocks into the trained descriptor extraction network to obtain feature descriptors.

[0047] Optionally, the descriptor extraction device further includes: a direction angle acquisition unit for acquiring the direction and angle of the key point; and a rotation and flattening unit for rotating and flattening the image block according to the direction and angle of the key point to obtain an aligned image block.

[0048] Optionally, the descriptor extraction device uses an end-to-end training method to train the image generation network, the feature extraction network and the descriptor extraction network.

[0049] According to another aspect of an embodiment of the present invention, a storage medium is provided, comprising a stored program, wherein when the program is executed, the device where the storage medium is located is controlled to execute any one of the above-mentioned image processing methods.

[0050] According to another aspect of an embodiment of the present invention, an electronic device is provided, comprising: a processor; and a memory for storing executable instructions of the processor; wherein the processor is configured to perform any one of the above-mentioned image processing methods by executing the executable instructions.

[0051] In an embodiment of the present invention, the following steps are performed: capturing a first original image and a second original image, wherein the first original image is one of an original target image and a background image, and the second original image is the other of the two; performing local brightness alignment processing on the first original image to obtain a first processed image; and obtaining a background-removed target image based on the first processed image and the second original image. This allows background textures contained in the original image to be removed in various environments, resulting in a clear target image. This solves the technical problem in the prior art of being unable to obtain a clear target image by removing background textures in different environments. BRIEF DESCRIPTION OF THE DRAWINGS

[0052] The drawings described herein are used to provide a further understanding of the present invention and constitute a part of this application. The exemplary embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation of the present invention. In the drawings:

[0053] Figure 1 is a flowchart of an optional image processing method according to an embodiment of the present invention;

[0054] Figure 2 is a flowchart of another optional image processing method according to an embodiment of the present invention;

[0055] Figure 3 is a flowchart of another optional image processing method according to an embodiment of the present invention;

[0056] Figure 4 is a flowchart of another optional image processing method according to an embodiment of the present invention;

[0057] Figure 5 is a flowchart of another optional image processing method according to an embodiment of the present invention;

[0058] Figure 6 is a flowchart of an optional deep learning-based image processing method according to an embodiment of the present invention;

[0059] Figure 7 is a flowchart of an optional descriptor extraction method according to an embodiment of the present invention;

[0060] Figure 8 is a structural block diagram of an optional image processing device according to an embodiment of the present invention;

[0061] Figure 9 is a structural block diagram of another optional image processing device according to an embodiment of the present invention;

[0062] Figure 10is a structural block diagram of another optional image processing device according to an embodiment of the present invention;

[0063] Figure 11 is a structural block diagram of another optional deep learning-based image processing device according to an embodiment of the present invention;

[0064] Figure 12 This is a structural block diagram of another optional descriptor extraction device according to an embodiment of the present invention. DETAILED DESCRIPTION

[0065] In order to enable those skilled in the art to better understand the solutions of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the embodiments described are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of the present invention.

[0066] It should be noted that the terms "first", "second", etc. in the description and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that the order used in this way can be interchanged where appropriate, so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0067] The following is a flowchart of an optional image processing method according to an embodiment of the present invention. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than shown here.

[0068] It should be noted that, in order to explain more clearly, in different embodiments of the following image processing method, the target image with the background removed from the original target image is described as the first result image, the second result image, the third result image, the fourth result image, the fifth result image, and the sixth result image, respectively.

[0069] refer to Figure 1 , is a flow chart of an optional image processing method according to an embodiment of the present invention. Figure 1As shown, the image processing method includes the following steps:

[0070] S100 , capturing a first original image and a second original image, wherein the first original image is one of an original target image and a background image, and the second original image is the other of the original target image and the background image.

[0071] In an optional embodiment, the original target image is an image of the target object captured when the target object is pressed against the display screen of an electronic device, and the background image is a texture image of the display screen itself captured when a smooth simulant with a reflectivity close to that of the target object is pressed against the display screen of the electronic device. When the image processing method is specifically applied to fingerprint images, the target object may be a finger, and the simulant may be a skin-colored rubber block.

[0072] In another optional embodiment, since the distribution of the noise signal has a certain randomness, while the target image signal is relatively fixed, in order to improve the quality of the target image, a multi-frame image fusion method is also used to reduce the noise signal while maintaining the target image signal strength unchanged. For example, when the target fully presses the screen, multiple frames of target images are continuously captured, and then the multiple frames of target images are weightedly fused as a whole or locally based on the overall quality or local quality of the multiple frames of target images to obtain a fused target image as the original target image. Compared with capturing a single frame of the target image as the original target image, the original target image obtained using the multi-frame image fusion method has less overall noise, providing a better foundation for subsequent target image processing steps.

[0073] S102: Perform local brightness alignment processing on the first original image to obtain a first processed image.

[0074] Although the reflectivity of the simulated object is close to that of the target object, the actual reflectivity is still not completely the same as that of the simulated object, resulting in uneven local brightness between the first original image and the second original image. In an optional embodiment, step S102 includes:

[0075] S1020: Calculate the pixel value r of each pixel in the first original image and the pixel average value of all pixels in the neighborhood window of each pixel. And the pixel value b of each pixel in the second original image and the pixel average of all pixels in the neighborhood window of each pixel The neighborhood window can be set according to actual conditions. For example, when the image processing method is specifically applied to process fingerprint images, an 11*11 window range centered on each pixel is selected as the neighborhood window.

[0076] S1022: Based on the pixel value r of each pixel in the first original image and the pixel average value of all pixels in the neighborhood window of each pixel And the pixel average of all pixels in the neighborhood window of each pixel in the second original image A first processed image that is locally aligned with the second original image is obtained, wherein the pixel value r1 of each pixel in the first processed image can be obtained according to Formula 1:

[0077] Therefore, through the above steps S1020 and S1022, local brightness alignment processing can be performed on the first original image, so that the local brightness of the first processed image after the local brightness alignment processing is consistent with that of the second original image.

[0078] S104: Obtain a first result image based on the first processed image and the second original image.

[0079] In an optional embodiment, a subtraction operation can be performed on the first processed image and the second original image to obtain a first result image. For example, when the first original image is the original target image and the second original image is the background image, the first processed image is the original target image that has undergone local brightness alignment processing. Subtracting the second original image from the first processed image can obtain the first result image, which is the target image with the background removed. For another example, when the first original image is the background image and the second original image is the original target image, the first processed image is the background image that has undergone local brightness alignment processing. Subtracting the first processed image from the second original image can also obtain the first result image, which is also the target image with the background removed.

[0080] It should be noted that, in the embodiment of the present application, performing a subtraction operation on the first processed image and the second original image means performing a subtraction operation on the pixel value of each pixel point of the first processed image and the pixel value of each pixel point of the second original image.

[0081] The image processing method provided by the embodiment of the present invention can not only remove the main background texture, but also solve the problem of uneven local brightness between the background image and the original target image, so that the local brightness of the target image after removing the background is relatively uniform and clear. When specifically applied to processing fingerprint images, a clear fingerprint image with the background texture of the display screen removed can be obtained.

[0082] However, due to hardware imaging anomalies or the influence of the external environment, sometimes the overall brightness difference between the first original image and the second original image is large, and the first result image obtained after the above steps S100-S104 still has obvious background texture residue. Therefore, the overall brightness alignment process can be performed before or after the local brightness alignment process. Figure 2 , provides a flowchart of another optional image processing method according to an embodiment of the present invention. Figure 2 As shown, the image processing method includes the following steps:

[0083] S200 , capturing a first original image and a second original image, wherein the first original image is one of an original target image and a background image, and the second original image is the other of the original target image and the background image.

[0084] S202: Perform local brightness alignment processing on the first original image to obtain a first processed image.

[0085] S204: Perform overall brightness alignment processing on the first processed image to obtain a second processed image.

[0086] S206: Obtain a second result image based on the second processed image and the second original image.

[0087] The above steps S200, S202 and Figure 1 Steps S100 and S102 in the described embodiment are the same, for details, see Figure 1 The corresponding description will not be described in detail here. Figure 2 The described embodiments and Figure 1 The difference is that the image processing method further includes step S204, performing overall brightness alignment processing on the first processed image to obtain a second processed image, and step S206 is to obtain a second result image based on the second processed image that has undergone overall brightness alignment processing and the second original image.

[0088] In an optional embodiment, step S204 includes:

[0089] S2042: Calculate the maximum pixel value rmax and the minimum pixel value rmin of the first processed image, and the maximum pixel value bmax and the minimum pixel value bmin of the second original image respectively;

[0090] S2044: Obtain an overall brightness scale coefficient α and an overall brightness offset coefficient β of the first processed image relative to the second original image based on the maximum pixel value rmax and the minimum pixel value rmin of the first processed image and the maximum pixel value bmax and the minimum pixel value bmin of the second original image; wherein the overall brightness scale coefficient α can be calculated according to Formula 2: The overall brightness shift coefficient β can be calculated according to formula 3: β = (b min ·r max -b max ·r min ) / (r max -rmin );

[0091] S2046: Perform a linear transformation on the first processed image based on the overall brightness scale coefficient α and the overall brightness offset coefficient β to obtain a second processed image that is aligned with the overall brightness of the second original image; wherein the pixel value r2 of each pixel in the second processed image is obtained by linear transformation based on Formula 4, Formula 4: r2 = α·r + β.

[0092] In another optional embodiment, the second original image may also be subjected to overall brightness alignment processing, referring to Figure 3 , provides a flowchart of another optional image processing method according to an embodiment of the present invention. Figure 3 As shown, the image processing method includes the following steps:

[0093] S300 , capturing a first original image and a second original image, wherein the first original image is one of an original target image and a background image, and the second original image is the other of the original target image and the background image.

[0094] S302: Perform local brightness alignment processing on the first original image to obtain a first processed image.

[0095] S304: Perform overall brightness alignment processing on the second original image to obtain a third processed image.

[0096] S306: Obtain a third result image based on the first processed image and the third processed image.

[0097] In an optional embodiment, step S304 includes:

[0098] S3042: Calculate the maximum pixel value rmax and the minimum pixel value rmin of the first processed image, and the maximum pixel value bmax and the minimum pixel value bmin of the second original image respectively;

[0099] S3044: Obtain an overall brightness scale coefficient α′ and an overall brightness offset coefficient β′ of the second original image relative to the first processed image based on the maximum pixel value rmax and the minimum pixel value rmin of the first processed image, and the maximum pixel value bmax and the minimum pixel value bmin of the second original image; wherein the overall brightness scale coefficient α′ can be calculated according to Formula 2, Formula 4: The overall brightness shift coefficient β′ can be calculated according to Formula 3, Formula 5: β′=(r min b max -r max b min ) / (b max -b min );

[0100] S3046: Perform a linear transformation on the second original image based on the overall brightness scale coefficient α′ and the overall brightness offset coefficient β′ to obtain a third processed image that is aligned with the overall brightness of the first processed image; wherein the pixel value b2 of each pixel point in the third processed image is obtained by linear transformation based on Formula 4, Formula 4: b2 = α′·b + β′.

[0101] It should be noted that in the above embodiment, local brightness alignment is performed first, followed by global brightness alignment. However, the above step sequence is merely an example and not a strict limitation. Those skilled in the art may make corresponding adjustments and modifications based on the principles of the above embodiment, for example, performing global brightness alignment first and then local brightness alignment.

[0102] Furthermore, to avoid the impact of local noise on parameter calculation, in embodiments including an overall brightness alignment step, the image processing method may further include performing a smoothing process before the overall brightness alignment process. For example, smoothing may be performed on the first original image and the second original image; or on the first processed image and the second original image, and so on. Those skilled in the art may make reasonable adjustments based on the specific embodiment. Specifically, the smoothing process may be performed using methods such as mean filtering and Gaussian filtering.

[0103] In addition, due to changes in the external environment and different pressing methods, the background texture contained in the original target image may be significantly different from the background image, that is, there is deformation of the background texture. For example, when the external temperature changes significantly, the internal structure of the display screen will undergo different degrees of thermal expansion and contraction effects. At this time, the imaging of the background texture in the original target image is significantly different from the background image. For example, when the target image is captured, due to the different forces of the target object pressing the display screen, the display screen will undergo different degrees of deformation, and the distance between it and the imaging sensor will also change, thereby causing the imaging of the background texture to change. In this case, the background texture may not be removed through the above embodiment.

[0104] To address the problem of background texture deformation, in an optional embodiment, an image of the target object captured when the target object has just touched the screen but has not yet fully pressed the display screen is used as the background image. When the target object just touches the display screen, the screen is illuminated, but the target object has not yet been fully pressed. The target image signal in the sensor imaging is weak, while the background texture is clear. Therefore, the target object image captured at this moment can be approximately used as the background image. By precisely controlling the image acquisition moment, a background image that is substantially consistent with the background texture in the current original target image can be obtained, thereby effectively addressing the problem of background texture deformation. However, this method has high requirements for the timing of image acquisition. If the image is acquired too early, the screen has not yet been illuminated or the target object is far away from the screen and the reflection is insufficient, resulting in a weak imaging intensity. If the image is acquired too late, the target object has already been fully pressed, and a pure background image cannot be obtained.

[0105] In order to solve the problem of background texture deformation without strict requirements on the timing of image acquisition, shape alignment can be performed before or after local brightness alignment. Figure 4 , provides a flowchart of another optional image processing method according to an embodiment of the present invention. Figure 4 As shown, the image processing method includes the following steps:

[0106] S400 , capturing a first original image and a second original image, wherein the first original image is one of an original target image and a background image, and the second original image is the other of the original target image and the background image.

[0107] S402: Perform local brightness alignment processing on the first original image to obtain a first processed image.

[0108] S404: Perform shape alignment processing on the first processed image to obtain a fourth processed image.

[0109] S406: Obtain a fourth result image based on the fourth processed image and the second original image.

[0110] The above steps S400, S402 and Figure 1 The embodiments described are the same, for details, please refer to Figure 1 The corresponding description will not be described in detail here. Figure 4 The described embodiments and Figure 1 The difference is that the image processing method also includes step S404, performing shape alignment processing on the first processed image to obtain a fourth processed image, and step S406 is obtaining a fourth result image based on the fourth processed image after shape alignment processing and the second original image.

[0111] In another optional embodiment, before performing the shape alignment processing step, the following may be included: determining whether the background image is deformed; optionally, determining whether the background image is deformed includes: determining the similarity between the first processed image and the second original image, specifically, for example, the zero-mean normalized cross-correlation coefficient ZNCC (zero-mean normalized cross-correlation) may be used for judgment; when the similarity is less than a first threshold, determining that the background image is deformed.

[0112] In an optional embodiment, step S404 includes:

[0113] S4042: Calculate the position offset between each pixel in the first processed image and the corresponding target pixel in the second original image; optionally, the Lucas-Kanade algorithm may be used for calculation;

[0114] S4044: fitting the displacement parameters (Δx, Δy) and scale parameter 6 of the first processed image relative to the second original image based on the position offset; optionally, a least squares method may be used for fitting;

[0115] S4046: Perform shape alignment processing on the first processed image according to the displacement parameters (Δx, Δy) and the scale parameter δ to obtain a fourth processed image that is shape-aligned with the second original image.

[0116] In another optional embodiment, the second original image may also be subjected to shape alignment processing, referring to Figure 5 , provides a flowchart of another optional image processing method according to an embodiment of the present invention. Figure 5 As shown, the image processing method includes the following steps:

[0117] S500 , capturing a first original image and a second original image, wherein the first original image is one of an original target image and a background image, and the second original image is the other of the original target image and the background image.

[0118] S502: Perform local brightness alignment processing on the first original image to obtain a first processed image.

[0119] S504: Perform shape alignment processing on the second original image to obtain a fifth processed image.

[0120] S506: Obtain a fifth result image based on the first processed image and the fifth processed image.

[0121] The above steps S400, S402 and Figure 1 The embodiments described are the same, for details, please refer to Figure 1 The corresponding description will not be described in detail here. Figure 5The described embodiments and Figure 1 The difference is that the image processing method also includes step S504, performing shape alignment processing on the second original image to obtain a fifth processed image, and step S506 is the first processed image and the fifth processed image based on the shape alignment processing to obtain a fifth result image.

[0122] In an optional embodiment, step S504 includes:

[0123] S5042: Calculate the position offset between each pixel in the first processed image and the corresponding target pixel in the second original image; optionally, the Lucas-Kanade algorithm may be used for calculation;

[0124] S5044: fitting the displacement parameters (Δx′, Δy′) and scale parameter δ′ between the second original image and the first processed image based on the position offset; optionally, a least squares method may be used for fitting;

[0125] S5046: Perform shape alignment processing on the second original image according to the displacement parameters (Δx′, Δy′) and the scale parameter δ′ to obtain a fifth processed image that is shape-aligned with the first processed image.

[0126] It should be noted that, in the above embodiment, local brightness alignment is performed first, and then shape alignment is performed. However, the above step sequence is merely an example and not a strict limitation. Those skilled in the art may make corresponding adjustments and transformations based on the principles of the above embodiment. For example, shape alignment may be performed first, and then local brightness alignment. In an image processing method including an overall brightness alignment step, the order of the three steps, overall brightness alignment, local brightness alignment, and shape alignment, may be arranged in different ways. For example, overall brightness alignment may be performed first, and then local brightness alignment, and finally shape alignment. For another example, local brightness alignment may be performed first, and then overall brightness alignment, and finally shape alignment. For another example, shape brightness alignment may be performed first, and then overall brightness alignment, and finally local brightness alignment.

[0127] In an optional embodiment, the step of capturing the first original image and the second original image may further include capturing a candidate background image, that is, sampling multiple frames of target images from the time the target just touches the screen to the time the target fully presses the screen as the candidate background image. Corresponding to this step, the image processing method further includes obtaining a result image after removing the candidate background based on the candidate background image and the first original image or the second original image that has at least undergone local brightness alignment processing; and selecting Figure 1-Figure 5In the described embodiment, the better one of the result image after background removal and the result image after candidate background removal is used as the target image for background removal.

[0128] Due to the imaging characteristics of the sensor, the resulting background-removed target image may exhibit localized uneven quality, typically manifesting as better quality in the center and poorer quality in the edge regions. Therefore, the background-removed target image needs to be enhanced to improve the quality of the edge regions. Poor quality in the edge regions relative to the center is primarily due to lower contrast in the fingerprint ridges. Local contrast enhancement can be used to improve this, but this also amplifies existing noise. Therefore, denoising can be performed after local contrast enhancement. Optional denoising methods include Fast Non-Local Means Denoising and BM3D (block-matching and 3D filtering). After local contrast enhancement and denoising, a clearer target image is obtained.

[0129] The above image processing method can produce a target image with background removed. The target image is clear overall, has uniform overall and local brightness, low noise, and no noticeable background texture residue. It also exhibits good adaptability to changes in the external environment. Specifically, when applied to fingerprint image processing, a clear fingerprint image with the background texture of the display screen removed can be obtained. The fingerprint features clear lines, uniform brightness, and no noticeable background texture residue. This method can overcome the effects of background deformation caused by changes in the external environment and finger pressure, and has no strict requirements on the timing of fingerprint collection, demonstrating good adaptability.

[0130] Through the above Figure 1-Figure 5 The image processing method described above obtains a clear target image with background removed. If it is used as a reference image, the background in the image can also be removed using deep learning methods. Figure 6 , is a flowchart of an optional image processing method based on deep learning according to an embodiment of the present invention. Figure 6 As shown, the image processing method includes the following steps:

[0131] S600, capturing a third original image and a fourth original image; wherein the third original image is one of the original target image and the background image, and the fourth original image is the other of the original target image and the background image;

[0132] In an optional embodiment, the third original image is an image of a target object captured when the target object is pressed against the display screen of an electronic device, and the background image is a texture image of the display screen itself captured when a smooth simulant with a reflectivity close to that of the target object is pressed against the display screen of the electronic device. When the image processing method is specifically applied to fingerprint images, the target object may be a finger, and the simulant may be a skin-colored rubber block.

[0133] S602, obtaining an input image according to the third original image and the fourth original image;

[0134] In an optional embodiment, the third original image and the fourth original image may be used as multi-channel input images;

[0135] In another optional embodiment, obtaining the input image based on the third original image and the fourth original image includes: performing a subtraction operation on the third original image and the fourth original image, and using the operation result as the input image; specifically, the fourth original image can be subtracted from the third original image, that is, the image with the background removed can be obtained as the input image.

[0136] In another optional embodiment, obtaining the input image according to the third original image and the fourth original image includes: performing local brightness alignment processing on the third original image and the fourth original image, and using the processing result as the input image; specifically, the local brightness alignment processing can be performed according to the following example: Figure 1 Step S102 described in is implemented.

[0137] S604, training the initial image generation network to construct a trained image generation network; wherein the image generation network is trained with the template image as a reference object, and the template image is used as a reference object. Figure 1-Figure 5 The target image with background removed is obtained by the image processing method described in the embodiment.

[0138] In an optional embodiment, the image generation network is a U-net network. Specifically, the image generation network may include three parts: a first part including two convolutional modules, a second part including four simple residual modules, and a third part including two deconvolutional modules. This structure ensures that the output image size of the image generation network remains consistent with the input image.

[0139] In an optional embodiment, training the initial image generation network to construct the trained image generation network includes:

[0140] S6042: Acquire a first sample image and a second sample image;

[0141] S6044: Inputting the first sample image and the second sample image into the initial image generation network to obtain a first training image;

[0142] S6046: Input the first training image and the template image into a first loss function module, train the initial image generation network according to the first loss value output by the first loss function module, and construct a trained image generation network. The first loss function module may be an L2-LOSS function.

[0143] In another optional embodiment, training the initial image generation network and constructing the trained image generation network further includes:

[0144] S6048: Inputting the first training image and the template image into a feature extraction network to obtain high-level semantic features of the first training image and high-level semantic features of the template image;

[0145] In an optional embodiment, the feature extraction network may select a VGG network.

[0146] S6050: Inputting the high-level semantic features of the first training image and the high-level semantic features of the template image into a second loss function module, and training the image generation network and / or the feature extraction network according to a second loss value output by the second loss function module. The second loss function module may be an L2-LOSS function.

[0147] In this embodiment, by adding a feature extraction network to obtain high-level semantic features, and training the image generation network and / or feature extraction network through a second loss function, the final generated target image can be made closer to the template image.

[0148] S606: Input the input image to the trained image generation network to obtain a sixth result image.

[0149] According to the embodiment provided by the above steps S600-S606, the background texture in the third original image or the fourth original image can be removed by deep learning methods in various environments without strict requirements on the quality of the third original image or the fourth original image.

[0150] Of course, those skilled in the art will know that Figure 6 The corresponding deep learning-based image processing method can also directly use high-quality clear images as reference objects for training instead of using Figure 1-Figure 5 The background-removed result image obtained by the image processing method described in the embodiment is used as a reference object for training.

[0151] Through the above Figures 1-6 The image processing method described above obtains a clear background-removed target image, and thus, a more accurate descriptor can be extracted from the background-removed target image. Figure 7, is a flow chart of an optional descriptor extraction method according to an embodiment of the present invention. Figure 7 As shown, the descriptor extraction method includes the following steps:

[0152] S700, obtaining key points of a target image, wherein the target image is Figures 1-6 The target image with background removed obtained by the image processing method described in the embodiment;

[0153] In an optional embodiment, the SIFT algorithm is used to determine the key points of the target image.

[0154] S702, classifying multiple target images with the same key points into the same category, and marking the category information as image labels;

[0155] S704, taking the key point as the center, intercepting an image block from multiple target images having image labels representing the same category;

[0156] In an optional embodiment, when specifically applied to processing fingerprint images, an image block of size 32*32 may be captured with the key point as the center.

[0157] S706, training the initial descriptor extraction network to construct a trained descriptor extraction network;

[0158] In an optional embodiment, the descriptor extraction network is trained according to a third loss value output by a third loss function module. The third loss function can be implemented by a variety of functions, such as tiplet loss, N-pairs loss, histogram loss, contrastive loss, circle loss, etc.

[0159] S708: Input the image block into a trained descriptor extraction network to obtain a feature descriptor.

[0160] When specifically applied to processing fingerprint images, since fingerprints contain too much repetitive texture information, directly inputting image blocks into a trained descriptor extraction network will reduce the robustness of the descriptor extraction method. To address this issue, the present invention provides an optional embodiment. Before inputting the image blocks into the trained descriptor extraction network to obtain feature descriptors, the descriptor extraction method further includes: obtaining the direction and angle of key points; and rotating and flattening the image blocks according to the direction and angle of the key points to obtain aligned image blocks.

[0161] In an optional embodiment, the image generation network and the feature extraction network can be trained separately first, and then the descriptor extraction network can be trained. Finally, the image generation network, the feature extraction network and the descriptor extraction network can be trained using an end-to-end training method, and the network parameters can be fine-tuned to obtain a more robust feature descriptor.

[0162] According to another aspect of an embodiment of the present invention, an electronic device is provided, including: a processor; and a memory for storing executable instructions of the processor; wherein the processor is configured to perform any one of the above-mentioned image processing methods by executing the executable instructions.

[0163] According to another aspect of an embodiment of the present invention, a storage medium is provided. The storage medium includes a stored program, wherein when the program is running, the device where the storage medium is located is controlled to execute any one of the above-mentioned image processing methods.

[0164] According to another aspect of the present invention, an image processing device is provided. Figure 8 , is a structural block diagram of an optional image processing device according to an embodiment of the present invention. Figure 8 As shown, the image processing device 80 includes an image acquisition unit 800 , a local brightness alignment unit 802 and a result acquisition unit 804 .

[0165] The following is a detailed description of each unit included in the image processing device 80.

[0166] The image acquisition unit 800 is configured to acquire a first original image and a second original image, wherein the first original image is one of an original target image and a background image, and the second original image is the other of the original target image and the background image.

[0167] In an optional embodiment, the original target image refers to an image of the target object captured when the target object is pressed against the display screen of an electronic device, and the background image refers to a texture image of the display screen itself captured when a smooth simulant with a reflectivity close to that of the target object is pressed against the display screen of the electronic device. When the image acquisition unit 800 is specifically used to process fingerprint images, the target object may be a finger, and the simulant may be a skin-colored rubber block.

[0168] In another optional embodiment, since the distribution of the noise signal has a certain randomness, while the target image signal is relatively fixed, in order to improve the quality of the target image, the image acquisition unit 800 also uses a multi-frame image fusion method to reduce the noise signal while maintaining the target image signal strength unchanged. For example, when the target fully presses the screen, multiple frames of target images are continuously acquired, and then the multiple frames of target images are weightedly fused as a whole or locally based on the overall quality or local quality of the multiple frames of target images to obtain a fused target image as the original target image. Compared with acquiring a single frame of target image as the original target image, the original target image obtained using the multi-frame image fusion method has less overall noise, providing a better foundation for subsequent target image processing steps.

[0169] The local brightness alignment unit 802 is configured to perform local brightness alignment processing on the first original image to obtain a first processed image.

[0170] Although the reflectivity of the simulated object is close to that of the target object, the reflectivity of the actual object is still not completely the same as that of the simulated object, resulting in uneven local brightness between the first original image and the second original image. In an optional embodiment, the local brightness alignment unit 802 includes:

[0171] The first pixel value calculation unit 8022 is used to calculate the pixel value r of each pixel in the first original image and the pixel average value of all pixels in the neighborhood window of each pixel. And the pixel value b of each pixel in the second original image and the pixel average of all pixels in the neighborhood window of each pixel The neighborhood window can be set according to actual conditions. For example, when the image processing method is specifically applied to process fingerprint images, an 11*11 window range centered on each pixel is selected as the neighborhood window.

[0172] The first processed image obtaining unit 8024 is used to obtain the pixel value r of each pixel in the first original image and the pixel average value of all pixels in the neighborhood window of each pixel. And the pixel average of all pixels in the neighborhood window of each pixel in the second original image A first processed image aligned with the local brightness of the second original image is obtained, wherein the pixel value r1 of each pixel in the first processed image can be obtained according to Formula 1:

[0173] Therefore, through the above-mentioned pixel value calculation unit 8022 and the first processed image acquisition unit 8024, local brightness alignment processing can be implemented on the first original image, so that the local brightness of the first processed image after local brightness alignment processing is consistent with that of the second original image.

[0174] The result acquisition unit 804 is configured to acquire a first result image based on the first processed image and the second original image.

[0175] In an optional embodiment, the result acquisition unit 804 may perform a subtraction operation on the first processed image and the second original image to obtain a first result image. For example, when the first original image is the original target image and the second original image is the background image, the first processed image is the original target image that has undergone local brightness alignment processing. The result acquisition unit 804 subtracts the second original image from the first processed image to obtain a first result image, which is the target image with the background removed. For another example, when the first original image is the background image and the second original image is the original target image, the first processed image is the background image that has undergone local brightness alignment processing. The result acquisition unit 804 subtracts the first processed image from the second original image to obtain a first result image, which is also the target image with the background removed.

[0176] It should be noted that, in the embodiment of the present application, performing a subtraction operation on the first processed image and the second original image means performing a subtraction operation on the pixel value of each pixel point of the first processed image and the pixel value of each pixel point of the second original image.

[0177] The image processing device provided according to the embodiment of the present invention can not only remove the main background texture, but also solve the problem of uneven local brightness between the background image and the original target image, so that the local brightness of the target image after removing the background is relatively uniform and clear. When specifically applied to processing fingerprint images, a clear fingerprint image with the background texture of the display screen removed can be obtained.

[0178] However, due to hardware imaging anomalies or the influence of the external environment, sometimes the overall brightness difference between the first original image and the second original image obtained by the image acquisition unit 800 is large, and the first result image obtained after the above steps S102-S104 still has obvious background texture residue. Therefore, the overall brightness alignment process can be performed before or after the local brightness alignment process. Figure 9 , provides a structural block diagram of another optional image processing device according to an embodiment of the present invention. Figure 9 As shown, the image processing device 90 includes:

[0179] The image acquisition unit 900 is configured to acquire a first original image and a second original image, wherein the first original image is one of an original target image and a background image, and the second original image is the other of the original target image and the background image.

[0180] The local brightness alignment unit 902 is configured to perform local brightness alignment processing on the first original image to obtain a first processed image.

[0181] The overall brightness alignment unit 904 is configured to perform overall brightness alignment processing on the first processed image to obtain a second processed image.

[0182] Result acquisition unit 906: obtains a second result image based on the second processed image and the second original image.

[0183] The above-mentioned image acquisition unit 900, local brightness alignment unit 902 and Figure 8 The image acquisition unit 800 and the local brightness alignment unit 802 in the described embodiment are the same, and can be found in detail. Figure 8 The corresponding description will not be described in detail here. Figure 9 The described embodiments and Figure 8 The difference is that the image processing device 90 also includes an overall brightness alignment unit 904, which is used to perform overall brightness alignment processing on the first processed image to obtain a second processed image, and the result acquisition unit 906 obtains a second result image based on the second processed image that has undergone overall brightness alignment processing and the second original image.

[0184] In an optional embodiment, the overall brightness alignment unit 904 includes:

[0185] A second pixel value calculation unit 9042 is used to calculate the maximum pixel value rmax and the minimum pixel value rmin of the first processed image, and the maximum pixel value bmax and the minimum pixel value bmin of the second original image;

[0186] The overall brightness coefficient calculation unit 9044 is configured to obtain an overall brightness scale coefficient α and an overall brightness offset coefficient β of the first processed image relative to the second original image based on the maximum pixel value rmax and the minimum pixel value rmin of the first processed image and the maximum pixel value bmax and the minimum pixel value bmin of the second original image. The overall brightness scale coefficient α can be calculated according to Formula 2: The overall brightness shift coefficient β can be calculated according to formula 3: β = (b min ·r max -b max ·r min ) / (r max -r min );

[0187] The linear transformation unit 9046 is used to perform a linear transformation on the first processed image based on the overall brightness scale coefficient α and the overall brightness offset coefficient β to obtain a second processed image that is aligned with the overall brightness of the second original image; wherein the pixel value r2 of each pixel point in the second processed image is obtained by linear transformation based on Formula 4, Formula 4: r2 = α·r + β.

[0188] In another optional embodiment, the second original image may also be subjected to overall brightness alignment processing. In this case, the image processing device 90 includes an image acquisition unit 900, a local brightness alignment unit 902, an overall brightness alignment unit 904' and a result acquisition unit 906'. Figure 8 The image acquisition unit 800 and the local brightness alignment unit 802 in the described embodiment are the same, and can be found in detail. Figure 8 The corresponding description will not be described in detail here.

[0189] Different from the overall brightness alignment unit 904, in this embodiment, the overall brightness alignment unit 904' is used to perform overall brightness alignment processing on the second original image to obtain a third processed image; the result acquisition unit 906' is used to obtain a third result image based on the first processed image and the third processed image.

[0190] In an optional embodiment, the overall brightness alignment unit 904′ includes:

[0191] A second pixel value calculation unit 9042 is used to calculate the maximum pixel value rmax and the minimum pixel value rmin of the first processed image, and the maximum pixel value bmax and the minimum pixel value bmin of the second original image;

[0192] The overall brightness coefficient calculation unit 9044′ is used to obtain the overall brightness scale coefficient α′ and the overall brightness offset coefficient β′ of the second original image relative to the first processed image based on the maximum pixel value rmax and the minimum pixel value rmin of the first processed image and the maximum pixel value bmax and the minimum pixel value bmin of the second original image. The overall brightness scale coefficient α′ can be calculated according to Formula 2. Formula 4: The overall brightness shift coefficient β′ can be calculated according to Formula 3, Formula 5: β′=(r min b max -r max b min ) / (b max -b min );

[0193] The linear transformation unit 9046′ is used to perform a linear transformation on the second original image based on the overall brightness scale coefficient α′ and the overall brightness offset coefficient β′ to obtain a third processed image whose overall brightness is aligned with that of the first processed image; wherein the pixel value b2 of each pixel point in the third processed image is obtained by linear transformation based on Formula 4, Formula 4: b2 = α′·b + β′.

[0194] It should be noted that in the above embodiment, local brightness alignment is performed first, followed by global brightness alignment. However, the above step sequence is merely an example and not a strict limitation. Those skilled in the art may make corresponding adjustments and modifications based on the principles of the above embodiment, for example, performing global brightness alignment first and then local brightness alignment.

[0195] Furthermore, to avoid the impact of local noise on parameter calculation, in embodiments including an overall brightness alignment step, the image processing apparatus may further include a smoothing unit for performing smoothing prior to the overall brightness alignment. For example, smoothing may be performed on the first original image and the second original image; or on the first processed image and the second original image, and so on. Persons skilled in the art may make reasonable adjustments based on specific embodiments. Specifically, the smoothing unit may employ methods such as mean filtering and Gaussian filtering for smoothing.

[0196] In addition, due to changes in the external environment and different pressing methods, the background texture contained in the original target image may be significantly different from the background image, that is, there is deformation of the background texture. For example, when the external temperature changes significantly, the internal structure of the display screen will undergo different degrees of thermal expansion and contraction effects. At this time, the imaging of the background texture in the original target image is significantly different from the background image. For example, when the target image is captured, due to the different forces of the target object pressing the display screen, the display screen will undergo different degrees of deformation, and the distance between it and the imaging sensor will also change, thereby causing the imaging of the background texture to change. In this case, the background texture may not be removed through the above embodiment.

[0197] To address the problem of background texture deformation, in an optional embodiment, an image of the target object captured when the target object has just touched the screen but has not yet fully pressed the display screen is used as the background image. When the target object just touches the display screen, the screen is illuminated, but the target object has not yet been fully pressed. The target image signal in the sensor imaging is weak, while the background texture is clear. Therefore, the target object image captured at this moment can be approximately used as the background image. By precisely controlling the image acquisition moment, a background image that is substantially consistent with the background texture in the current original target image can be obtained, thereby effectively addressing the problem of background texture deformation. However, this method has high requirements for the timing of image acquisition. If the image is acquired too early, the screen has not yet been illuminated or the target object is far away from the screen and the reflection is insufficient, resulting in a weak imaging intensity. If the image is acquired too late, the target object has already been fully pressed, and a pure background image cannot be obtained.

[0198] In order to solve the problem of background texture deformation without strict requirements on the timing of image acquisition, shape alignment can be performed before or after local brightness alignment. Figure 10, provides a structural block diagram of another optional image processing device according to an embodiment of the present invention. Figure 10 As shown, the image processing device 100 includes the following steps:

[0199] The image acquisition unit 1000 is configured to acquire a first original image and a second original image, wherein the first original image is one of an original target image and a background image, and the second original image is the other of the original target image and the background image.

[0200] The local brightness alignment unit 1002 is configured to perform local brightness alignment processing on the first original image to obtain a first processed image.

[0201] The shape alignment unit 1004 is configured to perform shape alignment processing on the first processed image to obtain a fourth processed image.

[0202] Result acquisition unit 1006: obtains a fourth result image based on the fourth processed image and the second original image.

[0203] The above-mentioned image acquisition unit 1000, local brightness alignment unit 1002 and Figure 8 The embodiments described are the same, for details, please refer to Figure 8 The corresponding description will not be described in detail here. Figure 10 The described embodiments and Figure 8 The difference is that the image processing device 100 further includes a shape alignment unit 1004 for performing shape alignment processing on the first processed image to obtain a fourth processed image, and the result acquisition unit 1006 obtains a fourth result image based on the fourth processed image after the shape alignment processing and the second original image.

[0204] In another optional embodiment, the image processing device 100 further includes a deformation judgment unit for judging whether the background image is deformed before the shape alignment unit 1004 performs the shape alignment processing step; optionally, judging whether the background image is deformed includes: judging the similarity between the first processed image and the second original image, specifically, for example, the zero-mean normalized cross-correlation coefficient ZNCC (zero-mean normalized cross-correlation) can be used for judgment; when the similarity is less than a first threshold, judging that the background image is deformed.

[0205] In an optional embodiment, the shape alignment unit 1004 includes:

[0206] a position calculation unit 10042, configured to calculate a position offset from each pixel in the first processed image to a corresponding target pixel in the second original image; optionally, the Lucas-Kanade algorithm may be used for calculation;

[0207] A fitting unit 10044 is configured to fit the displacement parameters (Δx, Δy) and scale parameter δ of the first processed image relative to the second original image based on the position offset; optionally, a least squares method may be used for fitting;

[0208] The alignment unit 10046 is configured to perform shape alignment processing on the first processed image according to the displacement parameter (Δx, Δy) and the scale parameter δ to obtain a fourth processed image that is shape-aligned with the second original image.

[0209] In another optional embodiment, the second original image may also be subjected to shape alignment processing. In this case, the image processing apparatus 100 includes an image acquisition unit 1000, a local brightness alignment unit 1002, a shape alignment unit 1004' and a result acquisition unit 1006'. Figure 8 The image acquisition unit 800 and the local brightness alignment unit 802 in the described embodiment are the same, and can be found in detail. Figure 8 The corresponding description will not be described in detail here.

[0210] Different from the shape alignment unit 1004, in this embodiment, the shape alignment unit 1004' is used to perform shape alignment processing on the second original image to obtain a fifth processed image; and the result acquisition unit 1006' is used to obtain a fifth result image based on the first processed image and the fifth processed image.

[0211] In an optional embodiment, the shape alignment unit 1004' includes:

[0212] a position calculation unit 10042, configured to calculate a position offset from each pixel in the first processed image to a corresponding target pixel in the second original image; optionally, the Lucas-Kanade algorithm may be used for calculation;

[0213] A fitting unit 10044' is configured to fit the displacement parameters (Δx', Δy') and scale parameter δ' between the second original image and the first processed image according to the position offset; optionally, a least squares method may be used for fitting;

[0214] The alignment unit 10046 ′ is configured to perform shape alignment processing on the second original image according to the displacement parameter (Δx′, Δy′) and the scale parameter δ′ to obtain a fifth processed image that is shape-aligned with the first processed image.

[0215] It should be noted that, in the above embodiment, local brightness alignment is performed first, and then shape alignment is performed. However, the above step sequence is merely an example and not a strict limitation. Those skilled in the art may make corresponding adjustments and transformations based on the principles of the above embodiment. For example, shape alignment may be performed first, and then local brightness alignment. In an image processing method including an overall brightness alignment step, the order of the three steps, overall brightness alignment, local brightness alignment, and shape alignment, may be arranged in different ways. For example, overall brightness alignment may be performed first, and then local brightness alignment, and finally shape alignment. For another example, local brightness alignment may be performed first, and then overall brightness alignment, and finally shape alignment. For another example, shape brightness alignment may be performed first, and then overall brightness alignment, and finally local brightness alignment.

[0216] In an optional embodiment, the image acquisition unit may further include acquiring a candidate background image, that is, sampling multiple frames of target images from the time the target just touches the screen to the time the target fully presses the screen as the candidate background image. Corresponding to this step, the image processing device further includes obtaining a result image after removing the candidate background based on the candidate background image and the first original image or the second original image that has at least undergone local brightness alignment processing; and selecting Figures 8-10 In the described embodiment, the better one of the result image after background removal and the result image after candidate background removal is used as the target image for background removal.

[0217] Due to the imaging characteristics of the sensor itself, the target image with the background removed may have local uneven quality, which is usually manifested as better quality in the central area and poor quality in the edge area. Therefore, it is necessary to enhance the quality of the target image with the background removed to improve the quality of the edge area. Compared with the central area, the poor quality of the edge area is mainly manifested as low contrast of the fingerprint ridges. To this end, the image processing device may include a local contrast enhancement unit for improving the local contrast, but the original noise will also be amplified at the same time. Therefore, the image processing device may also include a denoising unit for performing denoising after local contrast enhancement. Optionally, denoising methods such as Fast Non-Local Means Denoising and BM3D (Block-matching and 3D filtering) can be used. After local contrast enhancement and denoising, a clearer target image can be obtained.

[0218] Using the above-mentioned image processing device, a target image with background removed can be obtained. The image is clear overall, has uniform overall and local brightness, low noise, and no noticeable background texture residue. Furthermore, the device exhibits good adaptability to changes in the external environment. Specifically, when applied to fingerprint image processing, the device can produce a clear fingerprint image with the background texture of the display screen removed. The fingerprint features clear lines, uniform brightness, and no noticeable background texture residue. The device can overcome the effects of background deformation caused by changes in the external environment and finger pressure, and has no strict requirements on the timing of fingerprint collection, demonstrating good adaptability.

[0219] Use the above Figures 8-10 The image processing device described can obtain a clear target image with background removed. If it is used as a reference image, the background in the image can also be removed using a deep learning method. Figure 11 , is a structural block diagram of an optional image processing device based on deep learning according to an embodiment of the present invention. Figure 11 As shown, the image processing device 110 includes:

[0220] The image acquisition unit 1100 is configured to acquire a third original image and a fourth original image; wherein the third original image is one of an original target image and a background image, and the fourth original image is the other of the original target image and the background image;

[0221] In an optional embodiment, the third original image is an image of a target object captured when the target object is pressed against the display screen of an electronic device, and the background image is a texture image of the display screen itself captured when a smooth simulant with a reflectivity close to that of the target object is pressed against the display screen of the electronic device. When the image processing method is specifically applied to fingerprint images, the target object may be a finger, and the simulant may be a skin-colored rubber block.

[0222] An input image acquisition unit 1102 acquires an input image according to the third original image and the fourth original image;

[0223] In an optional embodiment, the input image acquisition unit 1102 may use the third original image and the fourth original image as multi-channel input images;

[0224] In another optional embodiment, the input image acquisition unit 1102 is used to perform a subtraction operation on the third original image and the fourth original image, and use the operation result as the input image; specifically, the fourth original image can be subtracted from the third original image, that is, the image with the background removed can be obtained as the input image.

[0225] In another optional embodiment, the input image acquisition unit 1102 is configured to perform local brightness alignment processing on the third original image and the fourth original image, and use the processing result as the input image; specifically, the local brightness alignment processing can be performed according to the following example: Figure 1 Step S104 described in is implemented.

[0226] The network construction unit 1104 is used to train the initial image generation network and construct a trained image generation network; wherein the image generation network is trained with the template image as a reference object, and the template image is used as a reference object. Figure 1-Figure 5 The target image with background removed is obtained by the image processing method described in the embodiment.

[0227] In an optional embodiment, the image generation network is a U-net network. Specifically, the image generation network may include three parts: a first part including two convolutional modules, a second part including four simple residual modules, and a third part including two deconvolutional modules. This structure ensures that the output image size of the image generation network remains consistent with the input image.

[0228] In an optional embodiment, the network construction unit 1104 includes:

[0229] A sample acquisition unit 11042 is configured to acquire a first sample image and a second sample image;

[0230] The training image acquisition unit 11044 is configured to input the first sample image and the second sample image into the initial image generation network to obtain a first training image;

[0231] The first training unit 11046 is configured to input the first training image and the template image into a first loss function module, train the initial image generation network based on the first loss value output by the first loss function module, and construct a trained image generation network. The first loss function module may be an L2-LOSS function.

[0232] In another optional embodiment, the network construction unit 1104 further includes:

[0233] The feature extraction unit 11048 is configured to input the first training image and the template image into a feature extraction network to obtain high-level semantic features of the first training image and the template image. In an optional embodiment, the feature extraction network may be a VGG network.

[0234] The second training unit 11050 is configured to input the high-level semantic features of the first training image and the high-level semantic features of the template image into a second loss function module, and train the image generation network and / or the feature extraction network based on a second loss value output by the second loss function module. The second loss function module may be an L2-LOSS function.

[0235] In this embodiment, by adding a feature extraction network to obtain high-level semantic features, and training the image generation network and / or feature extraction network through a second loss function, the final generated target image can be made closer to the template image.

[0236] The result acquisition unit 1106 is configured to input the input image into the trained image generation network to obtain a sixth result image.

[0237] According to the above embodiment, the background texture in the third original image or the fourth original image can be removed by a deep learning method in various environments without strict requirements on the quality of the third original image or the fourth original image.

[0238] Of course, those skilled in the art will know that Figure 11 The corresponding image processing device based on deep learning can also directly use high-quality clear images as reference objects for training instead of using Figure 1-Figure 5 The background-removed result image obtained by the image processing method described in the embodiment is used as a reference object for training.

[0239] Through the above Figures 8-11 The image processing apparatus described herein obtains a clear target image with the background removed, and thus a more accurate descriptor can be extracted from the target image with the background removed. Figure 12 , is a structural block diagram of an optional descriptor extraction device according to an embodiment of the present invention. Figure 12 As shown, the descriptor extraction device 120 includes:

[0240] The key point acquisition unit 1200 is used to acquire the key points of the target image, wherein the target image is a Figures 1-6 The target image with background removed obtained by the image processing method described in the embodiment;

[0241] In an optional embodiment, the SIFT algorithm is used to determine the key points of the target image.

[0242] The classification unit 1202 classifies multiple target images with the same key points into the same category and annotates the category information as an image label;

[0243] The screenshot unit 1204 is configured to capture image blocks from multiple target images having image labels representing the same category, with the key point as the center;

[0244] In an optional embodiment, when specifically applied to processing fingerprint images, an image block of size 32*32 may be captured with the key point as the center.

[0245] A descriptor extraction network construction unit 1206 is used to train the initial descriptor extraction network and construct a trained descriptor extraction network;

[0246] In an optional embodiment, the descriptor extraction network is trained according to a third loss value output by a third loss function module. The third loss function can be implemented by a variety of functions, such as tiplet loss, N-pairs loss, histogram loss, contrastive loss, circle loss, etc.

[0247] The feature descriptor acquisition unit 1208 inputs the image block into a trained descriptor extraction network to obtain a feature descriptor.

[0248] When specifically applied to processing fingerprint images, since fingerprints have too much repeated texture information, if the image blocks are directly input into a trained descriptor extraction network, the robustness of the descriptor extraction method will be reduced. To address this problem, the descriptor extraction device 120 also includes: a direction angle acquisition unit for acquiring the direction and angle of the key point; and a rotation and flattening unit for rotating and flattening the image block according to the direction and angle of the key point to obtain an aligned image block.

[0249] In an optional embodiment, the image generation network and the feature extraction network can be trained separately first, and then the descriptor extraction network can be trained. Finally, the image generation network, the feature extraction network and the descriptor extraction network can be trained using an end-to-end training method, and the network parameters can be fine-tuned to obtain a more robust feature descriptor.

[0250] The serial numbers of the above embodiments of the present invention are for description only and do not represent the advantages or disadvantages of the embodiments.

[0251] In the above embodiments of the present invention, the description of each embodiment has its own focus. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0252] In the several embodiments provided in this application, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are only exemplary. For example, the division of the units can be a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of units or modules, which can be electrical or other forms.

[0253] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple units. Some or all of the units may be selected according to actual needs to achieve the purpose of the present embodiment.

[0254] In addition, the functional units in the various embodiments of the present invention may be integrated into a single processing unit, each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.

[0255] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, server or network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: U disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), mobile hard disk, magnetic disk or optical disk, etc. Various media that can store program codes.

[0256] The above is only a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present invention. These improvements and modifications should also be regarded as within the scope of protection of the present invention.

Claims

1. An image processing method, comprising: Acquire a first original image and a second original image, wherein the first original image is one of an original target image and a background image, and the second original image is the other of the original target image and the background image; Performing local brightness alignment processing on the first original image to obtain a first processed image; Obtaining a target image with background removed based on the first processed image and the second original image; The step of performing local brightness alignment processing on the first original image to obtain a first processed image includes: Calculating the pixel value of each pixel in the first original image and the pixel average of all pixels in a neighborhood window of each pixel, and the pixel value of each pixel in the second original image and the pixel average of all pixels in a neighborhood window of each pixel; The first processed image aligned with the local brightness of the second original image is obtained based on the pixel value of each pixel in the first original image and the pixel average of all pixels in the neighborhood window of each pixel, as well as the pixel average of all pixels in the neighborhood window of each pixel in the second original image. The pixel value of each pixel in the first processed image is the pixel value of the pixel in the first original image minus the pixel average of all pixels in the neighborhood window of the pixel in the first original image plus the pixel average of all pixels in the neighborhood window of the pixel in the second original image.

2. The image processing method according to claim 1, wherein: The original target image is an image of the target object captured when the target object is pressed on the display screen surface of the electronic device, and the background image is a texture image of the display screen itself captured when a simulated object with a reflectivity close to that of the target object and a smooth surface is pressed on the display screen surface of the electronic device.

3. The image processing method according to claim 2, wherein: The target object is a finger, and the simulation object is a skin-colored rubber block.

4. The image processing method according to claim 2, wherein: When the target object completely presses the screen, multiple frames of target object images are continuously captured, and the multiple frames of target object images are weightedly fused as a whole or locally according to the overall quality or local quality of the multiple frames of target object images to obtain the fused target image as the original target image.

5. The image processing method according to claim 1, wherein: The obtaining of the first result image based on the first processed image and the second original image includes: performing a subtraction operation on the first processed image and the second original image to obtain the first result image. The image processing method according to claim 1 , further comprising performing an overall brightness alignment process before or after the local brightness alignment process.

7. The image processing method according to claim 6, characterized in that: After performing local brightness alignment processing on the first original image to obtain a first processed image, performing overall brightness alignment processing on the first processed image to obtain a second processed image; The step of performing overall brightness alignment processing on the first processed image to obtain a second processed image includes: respectively calculating a maximum pixel value and a minimum pixel value of the first processed image, and a maximum pixel value and a minimum pixel value of the second original image; Obtaining an overall brightness scale coefficient and an overall brightness offset coefficient of the first processed image relative to the second original image based on a maximum pixel value and a minimum pixel value of the first processed image and a maximum pixel value and a minimum pixel value of the second original image; The first processed image is linearly transformed based on the overall brightness scale coefficient and the overall brightness shift coefficient to obtain the second processed image whose overall brightness is aligned with that of the second original image.

8. The image processing method according to claim 6, wherein: Performing overall brightness alignment processing on the second original image to obtain a third processed image includes: respectively calculating a maximum pixel value and a minimum pixel value of the first processed image, and a maximum pixel value and a minimum pixel value of the second original image; Obtaining an overall brightness scale coefficient and an overall brightness offset coefficient of the second original image relative to the first processed image based on a maximum pixel value and a minimum pixel value of the first processed image and a maximum pixel value and a minimum pixel value of the second original image; The second original image is linearly transformed based on the overall brightness scale coefficient and the overall brightness shift coefficient to obtain the third processed image whose overall brightness is aligned with that of the first processed image. 9 . The image processing method according to claim 6 , further comprising performing a smoothing process before the overall brightness alignment.

10. The image processing method according to claim 9, wherein: The smoothing process includes at least one of the following: mean filtering and Gaussian filtering.

11. The image processing method according to claim 2, wherein: The target object image captured when the target object just touches the screen but does not completely press the display screen is used as the background image.

12. The image processing method according to claim 1, further comprising: The shape alignment process is performed before or after the local brightness alignment process.

13. The image processing method according to claim 12, wherein: After performing local brightness alignment processing on the first original image to obtain a first processed image, performing shape alignment processing on the first processed image to obtain a fourth processed image; The first processed image is subjected to shape alignment processing to obtain a fourth processed image, comprising: Calculating a position offset between each pixel in the first processed image and a corresponding target pixel in the second original image; Fitting a displacement parameter and a scale parameter of the first processed image relative to the second original image according to the position offset; Performing shape alignment processing on the first processed image according to the displacement parameter and the scale parameter to obtain the fourth processed image that is aligned in shape with the second original image.

14. The image processing method according to claim 12, wherein: Performing shape alignment processing on the second original image to obtain a fifth processed image includes: Calculating a position offset between each pixel in the first processed image and a corresponding target pixel in the second original image; fitting a displacement parameter and a scale parameter of the second original image relative to the first processed image according to the position offset; Shape alignment processing is performed on the second original image according to the displacement parameter and the scale parameter to obtain the fifth processed image that is aligned in shape with the first processed image.

15. The image processing method according to claim 1, further comprising: The target image with the background removed is subjected to at least one of the following processing: local contrast enhancement, fast non-local mean denoising, and three-dimensional block matching filtering.

16. The image processing method according to claim 12, further comprising: Before performing the shape alignment process, it is determined whether the background image is deformed.

17. The image processing method according to claim 1, further comprising: A candidate background image is collected, wherein the candidate background image is a plurality of frames of target object images sampled from the time when the target object just contacts the screen to the time when the target object completely presses the screen.

18. The image processing method according to claim 17, further comprising: A result image after removing the candidate background is obtained based on the candidate background image and the first original image or the second original image that has undergone local brightness alignment processing.

19. The image processing method according to claim 18, further comprising: The better one of the result image after background removal and the result image after candidate background removal is used as the target image for background removal.

20. An image processing method, comprising: Acquiring a third original image and a fourth original image, wherein the third original image is one of an original target image and a background image, and the fourth original image is the other of the original target image and the background image; Obtaining an input image according to the third original image and the fourth original image; Training an initial image generation network to construct a trained image generation network; wherein the image generation network is trained using a template image as a reference object, and the template image is a target image obtained by using the image processing method according to any one of claims 1 to 19 and removing the background; Inputting the input image into a trained image generation network to obtain the target image with the background removed; The step of obtaining the input image according to the third original image and the fourth original image includes: performing local brightness alignment processing on the third original image and the fourth original image, and using the processing result as the input image, including: respectively calculating the pixel value of each pixel in the third original image and the pixel average of all pixels in a neighborhood window of each pixel, and the pixel value of each pixel in the fourth original image and the pixel average of all pixels in a neighborhood window of each pixel; The input image aligned with the local brightness of the fourth original image is obtained based on the pixel value of each pixel in the third original image and the pixel average of all pixels in the neighborhood window of each pixel, as well as the pixel average of all pixels in the neighborhood window of each pixel in the fourth original image, wherein the pixel value of each pixel in the input image is the pixel value of the pixel in the third original image minus the pixel average of all pixels in the neighborhood window of the pixel in the third original image plus the pixel average of all pixels in the neighborhood window of the pixel in the fourth original image.

21. The image processing method according to claim 20, wherein: Training the initial image generation network and building a trained image generation network include: Acquire a first sample image and a second sample image; Inputting the first sample image and the second sample image into the initial image generation network to obtain a first training image; The first training image and the template image are input into a first loss function module, the initial image generation network is trained according to a first loss value output by the first loss function module, and the trained image generation network is constructed.

22. The image processing method according to claim 21, characterized in that: Training the initial image generation network and building a trained image generation network also includes: Inputting the first training image and the template image into a feature extraction network to obtain high-level semantic features of the first training image and high-level semantic features of the template image; The high-level semantic features of the first training image and the high-level semantic features of the template image are input into a second loss function module, and the image generation network and / or the feature extraction network are trained according to the second loss value output by the second loss function module.

23. A descriptor extraction method comprising: Acquire key points of a target image, wherein the target image is the target image with the background removed obtained using the image processing method according to any one of claims 1 to 19; Classifying multiple target images with the same key points into the same category, and marking the category information as image labels; Taking the key point as the center, extracting an image block from a plurality of the target images having the image label representing the same category; Train the initial descriptor extraction network to construct a trained descriptor extraction network; Inputting the aligned image blocks into the trained descriptor extraction network to obtain feature descriptors; Before inputting the aligned image blocks into the trained descriptor extraction network to obtain feature descriptors, the descriptor extraction method further includes: Obtaining the direction and angle of the key point; The image blocks are rotated and flattened according to the directions and angles of the key points to obtain aligned image blocks.

24. The descriptor extraction method according to claim 23, characterized in that: The descriptor extraction network is trained using an end-to-end training method.

25. An image processing device comprising: An image acquisition unit, configured to acquire a first original image and a second original image, wherein the first original image is one of an original target image and a background image, and the second original image is the other of the original target image and the background image; a local brightness alignment unit, configured to perform local brightness alignment processing on the first original image to obtain a first processed image; A result acquisition unit, configured to obtain a first result image based on the first processed image and the second original image; Wherein, the local brightness alignment unit includes: a first pixel value calculation unit, configured to respectively calculate a pixel value of each pixel point in the first original image and a pixel average value of all pixels points in a neighborhood window of each pixel point, and a pixel value of each pixel point in the second original image and a pixel average value of all pixels points in a neighborhood window of each pixel point; The first processed image obtaining unit is used to obtain the first processed image aligned with the local brightness of the second original image based on the pixel value of each pixel point in the first original image and the pixel average of all pixels points in the neighborhood window of each pixel point, as well as the pixel average of all pixels points in the neighborhood window of each pixel point in the second original image, wherein the pixel value of each pixel point in the first processed image is the pixel value of the pixel point in the first original image minus the pixel average of all pixels points in the neighborhood window of the pixel point in the first original image plus the pixel average of all pixels points in the neighborhood window of the pixel point in the second original image.

26. The image processing device according to claim 25, wherein The result acquisition unit obtains the first result image by performing a subtraction operation on the first processed image and the second original image. 27 . The image processing apparatus according to claim 25 , further comprising an overall brightness alignment unit configured to perform an overall brightness alignment process before or after the local brightness alignment process.

28. The image processing device according to claim 27, wherein: The overall brightness alignment unit includes: a second pixel value calculation unit, configured to respectively calculate a maximum pixel value and a minimum pixel value of the first processed image, and a maximum pixel value and a minimum pixel value of the second original image; an overall brightness coefficient calculation unit, which obtains an overall brightness scale coefficient and an overall brightness offset coefficient of the first processed image relative to the second original image based on the maximum pixel value and the minimum pixel value of the first processed image and the maximum pixel value and the minimum pixel value of the second original image; A linear transformation unit performs a linear transformation on the first processed image based on the overall brightness scale coefficient and the overall brightness offset coefficient to obtain a second processed image aligned with the overall brightness of the second original image.

29. The image processing device according to claim 27, wherein The overall brightness alignment unit includes: a second pixel value calculation unit, configured to respectively calculate a maximum pixel value and a minimum pixel value of the first processed image, and a maximum pixel value and a minimum pixel value of the second original image; an overall brightness coefficient calculation unit, which obtains an overall brightness scale coefficient and an overall brightness offset coefficient of the second original image relative to the first processed image based on the maximum pixel value and the minimum pixel value of the first processed image and the maximum pixel value and the minimum pixel value of the second original image; A linear transformation unit performs a linear transformation on the second original image based on the overall brightness scale coefficient and the overall brightness offset coefficient to obtain a third processed image whose overall brightness is aligned with that of the first processed image. 30 . The image processing apparatus according to claim 27 , further comprising a smoothing processing unit configured to perform a smoothing process before the overall brightness alignment. 31 . The image processing apparatus according to claim 25 , further comprising a shape alignment unit configured to perform shape alignment processing before or after the local brightness alignment processing.

32. The image processing device according to claim 31, wherein The shape alignment unit includes: a position calculation unit, configured to calculate a position offset between each pixel in the first processed image and a corresponding target pixel in the second original image; a fitting unit, configured to fit a displacement parameter and a scale parameter of the first processed image relative to the second original image according to the position offset; An alignment unit is configured to perform shape alignment processing on the first processed image according to the displacement parameter and the scale parameter to obtain a fourth processed image that is aligned in shape with the second original image.

33. The image processing device according to claim 31, wherein The shape alignment unit includes: a position calculation unit, configured to calculate a position offset between each pixel in the first processed image and a corresponding target pixel in the second original image; a fitting unit, configured to fit a displacement parameter and a scale parameter of the second original image relative to the first processed image according to the position offset; An alignment unit is configured to perform shape alignment processing on the second original image according to the displacement parameter and the scale parameter to obtain a fifth processed image that is shape-aligned with the first processed image.

34. An image processing apparatus, comprising: an image acquisition unit configured to acquire a third original image and a fourth original image, wherein the third original image is one of an original target image and a background image, and the fourth original image is the other of the original target image and the background image; an input image acquiring unit, configured to acquire an input image according to the third original image and the fourth original image; a network construction unit for training an initial image generation network to construct a trained image generation network; wherein the image generation network is trained using a template image as a reference object, the template image being a target image with background removed obtained using the image processing method according to any one of claims 1 to 22; A result acquisition unit, inputting the input image into the trained image generation network to obtain the target image with the background removed; The input image acquisition unit is configured to perform local brightness alignment processing on the third original image and the fourth original image, and use the processing result as the input image, including: Calculating the pixel value of each pixel in the third original image and the pixel average of all pixels in a neighborhood window of each pixel, and the pixel value of each pixel in the fourth original image and the pixel average of all pixels in a neighborhood window of each pixel; The input image aligned with the local brightness of the fourth original image is obtained based on the pixel value of each pixel in the third original image and the pixel average of all pixels in the neighborhood window of each pixel, as well as the pixel average of all pixels in the neighborhood window of each pixel in the fourth original image, wherein the pixel value of each pixel in the input image is the pixel value of the pixel in the third original image minus the pixel average of all pixels in the neighborhood window of the pixel in the third original image plus the pixel average of all pixels in the neighborhood window of the pixel in the fourth original image.

35. The image processing device according to claim 34, wherein: The network construction unit includes: a sample acquisition unit, configured to acquire a first sample image and a second sample image; a training image acquisition unit, configured to input the first sample image and the second sample image into the initial image generation network to obtain a first training image; The first training unit is used to input the first training image and the template image into a first loss function module, train the initial image generation network according to the first loss value output by the first loss function module, and construct the trained image generation network.

36. The image processing device according to claim 35, wherein: The network construction unit further includes: a feature extraction unit, configured to input the first training image and the template image into a feature extraction network to obtain high-level semantic features of the first training image and high-level semantic features of the template image; The second training unit is used to input the high-level semantic features of the first training image and the high-level semantic features of the template image into a second loss function module, and train the image generation network and / or the feature extraction network according to the second loss value output by the second loss function module.

37. A descriptor extraction device comprising: a key point acquisition unit, configured to acquire key points of a target image, wherein the target image is the target image with the background removed obtained by using the image processing method according to any one of claims 1 to 22; a classification unit, configured to classify a plurality of target images having the same key points into the same category and mark the category information as an image label; a screenshot unit, configured to capture an image block centered at the key point from a plurality of the target images having the image labels representing the same category; A descriptor extraction network construction unit, used to train the initial descriptor extraction network and construct a trained descriptor extraction network; A direction and angle obtaining unit, configured to obtain the direction and angle of the key point; a rotation and flattening unit, configured to rotate and flatten the image block according to the direction and angle of the key point to obtain an aligned image block; The feature descriptor acquisition unit is used to input the aligned image blocks into the trained descriptor extraction network to obtain feature descriptors.

38. The descriptor extraction device according to claim 37, characterized in that: The image generation network, the feature extraction network and the descriptor extraction network are trained using an end-to-end training method.

39. A storage medium, characterized in that The storage medium includes a stored program, wherein when the program is executed, the device where the storage medium is located is controlled to execute the image processing method according to any one of claims 1 to 22.

40. An electronic device, characterized in that: include: processor; as well as a memory for storing executable instructions of the processor; The processor is configured to perform the image processing method according to any one of claims 1 to 22 by executing the executable instructions.

Citation Information

Patent Citations

  • Projection-type image display device, image projection method, and computer program

    CN103888700A

  • Pedestrian feature extraction method and device, electronic equipment and storage medium

    CN111027455A

  • Fingerprint image processing method and device, electronic equipment and computer readable medium

    CN111028176A

  • Outdoor augmented reality application method based on cross-source image matching

    CN111260794A