Image generation method, face recognition method, device, electronic device and medium
By combining the four-step phase shift method and Gaussian Laplace operator-guided filtering with the pinhole imaging model, the time-of-flight camera image generation algorithm is optimized, which solves the problem of insufficient image accuracy and achieves higher image generation accuracy.
Patent Information
- Application Number
- CN202210371370.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-04-08
- Publication Date
- 2025-09-12
- Estimated Expiration
- 2042-04-08
AI Technical Summary
The time-of-flight camera is not accurate enough in reconstructing images, and the image generation algorithm needs to be further optimized.
A four-step phase shift method is used to calculate the amplitude map and depth map, and the Laplacian of Gaussian operator is used for guided filtering. The pinhole imaging model is combined for correction to improve the image generation accuracy.
By adaptively adjusting the image region characteristics and preserving edge detail information, the accuracy of image generation is significantly improved.
Smart Images

Figure CN114782574B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of computer vision and can also be used in the financial field or other fields. More specifically, it relates to an image generation method, a face recognition method, an apparatus, a device, a medium, and a program product. Background Art
[0002] Three-dimensional vision technology is a key research focus, integrating machine vision with graphics processing. In recent years, with the rapid development of three-dimensional depth sensor technology, research in this field has transcended the conventional two-dimensional imaging paradigm, enabling analysis and interaction in three-dimensional space. Time-of-flight (ToF) cameras, one of the current mainstream three-dimensional perception technologies, offer advantages such as compact size, fast response time, simple algorithms, and high-frame-rate 3D image acquisition. Consequently, they are widely used in smart cars, robotics, security monitoring, and other fields. ToF cameras indirectly calculate the time of flight of light, and thus distance, by using the phase difference between the transmitted and received signals.
[0003] In the process of implementing the concept of the present disclosure, the inventors discovered that the accuracy of the current time-of-flight camera in reconstructing images is insufficient, and the image generation algorithm needs to be further optimized. Summary of the Invention
[0004] In view of the above problems, the present disclosure provides an image generation method, a face recognition method, an apparatus, a device, a medium, and a program product.
[0005] According to a first aspect of the present disclosure, a method for generating an image based on a time-of-flight camera is provided, comprising: acquiring raw data captured by the time-of-flight camera; calculating an amplitude map and a depth map based on the raw data using a four-step phase shift method; and using the amplitude map as a guide map and performing guided filtering on the depth map based on a Laplacian of Gaussian operator to obtain a target image.
[0006] According to an embodiment of the present disclosure, the step of performing guided filtering on the depth map based on the Laplacian of Gaussian operator to obtain a target image includes: constructing a weight factor of an adaptive regularization parameter based on the Laplacian of Gaussian operator; based on the weight factor, calculating a parameter expression of a guided filtering algorithm using the least squares method; based on the parameter expression of the guided filtering algorithm, obtaining a guided filtering algorithm optimized based on the Laplacian of Gaussian operator; and performing guided filtering on the depth map using the guided filtering algorithm optimized based on the Laplacian of Gaussian operator to obtain a target image.
[0007] According to an embodiment of the present disclosure, the amplitude graph includes a first amplitude graph and a second amplitude graph, the frequency of the first amplitude graph is smaller than that of the second amplitude graph, and the step of using the amplitude graph as a guide graph includes: using the second amplitude graph as a guide graph.
[0008] According to an embodiment of the present disclosure, before the step of using the second amplitude map as a guide map, the method further includes: performing joint bilateral filtering processing on the second amplitude map using the first amplitude map.
[0009] According to an embodiment of the present disclosure, before the step of performing guided filtering on the depth map based on the Laplacian of Gaussian operator, the method further includes: correcting the depth map using a pinhole imaging model.
[0010] A second aspect of the present disclosure provides a face recognition method for an automated teller machine, the automated teller machine including a time-of-flight camera, the method comprising: obtaining authorization from a user to enter a face; after obtaining authorization from the user to enter a face, using the time-of-flight camera of the automated teller machine to collect face information; based on the face information, generating a target face image using the above-mentioned method; comparing the target face image with a pre-stored standard face image; and determining that the face recognition result is passed when the comparison results are consistent.
[0011] A third aspect of the present disclosure provides a device for generating a target image based on a time-of-flight camera, comprising: a first acquisition module for acquiring raw data captured by the time-of-flight camera; a calculation module for calculating an amplitude map and a depth map based on the raw data using a four-step phase shift method; and a first image generation module for using the amplitude map as a guide map and performing guided filtering on the depth map based on a Laplacian of Gaussian operator to obtain a target image.
[0012] A fourth aspect of the present disclosure provides a face recognition device for an automated teller machine, comprising: a second acquisition module for obtaining a user's authorization to enter a face; a collection module for collecting face information using a time-of-flight camera of the automated teller machine after obtaining the user's authorization to enter a face; a second image generation module for generating a target face image based on the face information using any of the above methods; a comparison module for comparing the target face image with a pre-stored standard face image; and a result output module for determining that the face recognition result is passed when the comparison results are consistent.
[0013] The fifth aspect of the present disclosure provides an electronic device, comprising: one or more processors; and a memory for storing one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors execute the above method.
[0014] The sixth aspect of the present disclosure further provides a computer-readable storage medium having executable instructions stored thereon, which, when executed by a processor, causes the processor to execute the above method.
[0015] The seventh aspect of the present disclosure further provides a computer program product, comprising a computer program, which implements the above method when executed by a processor. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] The above contents and other objects, features and advantages of the present disclosure will become more apparent through the following description of the embodiments of the present disclosure with reference to the accompanying drawings, in which:
[0017] Figure 1 Schematically illustrates an application scenario diagram of the image generation method, apparatus, device, medium, and program product according to an embodiment of the present disclosure;
[0018] Figure 2 The flowchart of the image generation method according to the embodiment of the present disclosure is schematically shown;
[0019] Figure 3 Schematically illustrates imaging models between different coordinate systems in a pinhole imaging model according to an embodiment of the present disclosure;
[0020] Figure 4 Schematically illustrates the relationship between pixel coordinates and image coordinates in a pinhole imaging model according to an embodiment of the present disclosure;
[0021] Figure 5 The following schematically shows a flow chart of a face recognition method according to an embodiment of the present disclosure;
[0022] Figure 6 Schematically shows a structural block diagram of an image generating device according to an embodiment of the present disclosure;
[0023] Figure 7 Schematically shows a structural block diagram of a face recognition device according to an embodiment of the present disclosure; and
[0024] Figure 8 The block diagram schematically shows an electronic device suitable for implementing the image generation method and the face recognition method according to an embodiment of the present disclosure. DETAILED DESCRIPTION
[0025] Hereinafter, embodiments of the present disclosure will be described with reference to the accompanying drawings. However, it should be understood that these descriptions are merely exemplary and are not intended to limit the scope of the present disclosure. In the detailed description below, for ease of explanation, many specific details are set forth to provide a comprehensive understanding of the embodiments of the present disclosure. However, it is apparent that one or more embodiments may also be implemented without these specific details. In addition, in the following description, descriptions of well-known structures and technologies are omitted to avoid unnecessary confusion of the concepts of the present disclosure.
[0026] The terms used herein are only for describing specific embodiments and are not intended to limit the present disclosure. The terms "comprise," "include," etc. used herein indicate the presence of the features, steps, operations, and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, or components.
[0027] All terms used herein (including technical and scientific terms) have the meanings commonly understood by those skilled in the art unless otherwise defined. It should be noted that the terms used herein should be interpreted as having a meaning consistent with the context of this specification and should not be interpreted in an idealized or overly rigid manner.
[0028] When expressions such as "at least one of A, B and C, etc." are used, they should generally be interpreted in accordance with the meaning of the expression commonly understood by those skilled in the art (for example, "a system having at least one of A, B and C" should include but is not limited to a system having A alone, B alone, C alone, A and B, A and C, B and C, and / or A, B, C, etc.).
[0029] Three-dimensional vision technology is a key research hotspot that integrates machine vision and graphics processing. In recent years, with the rapid development of three-dimensional depth sensor technology, research in three-dimensional vision technology has broken through the conventional thinking of two-dimensional imaging and enabled analysis and interaction in three-dimensional space. As one of the current mainstream three-dimensional perception technologies, Time of Flight (ToF) cameras offer advantages such as compact structure, fast response time, simple algorithms, and high frame rate acquisition of three-dimensional images. Therefore, they are widely used in smart cars, robotics, security monitoring, and other fields. ToF cameras indirectly calculate the time of flight of light and, therefore, distance by using the phase difference between the transmitted and received signals. However, when reconstructing high-precision images, the measurement results and accuracy of ToF cameras are affected by many factors, including those within the camera system and the external environment. To obtain more accurate distance information, research on depth image optimization algorithms for ToF cameras is particularly important.
[0030] The sources of errors in time-of-flight cameras can be roughly divided into two categories. The first is the error caused by the camera hardware itself, such as pinhole imaging error, pixel response non-uniformity and odd harmonics, which are called systematic errors. Due to their high frequency of occurrence and relatively fixed form, the measured distance is generally corrected through the pinhole imaging model and the distance error model; the second is the error caused by the influence of uncertain factors such as the lighting, material, motion and color of the measured object, which is called non-systematic error. This type of error is random and non-fixed, and it is difficult to use a unified standard or model for error correction. The embodiment of the present disclosure proposes to further optimize the depth image through filtering methods to improve the accuracy of image generation.
[0031] Considering that time-of-flight cameras can simultaneously acquire depth maps and amplitude maps, the amplitude map directly reflects the amount of light received by the receiver from the transmitted modulated light signal, excluding ambient light and irrelevant stray light. It also indirectly reflects the reliability of the distance measurement; generally, larger amplitude values indicate higher reliability. Furthermore, guided filtering is an adaptive filtering method that uses a dynamic filter kernel to guide the input image based on a guide image. This method fully utilizes the overall features and edge details of both the filtered image and the guide image. Therefore, a guided filtering algorithm based on amplitude images is proposed to reduce the non-systematic errors of time-of-flight cameras.
[0032] Based on this, an embodiment of the present disclosure provides an image generation method based on a time-of-flight camera, including: obtaining raw data captured by the time-of-flight camera; based on the raw data, calculating an amplitude map and a depth map using a four-step phase shift method; and using the amplitude map as a guide map, performing guided filtering on the depth map based on a Gaussian Laplacian operator to obtain a target image.
[0033] In addition, an embodiment of the present disclosure also provides a face recognition method for an automatic teller machine, wherein the automatic teller machine includes a time-of-flight camera, and the method includes: obtaining a user's authorization to enter a face; after obtaining the user's authorization to enter a face, using the time-of-flight camera of the automatic teller machine to collect face information; based on the face information, generating a target face image using the above-mentioned method; comparing the target face image with a pre-stored standard face image; and when the comparison results are consistent, determining that the face recognition result is passed.
[0034] It should be noted that the method and apparatus determined in the present disclosure can be used for image generation in the financial field, and can also be used for image generation in any field other than the financial field. The application field of the method and apparatus for image generation disclosed in the present disclosure is not limited.
[0035] In the technical solution disclosed herein, the collection, storage, use, processing, transmission, provision, disclosure and application of user personal information involved comply with the provisions of relevant laws and regulations, take necessary confidentiality measures, and do not violate public order and good morals.
[0036] In the technical solution disclosed herein, the user's authorization or consent is obtained before obtaining or collecting the user's personal information.
[0037] Figure 1 The application scenario diagram of the image generation method, apparatus, device, medium and program product according to the embodiments of the present disclosure is schematically shown.
[0038] like Figure 1 As shown, the application scenario 100 according to this embodiment may include terminal devices 101, 102, 103, a network 104, and a server 105. The network 104 is used as a medium for providing a communication link between the terminal devices 101, 102, 103 and the server 105. The network 104 may include various connection types, such as wired or wireless communication links or optical fiber cables.
[0039] Users can use terminal devices 101, 102, and 103 to interact with server 105 via network 104 to receive or send messages, etc. Various communication client applications can be installed on terminal devices 101, 102, and 103, such as shopping applications, web browser applications, search applications, instant messaging tools, email clients, social platform software, etc. (only as examples).
[0040] The terminal devices 101 , 102 , and 103 may be various electronic devices having a display screen and supporting web browsing, including but not limited to smart phones, tablet computers, laptop computers, and desktop computers.
[0041] The server 105 may be a server that provides various services, such as a background management server (for example only) that supports websites browsed by users using the terminal devices 101, 102, and 103. The background management server may analyze and process received data such as user requests, and feed back processing results (such as web pages, information, or data obtained or generated according to user requests) to the terminal device.
[0042] It should be noted that the image generation method provided in the embodiments of the present disclosure can generally be executed by the server 105. Accordingly, the image generation apparatus provided in the embodiments of the present disclosure can generally be set in the server 105. The image generation method provided in the embodiments of the present disclosure can also be executed by a server or server cluster that is different from the server 105 and can communicate with the terminal devices 101, 102, 103 and / or the server 105. Accordingly, the image generation apparatus provided in the embodiments of the present disclosure can also be set in a server or server cluster that is different from the server 105 and can communicate with the terminal devices 101, 102, 103 and / or the server 105.
[0043] It should be understood that Figure 1 The number of terminal devices, networks and servers in the embodiment is merely illustrative. Any number of terminal devices, networks and servers may be provided as required.
[0044] The following will be based on Figure 1 The scene described by Figure 2 The image generation method of the disclosed embodiment is described in detail.
[0045] Figure 2 The flowchart of the image generating method according to the embodiment of the present disclosure is schematically shown.
[0046] like Figure 2 As shown, the image generating method of this embodiment includes operations S201 to S203.
[0047] In operation S201 , raw data captured by a time-of-flight camera is acquired.
[0048] According to an embodiment of the present disclosure, the raw data includes the phase difference between the emission signal and the reception signal of the time-of-flight camera. Based on this phase difference, the flight time of light can be indirectly calculated, and then the distance can be calculated, thereby reconstructing the image.
[0049] In operation S202 , an amplitude map and a depth map are calculated based on the original data using a four-step phase shift method.
[0050] The four-step phase shift method uses four sampling calculation windows for measurement. Each calculation window has a phase delay of 90° (0°, 90°, 180°, and 270°). The raw data collected by the receiving camera are Q0, Q1, Q2, and Q3 respectively. The depth value d is calculated according to the following formula.
[0051]
[0052]
[0053] Where: arctan(*) represents the inverse tangent function, modulation frequency f, and speed of light c.
[0054] The confidence calculation formula of the amplitude graph is:
[0055] confidence=abs(Q3-Q1)+abs(Q0-Q2)
[0056] Where: abs(*) represents the absolute value, Q0, Q1, Q2, and Q3 represent the raw data of the four phases collected by the camera.
[0057] According to an embodiment of the present disclosure, the amplitude map includes a first amplitude map and a second amplitude map, the frequency of the first amplitude map being lower than that of the second amplitude map, and the step of using the amplitude map as a guide map includes: using the second amplitude map as a guide map. A time-of-flight camera generally uses two modulation frequency modes for data acquisition. Different frequencies produce amplitude maps with different light intensities. Therefore, when acquiring data using a time-of-flight camera, images with different frequency amplitudes can be obtained. Low and high frequencies are determined by comparing the magnitudes of the two modulation frequency modes used by the time-of-flight camera. The amplitude map with the smaller frequency amplitude is the low-frequency amplitude map, while the amplitude map with the larger frequency amplitude is the high-frequency amplitude map. Because the frequency of the first amplitude map is lower than that of the second amplitude map, the first amplitude map is the low-frequency amplitude map, and the second amplitude map is the high-frequency amplitude map. For example, if the frequency amplitudes of two images captured are 40 and 60, respectively, the first amplitude map has a frequency amplitude of 40, i.e., a low-frequency amplitude map; the second amplitude map has a frequency amplitude of 60, i.e., a high-frequency amplitude map.
[0058] It's important to note that the amplitude image serves as a measure of confidence, and its quality directly determines the effectiveness of guided filtering on the depth map. Compared to low-frequency amplitude images, high-frequency amplitude images provide more distinct edge information and exhibit more pronounced variations in light intensity across different regions. Therefore, high-frequency amplitude images are chosen as the guiding image in the guided filtering algorithm.
[0059] According to an embodiment of the present disclosure, before the step of using the second amplitude map as a guide map, the method further includes: performing joint bilateral filtering processing on the second amplitude map using the first amplitude map.
[0060] It should be noted that although the high-frequency amplitude map can obtain more obvious edge information and the intensity of light in different areas varies more significantly, the low-frequency amplitude map can obtain stronger light signal intensity and can more accurately reflect the reliability of distance measurement in the low-frequency mode. Therefore, using the low-frequency amplitude map to perform joint bilateral filtering on the high-frequency amplitude map can further improve the reliability of image information.
[0061] In operation S203 , the amplitude map is used as a guide map, and guided filtering is performed on the depth map based on a Laplacian of Gaussian operator to obtain a target image.
[0062] According to an embodiment of the present disclosure, the step of performing guided filtering on the depth map based on the Laplacian of Gaussian operator to obtain a target image includes: constructing a weight factor of an adaptive regularization parameter based on the Laplacian of Gaussian operator; calculating a parameter expression of the guided filtering algorithm using the least squares method based on the weight factor; obtaining a guided filtering algorithm optimized based on the Laplacian of Gaussian operator based on the parameter expression of the guided filtering algorithm; and performing guided filtering on the depth map using the guided filtering algorithm optimized based on the Laplacian of Gaussian operator to obtain the target image. The guided filtering algorithm optimized based on the Laplacian of Gaussian operator can adapt to the characteristics of different image regions and preserve edge detail information, thereby improving the accuracy of the generated image.
[0063] In the classical guided filtering algorithm, there is the following relationship between the output image D and the guided image amp, where k represents the neighborhood window ω with a radius of r k Inner center pixel position:
[0064]
[0065] Where: a k , b k is the neighborhood window ω k An internal fixed constant; amp i is the pixel value of a pixel position i in the guide image within the neighborhood window.
[0066] The core of guided filtering lies in a k and b k The optimal solution is solved. The least squares method is used to calculate the difference between the input image I and the output image D to be the smallest. The minimum cost function is expressed as:
[0067] Where: ε is the regularization parameter, which is a constant used to prevent a k Too large to maintain data stability.
[0068] The cost function is about a k and b k The function can be obtained by k and b k Find the partial derivatives to find the optimal solution, that is:
[0069]
[0070] By setting the partial derivative in formula (3) to 0, we can obtain the optimal solution ak and b k :
[0071]
[0072] Where: |ω| is the number of elements in the neighborhood window, |ω|=(2r+1) 2 ; They represent the average value in the neighborhood window of the guide image amp and the average value in the neighborhood window of the input image I respectively; is the standard deviation of the neighborhood window of the guidance image amp.
[0073] The same pixel position i can be included in multiple windows. Because the center positions of multiple windows are different, the same pixel position a k and b k The values of are different, so the same pixel position i calculates a k and b k , it is necessary to calculate a in the neighborhood window with position i as the center k and b k Average value and So solve D i It can be expressed as:
[0074]
[0075] It can be seen from the above classic guided filtering algorithm that the use of a unified regularization parameter in the entire image processing process fails to adapt to the characteristics of different image regions, resulting in the loss of some edge detail information.
[0076] In the embodiments of the present disclosure, based on the idea of improving the guided filtering method through the adaptability of the regularization parameter, and taking into account the characteristics of the Laplacian operator that can significantly highlight the changes in different regions and is sensitive to noise, it is proposed to improve the guided filtering algorithm based on the idea of the Gaussian Laplacian operator, and the distance information and local variance of the pixel position in the neighborhood window are used as weight factors of the adaptive regularization parameter. The minimum cost function of this method changes from formula (2) to:
[0077]
[0078] Where: L amp (i) is the weight factor of the adaptive regularization parameter based on the Laplacian of Gaussian operator.
[0079] The Laplace operator is an operator that calculates the second-order partial derivative based on the Gaussian kernel function. It is a combination of the Gaussian function and the Laplace operator, and has the smoothing properties of the Gaussian function and the edge detection properties of the Laplace operator. Assuming that the two-dimensional Gaussian kernel function with a standard deviation of σ is expressed as:
[0080]
[0081] Solving the second-order partial derivative of the Gaussian kernel function, we can get the Gaussian Laplace convolution kernel, the formula is as follows:
[0082]
[0083] The distance information and local variance of the pixel position in the neighborhood window are referenced to the Gaussian Laplacian operator convolution kernel, which can be expressed as:
[0084]
[0085] Where: |dis| is the position (x, y) of other pixels i in the neighborhood window relative to the position k of the window center (x c ,y c ), is the variance of a point in the neighborhood window of the guidance image amp.
[0086] Based on the above Gaussian Laplace operator ΔG σ (i) Based on the ΔG between a certain pixel in the window and the center pixel σ (i) The weight factor of the adaptive regularization parameter is expressed as:
[0087]
[0088] Where: σ is the regularization factor, with a value of 0.1*max(amp(i))
[0089] Based on the weight factor, the parameter a of the guided filtering algorithm is obtained using the least squares method. k and b k The expression of the optimal solution is:
[0090]
[0091] The a that guides the filtering algorithm k and b k The expression of the optimal solution is brought into the above-mentioned classic guided filtering algorithm to obtain a guided filtering algorithm based on Gaussian Laplace operator optimization.
[0092] The depth map is guided filtered using the guided filtering algorithm optimized based on the Laplacian of Gaussian operator, and the input image I is convolved according to the optimized formula (5) to obtain a target image.
[0093] According to an embodiment of the present disclosure, before performing guided filtering on the depth map based on the Laplacian of Gaussian operator, the method further includes correcting the depth map using a pinhole imaging model. During time-of-flight camera imaging, pinhole imaging errors are caused by the camera hardware itself, and the pinhole imaging model is used to correct the measured distance.
[0094] Figure 3 The imaging model between different coordinate systems in the pinhole imaging model according to an embodiment of the present disclosure is schematically shown. Figure 4 The relationship between pixel coordinates and image coordinates in the pinhole imaging model according to an embodiment of the present disclosure is schematically shown.
[0095] The specific correction method of the pinhole imaging model is: modeling through the image imaging model between different coordinate systems, such as Figure 3 As shown, the camera coordinate system O c -X c Y c Z c , the image coordinate system o-xy, the coordinates of the point P in the world coordinate system in the camera coordinate system are P(X c , Y c , Z c ), the coordinates in the image coordinate system are p(x, y), and the camera focal length f is the distance between the camera coordinate origin and the image coordinate origin, that is, f = O c o. From this model, we can see that there are triangle similarity relationships such as ABO C and oCO C Similar, PBO C and pCO C Similar, so the measured distance and the corrected distance satisfy the following relationship:
[0096]
[0097] Where: O c P is the distance measured by the camera; P is the correction distance; P is the distance measured by the right triangle po c The hypotenuse distance can be calculated.
[0098] like Figure 4 As shown, the pixel coordinate system o uv In the relationship between -uv and the image coordinate system o-xy, the image coordinate origin is located at the exact center of the pixel coordinate, so it can be deduced:
[0099]
[0100] Where dx and dy represent the actual distance of each pixel in the x and y directions of the camera, respectively, in mm; u0 and v0 represent the coordinates of the center point of the pixel coordinate system in the ideal case, and cx and cy are actually obtained by camera calibration.
[0101] Therefore, the distance after imaging correction obtained based on the camera measurement distance is:
[0102]
[0103] The image generation method based on a time-of-flight camera provided by the embodiments of the present disclosure optimizes the classical guided filtering algorithm based on the Laplacian of Gaussian operator, can adapt to the characteristics of different areas of the image, preserve edge detail information, and thus improve the accuracy of the generated image.
[0104] Current ATMs rely on passwords for security verification. Manually entering passwords is not very secure, and leaks can easily lead to financial losses and financial risks. Therefore, a time-of-flight camera-based image generation method can be used to perform facial recognition on ATMs. This replaces manual password entry for verification, improving the security of ATM transactions.
[0105] An embodiment of the present disclosure provides a face recognition method for an automated teller machine (ATM), wherein the ATM includes a time-of-flight camera. The method includes: obtaining a user's authorization to enter a face; after obtaining the user's authorization to enter a face, using the ATM's time-of-flight camera to collect facial information; based on the facial information, generating a target facial image using the above-mentioned method; comparing the target facial image with a pre-stored standard facial image; and determining that the facial recognition result is passed when the comparison results are consistent.
[0106] Figure 5 The flowchart of the face recognition method according to the embodiment of the present disclosure is schematically shown.
[0107] like Figure 5 As shown, the face recognition method for an ATM in this embodiment includes operations S501 to S505.
[0108] In operation S501, the user's authorization for recording a face is obtained.
[0109] In operation S502, after obtaining authorization from the user to enter a face, facial information is collected using a time-of-flight camera of the ATM.
[0110] In an embodiment of the present disclosure, before obtaining the user's information, the user's consent or authorization may be obtained. For example, before operation S502, a request to obtain the user's information may be issued to the user. If the user agrees or authorizes the acquisition of the user's information, operation S502 is performed.
[0111] In operation S503 , based on the facial information, a target facial image is generated using the above-mentioned image generation method based on a time-of-flight camera.
[0112] In operation S504, the target facial image is compared with a pre-stored standard facial image.
[0113] In operation S505 , when the comparison results are consistent, the face recognition result is determined to be passed.
[0114] The face recognition method for ATMs provided by the embodiments of the present disclosure reconstructs high-precision three-dimensional face imaging by utilizing an image generation method based on a time-of-flight camera. The method is used for password verification at ATMs through face recognition, replacing the manual verification method of entering a bank card password, thereby greatly improving the security and efficiency of ATM business operations.
[0115] Based on the above-mentioned image generation method based on a time-of-flight camera, the present disclosure also provides an image generation device based on a time-of-flight camera. Figure 6 The device is described in detail.
[0116] Figure 6 The structural block diagram of the image generating device according to an embodiment of the present disclosure is schematically shown.
[0117] like Figure 6 As shown, the image generation device 600 based on a time-of-flight camera of this embodiment includes a first acquisition module 610 , a calculation module 620 and a first image generation module 630 .
[0118] The first acquisition module 610 is used to acquire the raw data captured by the time-of-flight camera. In one embodiment, the first acquisition module 610 can be used to perform the operation S201 described above, which will not be repeated here.
[0119] The calculation module 620 is configured to calculate the amplitude map and the depth map based on the raw data using a four-step phase shift method. In one embodiment, the calculation module 620 may be configured to perform the operation S202 described above, which will not be described in detail here.
[0120] The first image generation module 630 is configured to use the amplitude map as a guide map and perform guided filtering on the depth map based on the Laplacian of Gaussian operator to obtain a target image. In one embodiment, the first image generation module 630 may be configured to perform the operation S203 described above, which will not be described in detail here.
[0121] Based on the above-mentioned face recognition method for an ATM, the present disclosure also provides a face recognition device for an ATM. Figure 7 The device is described in detail.
[0122] Figure 7 The structural block diagram of the face recognition device according to an embodiment of the present disclosure is schematically shown.
[0123] like Figure 7 As shown, the face recognition device 700 for an ATM in this embodiment includes a second acquisition module 710 , a collection module 720 , a second image generation module 730 , a comparison module 740 and a result output module 750 .
[0124] The second acquisition module 710 is used to obtain the user's authorization to record the face. In one embodiment, the second acquisition module 710 can be used to perform the operation S501 described above, which will not be repeated here.
[0125] The acquisition module 720 is used to acquire facial information using the time-of-flight camera of the ATM after obtaining authorization from the user to enter the facial information. In one embodiment, the acquisition module 720 can be used to perform the operation S502 described above, which will not be repeated here.
[0126] In embodiments of the present disclosure, the user's consent or authorization may be obtained before obtaining the user's information. For example, before the acquisition module 720 acquires facial information, a request to acquire the user's information may be issued to the user. If the user agrees or authorizes the acquisition of the user's information, the acquisition module 720 acquires the facial information.
[0127] The second image generation module 730 is used to generate a target face image based on the face information using the target image generation method. In one embodiment, the second image generation module 730 can be used to perform the operation S503 described above, which will not be repeated here.
[0128] The comparison module 740 is used to compare the target face image with a pre-stored standard face image. In one embodiment, the comparison module 740 can be used to perform the operation S504 described above, which will not be described in detail here.
[0129] The result output module 750 is used to determine that the face recognition result is passed when the comparison results are consistent. In one embodiment, the result output module 750 can be used to perform the operation S505 described above, which will not be repeated here.
[0130] According to an embodiment of the present disclosure, the first acquisition module 610, the calculation module 620, the first image generation module 630, the second acquisition module 710, the acquisition module 720, the second image generation module 730, the comparison module 740, and the result output module 750 can be implemented in a single module, or any one of these modules can be split into multiple modules. Alternatively, at least part of the functionality of one or more of these modules can be combined with at least part of the functionality of other modules and implemented in a single module. According to an embodiment of the present disclosure, at least one of the first acquisition module 610, the calculation module 620, the first image generation module 630, the second acquisition module 710, the acquisition module 720, the second image generation module 730, the comparison module 740, and the result output module 750 can be at least partially implemented as a hardware circuit, such as a field programmable gate array (FPGA), a programmable logic array (PLA), a system on a chip, a system on a substrate, a system on a package, an application specific integrated circuit (ASIC), or can be implemented by hardware or firmware through any other reasonable means of circuit integration or packaging, or can be implemented in any one of the three implementation methods of software, hardware, and firmware, or in any appropriate combination of any of them. Alternatively, at least one of the first acquisition module 610, the calculation module 620, the first image generation module 630, the second acquisition module 710, the acquisition module 720, the second image generation module 730, the comparison module 740, and the result output module 750 can be at least partially implemented as a computer program module, which can perform the corresponding function when executed.
[0131] Figure 8 The block diagram schematically shows an electronic device suitable for implementing the image generation method and the face recognition method according to an embodiment of the present disclosure.
[0132] like Figure 8As shown, the electronic device 800 according to an embodiment of the present disclosure includes a processor 801, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 802 or a program loaded from a storage part 808 into a random access memory (RAM) 803. The processor 801 may include, for example, a general-purpose microprocessor (e.g., a CPU), an instruction set processor and / or a related chipset and / or a special-purpose microprocessor (e.g., an application-specific integrated circuit (ASIC)), etc. The processor 801 may also include an onboard memory for caching purposes. The processor 801 may include a single processing unit or multiple processing units for performing different actions of the method flow according to an embodiment of the present disclosure.
[0133] Various programs and data required for the operation of the electronic device 800 are stored in the RAM 803. The processor 801, ROM 802, and RAM 803 are connected to each other via a bus 804. The processor 801 executes the various operations of the method flow according to the embodiment of the present disclosure by executing the programs in the ROM 802 and / or RAM 803. It should be noted that the programs may also be stored in one or more memories other than the ROM 802 and RAM 803. The processor 801 may also execute the various operations of the method flow according to the embodiment of the present disclosure by executing the programs stored in the one or more memories.
[0134] According to an embodiment of the present disclosure, the electronic device 800 may further include an input / output (I / O) interface 805, which is also connected to the bus 804. The electronic device 800 may further include one or more of the following components connected to the I / O interface 805: an input portion 806 including a keyboard, a mouse, etc.; an output portion 807 including devices such as a cathode ray tube (CRT), a liquid crystal display (LCD), and a speaker; a storage portion 808 including a hard disk; and a communication portion 809 including a network interface card such as a LAN card or a modem. The communication portion 809 performs communication processing via a network such as the Internet. A drive 810 is also connected to the I / O interface 805 as needed. A removable medium 811, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed in the drive 810 as needed, so that a computer program read therefrom can be installed into the storage portion 808 as needed.
[0135] The present disclosure also provides a computer-readable storage medium, which may be included in the device / apparatus / system described in the above embodiments, or may exist independently and not be incorporated into the device / apparatus / system. The computer-readable storage medium carries one or more programs, and when executed, implements the method according to the embodiments of the present disclosure.
[0136] According to an embodiment of the present disclosure, a computer-readable storage medium may be a non-volatile computer-readable storage medium, for example, it may include but is not limited to: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In the present disclosure, a computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in combination with an instruction execution system, apparatus, or device. For example, according to an embodiment of the present disclosure, a computer-readable storage medium may include the ROM 802 and / or RAM 803 described above and / or one or more memories other than ROM 802 and RAM 803.
[0137] The embodiments of the present disclosure also include a computer program product, which includes a computer program containing program code for executing the method shown in the flowchart. When the computer program product is run in a computer system, the program code is used to cause the computer system to implement the method of the embodiments of the present disclosure.
[0138] The computer program executes the above functions defined in the system / device of the embodiment of the present disclosure when the computer program is executed by the processor 801. According to the embodiment of the present disclosure, the system, device, module, unit, etc. described above can be implemented by a computer program module.
[0139] In one embodiment, the computer program may be stored on a tangible storage medium such as an optical storage device or a magnetic storage device. In another embodiment, the computer program may be transmitted and distributed in the form of a signal on a network medium, downloaded and installed via the communication portion 809, and / or installed from a removable medium 811. The program code contained in the computer program may be transmitted using any appropriate network medium, including but not limited to wireless, wired, or any suitable combination thereof.
[0140] In such an embodiment, the computer program can be downloaded and installed from a network via the communication section 809, and / or installed from a removable medium 811. When the computer program is executed by the processor 801, the above-described functions defined in the system of the embodiment of the present disclosure are performed. According to the embodiment of the present disclosure, the systems, devices, means, modules, units, etc. described above can be implemented by computer program modules.
[0141] According to an embodiment of the present disclosure, the program code for executing the computer program provided by the embodiment of the present disclosure can be written in any combination of one or more programming languages. Specifically, these computer programs can be implemented using high-level procedural and / or object-oriented programming languages, and / or assembly / machine languages. Programming languages include, but are not limited to, languages such as Java, C++, python, "C" or similar programming languages. The program code can be executed entirely on the user computing device, partially on the user device, partially on a remote computing device, or entirely on a remote computing device or server. In cases involving a remote computing device, the remote computing device can be connected to the user computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computing device (for example, using an Internet service provider to connect via the Internet).
[0142] The flowcharts and block diagrams in the accompanying drawings illustrate the possible implementation architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present disclosure. In this regard, each box in the flowchart or block diagram can represent a module, program segment, or a part of code, and the above-mentioned module, program segment, or a part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in an order different from that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram or flowchart, and the combination of boxes in the block diagram or flowchart, can be implemented with a dedicated hardware-based system that performs the specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.
[0143] Those skilled in the art will appreciate that the features described in the various embodiments and / or claims of this disclosure may be combined and / or coupled in various ways, even if such combinations and / or couplings are not explicitly described in this disclosure. In particular, the features described in the various embodiments and / or claims of this disclosure may be combined and / or coupled in various ways without departing from the spirit and teachings of this disclosure. All such combinations and / or couplings are intended to fall within the scope of this disclosure.
[0144] The embodiments of the present disclosure are described above. However, these embodiments are for illustrative purposes only and are not intended to limit the scope of the present disclosure. Although each embodiment has been described separately above, this does not mean that the measures in each embodiment cannot be used in combination to advantage. The scope of the present disclosure is defined by the appended claims and their equivalents. Without departing from the scope of the present disclosure, those skilled in the art may make various substitutions and modifications, which should all fall within the scope of the present disclosure.
Claims
1. An image generation method based on a time-of-flight camera, characterized in that: include: Get the raw data captured by the time-of-flight camera; Based on the original data, calculating an amplitude map and a depth map using a four-step phase shift method, wherein the amplitude map includes a first amplitude map and a second amplitude map, and the frequency of the first amplitude map is smaller than that of the second amplitude map; and Using the amplitude map as a guide map, the depth map is guided filtered based on the Gaussian Laplacian operator to obtain a target image. The step of using the amplitude map as a guide map includes: performing joint bilateral filtering on the second amplitude map using the first amplitude map; The second amplitude map after joint bilateral filtering is used as the guide map. And wherein, the step of performing guided filtering on the depth map based on the Gaussian Laplacian operator to obtain the target image includes: Constructing a weight factor of an adaptive regularization parameter based on the Laplace operator of Gaussian, specifically including: improving a guided filtering algorithm based on the Laplace operator of Gaussian, and using distance information and local variance of pixel positions in a neighborhood window as a weight factor of the adaptive regularization parameter; Based on the weight factors, a parameter expression of the guided filtering algorithm is calculated using a least squares method; Based on the parameter expression of the guided filtering algorithm, a guided filtering algorithm based on Gaussian Laplace operator optimization is obtained; and The depth map is guided filtered using the guided filtering algorithm optimized based on the Laplacian of Gaussian operator to obtain a target image.
2. The method according to claim 1, characterized in that Before the step of performing guided filtering on the depth map based on the Laplacian of Gaussian operator, the method further includes: The depth map is corrected using a pinhole imaging model.
3. A face recognition method for an automatic teller machine, wherein the automatic teller machine includes a time-of-flight camera, characterized in that: The method comprises: Obtain the user's authorization to record the face; After obtaining the user's authorization to enter their face, the ATM's time-of-flight camera is used to collect facial information; Based on the facial information, generating a target facial image using the method of claim 1 or 2; Comparing the target face image with a pre-stored standard face image; and When the comparison results are consistent, the face recognition result is determined to be passed.
4. A device for generating a target image based on a time-of-flight camera, comprising: A first acquisition module is used to acquire raw data captured by the time-of-flight camera; a calculation module, configured to calculate an amplitude map and a depth map based on the raw data using a four-step phase shift method, wherein the amplitude map includes a first amplitude map and a second amplitude map, and the frequency of the first amplitude map is smaller than that of the second amplitude map; and A first image generation module is configured to use the amplitude map as a guide map and perform guided filtering on the depth map based on a Laplacian of Gaussian operator to obtain a target image. The step of using the amplitude map as a guide map includes: performing joint bilateral filtering on the second amplitude map using the first amplitude map; The second amplitude map after joint bilateral filtering is used as the guide map. And wherein, the step of performing guided filtering on the depth map based on the Gaussian Laplacian operator to obtain the target image includes: Constructing a weight factor of an adaptive regularization parameter based on the Laplace operator of Gaussian, specifically including: improving a guided filtering algorithm based on the Laplace operator of Gaussian, and using distance information and local variance of pixel positions in a neighborhood window as a weight factor of the adaptive regularization parameter; Based on the weight factors, a parameter expression of the guided filtering algorithm is calculated using a least squares method; Based on the parameter expression of the guided filtering algorithm, a guided filtering algorithm based on Gaussian Laplace operator optimization is obtained; and The depth map is guided filtered using the guided filtering algorithm optimized based on the Laplacian of Gaussian operator to obtain a target image.
5. A face recognition device for an automatic teller machine, comprising: The second acquisition module is used to obtain the user's authorization to enter the face; The acquisition module is used to collect facial information using the time-of-flight camera of the ATM after obtaining the user's authorization to enter the facial information; A second image generation module, configured to generate a target facial image based on the facial information using the method of claim 1 or 2; a comparison module, configured to compare the target face image with a pre-stored standard face image; and The result output module is used to determine that the face recognition result is passed when the comparison results are consistent.
6. An electronic device comprising: one or more processors; a storage device for storing one or more programs, When the one or more programs are executed by the one or more processors, the one or more processors are enabled to execute the method according to any one of claims 1 to 3.
7. A computer-readable storage medium having executable instructions stored thereon, which, when executed by a processor, causes the processor to perform the method according to any one of claims 1 to 3.
8. A computer program product, comprising a computer program, wherein when the computer program is executed by a processor, the method according to any one of claims 1 to 3 is implemented.
Citation Information
Patent Citations
Method and system for real-time motion artifact handling and noise removal for tof sensor images
CN107743638A
Artificial intelligence identification system and use method
CN112633896A
ITOF depth camera calibration and depth optimization method
CN113096189A