A photographing apparatus and a photographing method
Patent Information
- Application Number
- CN202611066888.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-17
- Publication Date
- 2026-09-25
AI Technical Summary
[0006]本申请旨在解决现有自助拍照设备因高度差由单一机构承担,导致相机光轴与用户双眼不齐平而引起照片透视变形的问题
[0009]本申请实施例提供的拍摄设备及拍摄方法的至少一个优势是:通过视觉感知模块确定人脸中心并获取表征证件照构图目标位置的预设参考点,由控制模块根据人脸中心与预设参考点在竖直方向上的像素偏差确定实际高度差,进而驱动可升降摄像头与可升降拍照椅沿竖直方向协同相向运动,并使二者的高度调节量之和等于实际高度差,从而由摄像头与拍照椅共同分担该高度差,减小了单一机构的调节行程,同时使相机光轴与用户双眼趋于齐平,避免了因相机光轴相对双眼偏高或偏低所引起的照片透视变形,提升了证件照的成片质量。
Smart Images

Figure CN122824976A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of self-service photography technology, and in particular to a shooting device and shooting method. Background Technology
[0002] Passport photos have strict requirements for composition; the position and proportion of the face in the photo, as well as the relative height of the camera and the subject's eyes, must all meet the standards. To meet the growing demand in scenarios such as digital identity authentication, self-service photo booths are widely used in public security, transportation, medical, and educational institutions, allowing users to take their passport photos independently without supervision.
[0003] In self-service photo-taking scenarios, users vary significantly in height and sitting position, and the relative height between the camera and the user directly determines the vertical position of the face in the frame. If the camera is higher or lower than the user's face, the face will deviate from the required target position in the frame, making it difficult to obtain a photo that meets the composition standards for ID photos.
[0004] To ensure that the face is aligned with the target position in the image, existing self-service photo booths typically use a one-sided adjustment method, that is, only a height-adjustable photo chair or only a height-adjustable camera is provided. By adjusting the user's sitting height or the camera height, the face is moved to the required position in the image.
[0005] However, the above-mentioned unilateral adjustment method has at least the following problems: Since the relative height difference between the camera and the user's face needs to be completely eliminated by the photo chair or the camera alone, when there are large differences in the height and sitting height of different users, the adjustment stroke that a single mechanism needs to complete is large. This not only increases the adjustment time, but also, after the face is aligned with the target position of the image, the relative height between the camera and the user's face is often not at the same level. The optical axis of the camera is higher or lower than the user's eyes, resulting in perspective distortion of the photos taken from an upward or downward perspective, which affects the quality of the ID photo. Summary of the Invention
[0006] This application aims to solve the problem of perspective distortion in photos caused by existing self-service photo-taking equipment where the height difference is borne by a single institution, resulting in the camera's optical axis not being aligned with the user's eyes.
[0007] In a first aspect, this application provides a shooting device, including: A height-adjustable photo chair for supporting users; A retractable camera is used to capture images of the user. A visual perception module, connected to the liftable camera, is used to perform face detection on the image to determine the center of the face in the image and obtain a preset reference point; the preset reference point represents the target position of the face center required for the composition of the ID photo in the image; A control module, connected to the visual perception module and to the liftable camera and the liftable photo chair respectively, is used to determine the actual height difference based on the pixel deviation between the center of the face and the preset reference point in the vertical direction, and drive the liftable camera and the liftable photo chair to move in opposite directions in the vertical direction according to the actual height difference, so that the center of the face in the image approaches the preset reference point. Wherein, the sum of the height adjustment amount of the liftable camera and the height adjustment amount of the liftable photo chair is equal to the actual height difference.
[0008] Secondly, this application also provides a shooting method, including: Get the first image of the seated user; Face detection is performed on the first image to determine the face center in the first image, and a preset reference point is obtained; the preset reference point represents the target position of the face center required for the composition of the ID photo in the image; The actual height difference is determined based on the pixel deviation between the face center and the preset reference point in the vertical direction, and the liftable camera and the liftable photo chair are driven to move in opposite directions in the vertical direction according to the actual height difference, so that the face center in the first image approaches the preset reference point; wherein, the sum of the height adjustment amount of the liftable camera and the height adjustment amount of the liftable photo chair is equal to the actual height difference; A second image of the user after height adjustment is acquired. The head pose is estimated in the second image to obtain the Euler angles of the head pose. The Euler angles are fused with the pixel coordinates of facial key points in the second image to generate a pose guidance instruction. The pose guidance instruction is then output to guide the user to adjust their pose. Once the user's posture meets the requirements, the target image is acquired and output.
[0009] At least one advantage of the shooting device and shooting method provided in this application embodiment is that: the center of the face is determined by the visual perception module and a preset reference point representing the target position of the ID photo composition is obtained; the actual height difference is determined by the control module based on the pixel deviation between the center of the face and the preset reference point in the vertical direction; and then the liftable camera and the liftable photo chair are driven to move in opposite directions in the vertical direction in coordination, and the sum of the height adjustment of the two is equal to the actual height difference. Thus, the height difference is shared by the camera and the photo chair, reducing the adjustment stroke of a single mechanism. At the same time, the optical axis of the camera is made to be level with the user's eyes, avoiding the perspective distortion of the photo caused by the optical axis of the camera being too high or too low relative to the eyes, and improving the quality of the ID photo. Attached Figure Description
[0010] One or more embodiments are illustrated by way of example with reference numerals in the accompanying drawings. These illustrations do not constitute a limitation on the embodiments. Elements with the same reference numerals in the drawings are denoted as similar elements. Unless otherwise stated, the figures in the drawings are not to be limited by scale.
[0011] Figure 1 A schematic diagram of the structure of a shooting device provided for an embodiment of the present invention; Figure 2 A schematic diagram of another imaging device provided for an embodiment of the present invention; Figure 3 A schematic diagram illustrating the height adjustment principle of the shooting device provided for an embodiment of the present invention; Figure 4 A schematic flowchart of a shooting method provided for an embodiment of the present invention; Figure 5 A schematic diagram of the posture guidance process in the shooting method provided by the embodiments of the present invention; Figure 6 This is a schematic diagram illustrating the process of judging the quality and compliance of a target image in the shooting method provided in the embodiments of the present invention. Detailed Implementation
[0012] To facilitate understanding of this application, a more detailed description is provided below with reference to the accompanying drawings and specific embodiments. It should be noted that when an element is described as being "fixed to" another element, it can be directly on the other element, or one or more intermediate elements may exist between them. When an element is described as being "connected" to another element, it can be directly connected to the other element, or one or more intermediate elements may exist between them. The terms "upper," "lower," "inner," "outer," "bottom," etc., used in this specification indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings, and are only for the convenience of describing this application and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation, and therefore should not be construed as a limitation of this application. Furthermore, the terms "first," "second," "third," etc., are used for descriptive purposes only and should not be construed as indicating or implying relative importance.
[0013] Unless otherwise defined, all technical and scientific terms used in this specification have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used in this specification is for the purpose of describing particular embodiments only and is not intended to limit the scope of this application. The term "and / or" as used in this specification includes any and all combinations of one or more of the associated listed items.
[0014] Furthermore, the technical features involved in the different embodiments of this application described below can be combined with each other as long as they do not conflict with each other.
[0015] This application provides a shooting device 100, which can be a self-service photo taking device installed in places such as public security, transportation, medical care, and schools, for users to take ID photos by themselves when no one is on duty.
[0016] Referring to Figure 1, the shooting device 100 includes a height-adjustable shooting chair 10, a height-adjustable camera 20, a visual perception module 30, and a control module 40.
[0017] The height-adjustable photo chair 10 is used to support the user and can be raised and lowered in the vertical direction. When the height-adjustable photo chair 10 is raised and lowered, it moves the user sitting on it in the vertical direction, thereby changing the height of the user's face in the vertical direction.
[0018] The liftable camera 20 is used to acquire images of the user and can be raised and lowered in the vertical direction, thereby changing the height of the liftable camera 20 in the vertical direction.
[0019] In one embodiment, the height-adjustable photo chair 10 and the height-adjustable camera 20 can be raised and lowered by their respective motors via a transmission mechanism. The two motors are respectively connected to the control module 40, which controls the direction and stroke of their raising and lowering.
[0020] The visual perception module 30 is connected to the liftable camera 20 and is used to perform face detection on the image acquired by the liftable camera 20 to determine the center of the face in the image and obtain a preset reference point.
[0021] The face center refers to the central position of a face in an image. In one implementation, the visual perception module 30 performs face detection on the image to obtain a face detection box that outlines the face, and uses the center of the face detection box as the face center.
[0022] The preset reference point refers to the target position of the face center in the image, as required by the composition of an ID photo. When the face center coincides with the preset reference point, the user's face is in a position that meets the composition requirements of an ID photo in the image; when the face center deviates from the preset reference point, the user's face is in a position that does not meet the composition requirements of an ID photo in the image.
[0023] The control module 40 is connected to the visual perception module 30, and also to the liftable camera 20 and the liftable photo chair 10. The control module 40 is used to drive the liftable camera 20 and the liftable photo chair 10 to move vertically according to the face center determined by the visual perception module 30 and the preset reference point, so as to adjust the position of the user's face in the image and make the face center approach the preset reference point.
[0024] In one embodiment, the camera device 100 adopts a layered architecture. The camera device 100 includes an application layer, a business logic layer, an algorithm layer, a hardware control layer, and a data layer, which are arranged sequentially.
[0025] The application layer provides users with services such as user interface, payment, and photo output; the business logic layer implements functions such as visual perception, height adjustment control, and posture guidance interaction, and the visual perception module 30 and control module 40 can be set in the business logic layer; the algorithm layer provides the business logic layer with computing capabilities such as face detection, posture estimation, and quality assessment; the hardware control layer implements the control of hardware such as the height-adjustable camera 20, the height-adjustable photography chair 10, and supplementary lighting; and the data layer provides data storage and support for the shooting device 100.
[0026] In one embodiment, the imaging device 100 is further provided with a supplementary lighting system. The supplementary lighting system is used to provide supplementary lighting when the retractable camera 20 acquires images, and automatically adjusts the brightness of the supplementary lighting according to the intensity of the ambient light, so as to acquire images with appropriate brightness under different ambient light conditions.
[0027] Referring to Figure 2, in one embodiment, the imaging device 100 also includes a pressure sensor 50 and a database 60.
[0028] A pressure sensor 50 is installed on the height-adjustable photo chair 10 and connected to the visual perception module 30. The pressure sensor 50 is used to detect the pressure on the height-adjustable photo chair 10. On the one hand, it can determine whether the user is seated based on the detected pressure, and on the other hand, it can determine the user's weight based on the detected pressure.
[0029] Database 60 stores the correspondence between user features and composition parameters. In one embodiment, database 60 may include an image database storing images, a user feature database storing user features, and a reference point database storing composition parameters.
[0030] The specific method of obtaining the preset reference point by combining the pressure sensor 50 and the database 60 with the visual perception module 30 will be explained later.
[0031] The visual perception module 30 and the control module 40 can be implemented by the processor in the shooting device 100 or by discrete circuits; this application does not limit either approach.
[0032] Referring to Figure 3, after the user sits in the height-adjustable photo chair 10, the shooting device 100 detects that the user is seated and triggers the subsequent shooting process. In one embodiment, the pressure sensor 50 can determine whether the user is seated based on the detected pressure; when triggering the shooting process, at least one of the following methods can be used: touchscreen button triggering, voice command triggering, or automatic triggering after a preset duration.
[0033] The retractable camera 20 acquires images of the user. In one embodiment, before performing face detection on the image, the image can be preprocessed. The preprocessing includes at least one of noise reduction, color balancing, and exposure correction to provide a more consistent image for subsequent face detection.
[0034] The visual perception module 30 performs face detection on the image acquired by the liftable camera 20, determines the center of the face in the image, and obtains a preset reference point.
[0035] In one implementation, the face detection model used by the visual perception module 30 for face detection can be an improved RetinaFace model to detect the main faces in multi-person scenes. Furthermore, the face detection model can employ a lightweight network that has undergone pruning and quantization to run locally in real-time on the capturing device 100 without relying on an additional graphics processor.
[0036] In another embodiment, the face detection model can also use MobileNetV3 as the backbone network and combine it with an SSD detection head to detect faces in images. In one embodiment, the face detection model is pruned and quantized, and its model size is less than 5MB. It can run at a speed of more than 30 frames per second on the local central processing unit of the shooting device, thereby achieving real-time face detection without relying on a graphics processor.
[0037] For clarity, the following example is provided: Assume the image has its top-left corner as the origin, and the ordinate increases vertically downwards. A larger ordinate value indicates a lower position within the image. In this example, the ordinate of the preset reference point is 240 pixels; after the user is seated, the ordinate of the center of the face in the image acquired by the retractable camera 20 is 340 pixels.
[0038] The control module 40 determines the corresponding actual height difference based on the pixel deviation between the face center and the preset reference point in the vertical direction.
[0039] Pixel deviation refers to the difference in vertical coordinates between the face center and a preset reference point. Continuing with the previous example, the difference between the face center's vertical coordinate of 340 pixels and the preset reference point's vertical coordinate of 240 pixels is 100 pixels, meaning the pixel deviation is 100 pixels. Because the face center's vertical coordinate is greater than the preset reference point's vertical coordinate, the face center is located below the preset reference point in the image.
[0040] The actual height difference refers to the actual physical distance that the user's face needs to move vertically to move the center of the face in the image to a preset reference point. The control module 40 converts the pixel deviation into the actual height difference according to a pre-determined correspondence between image pixels and actual physical height. Continuing with the previous example, converting a pixel deviation of 100 pixels according to the above correspondence yields an actual height difference of 60 millimeters.
[0041] The control module 40, based on visual servoing, maps the pixel deviation between the center of the face in the image and a preset reference point to drive the liftable camera 20 and the liftable photography chair 10. Specifically, according to the actual height difference, the control module 40 drives the liftable camera 20 and the liftable photography chair 10 to move in opposite directions in the vertical direction, and makes the sum of the height adjustment of the liftable camera 20 and the height adjustment of the liftable photography chair 10 equal to the actual height difference, thereby making the center of the face in the image approach the preset reference point.
[0042] The height adjustment range refers to the distance that the height-adjustable camera 20 or the height-adjustable photo chair 10 moves in the vertical direction.
[0043] Coordinated moving towards each other refers to the simultaneous vertical movement of the liftable camera 20 and the liftable photo chair 10 towards each other.
[0044] Continuing with the previous example, the center of the face in the image is located below a preset reference point. To move the center of the face upwards to the preset reference point, the control module 40 drives the liftable camera 20 downwards and simultaneously drives the liftable photo chair 10 upwards, bringing the liftable camera 20 and the liftable photo chair 10 closer together. The sum of the height adjustment amounts of the liftable camera 20 and the liftable photo chair 10 equals the actual height difference of 60 mm. The upward movement of the liftable photo chair 10 raises the user's face accordingly, causing the center of the face to move upwards in the image, approaching the preset reference point.
[0045] In this way, the height difference between the user's face and the liftable camera 20 is shared by the liftable camera 20 and the liftable photo chair 10. Compared to methods where the liftable photo chair 10 or the liftable camera 20 alone bears the height difference, the adjustment range required for each of the liftable camera 20 and the liftable photo chair 10 is reduced.
[0046] In one embodiment, the height adjustment amount of the liftable camera 20 and the height adjustment amount of the liftable photo chair 10 are both equal to half of the actual height difference, so that the optical axis of the liftable camera 20 is collinear with the horizontal line where the user's eyes are located after the coordinated opposite movement.
[0047] The optical axis refers to the optical axis of the lens of the liftable camera 20. When the optical axis of the liftable camera 20 is collinear with the horizontal line where the user's eyes are located, the liftable camera 20 is directly facing the user's eyes, and the lens neither looks down nor up at the user.
[0048] Continuing with the previous example, the actual height difference is 60 mm. The height-adjustable camera 20 moves downwards by 30 mm, and the height-adjustable photo chair 10 moves upwards by 30 mm. The sum of their height adjustments is 60 mm, which equals the actual height difference. The height-adjustable photo chair 10 moves upwards by 30 mm, raising the user's eyes by 30 mm; the height-adjustable camera 20 moves downwards by 30 mm, lowering it by 30 mm. After the movements are complete, the height-adjustable camera 20 and the user's eyes are vertically close to each other, and the optical axis of the height-adjustable camera 20 is exactly collinear with the horizontal line where the user's eyes are located.
[0049] Therefore, while aligning the center of the face with the preset reference point, the optical axis of the liftable camera 20 is aligned with the user's eyes, avoiding perspective distortion caused by the optical axis of the liftable camera 20 being too high or too low relative to the user's eyes, thus improving the quality of the ID photo.
[0050] In one implementation, during the coordinated opposite-direction movement, the control module 40 controls the liftable camera 20 to continuously acquire images, determines the residual pixel deviation in real time based on the continuously acquired images, and iteratively drives the liftable camera 20 and the liftable photo chair 10 to move until the residual pixel deviation falls within a preset threshold range.
[0051] Residual pixel deviation refers to the difference in vertical coordinates between the center of the face in the currently acquired image and the preset reference point during the cooperative opposite motion process.
[0052] Specifically, each time the control module 40 drives the liftable camera 20 and the liftable photo chair 10 to move, the visual perception module 30 re-determines the face center in the currently acquired image, and the control module 40 re-determines the residual pixel deviation accordingly. When the residual pixel deviation does not fall within the preset threshold range, the control module 40 continues to drive the liftable camera 20 and the liftable photo chair 10 to move according to the residual pixel deviation. When the residual pixel deviation falls within the preset threshold range, the control module 40 stops driving.
[0053] Through the aforementioned real-time feedback correction method, even if there is a deviation between the actual height difference and the actual required adjustment amount, this deviation can be gradually eliminated through multiple iterations, making the center of the face accurately approach the preset reference point. Furthermore, the control module 40 can drive the liftable camera 20 and the liftable photography chair 10 to move with a small stroke in each iteration, ensuring smooth movement of the liftable camera 20 and the liftable photography chair 10.
[0054] In one implementation, the control module 40 determines the actual height difference based on the pixel deviation by multiplying the pixel deviation by a preset scaling factor to obtain the actual height difference.
[0055] The preset scaling factor refers to the actual physical height corresponding to a unit pixel in the vertical direction of the image, which is obtained by camera calibration of the liftable camera 20.
[0056] Continuing with the previous example, let the preset scaling factor be 0.6 mm / pixel. The control module 40 multiplies the pixel deviation of 100 pixels by the preset scaling factor of 0.6 mm / pixel to obtain an actual height difference of 60 mm, consistent with the previous example.
[0057] In one embodiment, the visual perception module 30 is further configured to determine the pixel width of the face detection box in the image; the control module 40 is further configured to correct a preset scaling factor based on the ratio of the pixel width to the width of the reference face detection box during calibration, and to determine the actual height difference using the corrected scaling factor.
[0058] The pixel width of the face detection bounding box refers to the width of the face detection bounding box obtained by the visual perception module 30 in the image. The baseline face detection bounding box width refers to the pixel width of the face detection bounding box when calibrating the liftable camera 20.
[0059] The closer the user's face is to the retractable camera 20, the larger the pixel width of the face detection box in the image. Therefore, the pixel width of the face detection box can reflect the distance between the user's face and the retractable camera 20. When the user leans forward or backward, causing this distance to deviate from the calibrated distance, the actual physical height corresponding to each pixel in the vertical direction of the image also changes. By correcting the preset scaling factor through the pixel width of the face detection box, the calculated actual height difference can be made more accurate.
[0060] In one implementation, the control module 40 multiplies a preset scaling factor by the ratio of the width of the reference face detection box to the pixel width of the current face detection box to obtain a corrected scaling factor, and then multiplies the pixel deviation by the corrected scaling factor to obtain the actual height difference.
[0061] Continuing with the previous example, let's assume the baseline face detection box width during calibration is 200 pixels. When the distance between the user's face and the liftable camera 20 is consistent with the calibration, the pixel width of the face detection box in the current image is also 200 pixels. The control module 40 multiplies the preset scaling factor of 0.6 mm / pixel by the ratio of the baseline face detection box width of 200 pixels to the current pixel width of 200 pixels, obtaining a corrected scaling factor that is still 0.6 mm / pixel. Then, the pixel deviation of 100 pixels is multiplied by the corrected scaling factor to obtain an actual height difference of 60 mm, consistent with the previous example.
[0062] In another scenario, when a user leans forward, bringing their face closer to the retractable camera 20, the pixel width of the face detection box in the current image increases. For example, if the current pixel width is 220 pixels, the control module 40 multiplies a preset scaling factor of 0.6 mm / pixel by the ratio of the baseline face detection box width of 200 pixels to the current pixel width of 220 pixels, resulting in a corrected scaling factor of approximately 0.545 mm / pixel. Then, the pixel deviation of 100 pixels is multiplied by the corrected scaling factor to obtain an actual height difference of approximately 54.5 mm. Thus, when the user leans forward and deviates from the calibrated distance, the calculated actual height difference is more accurate by correcting the preset scaling factor.
[0063] In one embodiment, referring to FIG2, the imaging device 100 further includes a pressure sensor 50 and a database 60. The visual perception module 30 obtains the preset reference point by acquiring the user's user characteristics and matching the preset reference point from the database 60 based on the user characteristics.
[0064] User characteristics refer to features used to characterize a user's body shape. User characteristics include at least one of the following: weight measured by pressure sensor 50, gender and age range estimated based on facial regions in an image, and sitting height range estimated based on the size and position of facial regions in an image.
[0065] Specifically, pressure sensor 50 detects the pressure applied to the height-adjustable photo chair 10 when the user sits down, thus obtaining the user's weight; visual perception module 30 estimates the user's gender and age group based on the facial region in the image; visual perception module 30 also estimates the user's sitting height based on the size and position of the facial region in the image. Database 60 stores the correspondence between user features and composition parameters. Based on the acquired user features, visual perception module 30 matches the composition parameters corresponding to the user features from database 60, and determines a preset reference point based on the matched composition parameters.
[0066] Therefore, for users with different body shapes, the visual perception module 30 can match a preset reference point that is suitable for them, enabling the shooting device 100 to provide a suitable composition for users with different body shapes. In one embodiment, the database 60 also updates the correspondence between user characteristics and composition parameters based on user characteristics and composition parameters accumulated from historical shooting, so as to gradually improve the fit between the matched preset reference points and users.
[0067] In one embodiment, before acquiring an image for determining the center of a face, the control module 40 first drives the liftable camera 20 to a preset initial height and acquires an initial image with a preset field of view. When the visual perception module 30 fails to detect a face from the initial image, the control module 40 drives at least one of the liftable camera 20 and the liftable photo chair 10 to move by a preset step size and repeatedly acquire the image until a face is detected.
[0068] The preset initial height refers to the height to which the liftable camera 20 is driven before acquiring an image used to determine the center of the face.
[0069] Because users have different heights and sitting heights, if the initial height of the height-adjustable camera 20 remains fixed, the faces of users whose height or sitting height deviates from the normal range may not fall into the image acquired by the height-adjustable camera 20, thus failing to detect their faces. By first driving the height-adjustable camera 20 to a preset initial height and acquiring an initial image with a preset field of view, the faces of users within a larger height range can fall into the initial image. When a face is not detected from the initial image, by driving at least one of the height-adjustable camera 20 and the height-adjustable photo chair 10 to move by a preset step size and repeatedly acquiring images, the face can be further captured in the image, thus ensuring that subsequent height adjustments can proceed normally.
[0070] In one embodiment, the control module 40 can roughly determine the user's body shape based on the weight measured by the pressure sensor 50, and determine a preset initial height accordingly, so that the initial image is easier to detect the face. In one embodiment, the user height range applicable to the shooting device 100 is 140 cm to 195 cm.
[0071] Based on the aforementioned shooting device 100, this application embodiment also provides a shooting method. This shooting method can be executed by the aforementioned shooting device 100, and the face detection, acquisition of preset reference points, determination of actual height difference, and coordinated opposite movement involved can all adopt the corresponding methods in the aforementioned embodiments, which will not be repeated below.
[0072] In one embodiment, the shooting device 100 is equipped with a process management machine for managing the shooting process. The process management machine has six states: standby, seating detection, initial positioning, height adjustment, posture guidance, and final shooting, and automatically transitions between these six states during the shooting process.
[0073] When no user is using the device, the process management unit is in standby mode. When a user is detected sitting down, the device switches from standby mode to seating detection mode and triggers the subsequent shooting process.
[0074] Referring to Figure 4, the imaging method provided in this embodiment of the application specifically includes the following steps: Step S101: Obtain the first image of the seated user.
[0075] The first image refers to the image acquired to determine the center of the face and adjust its height. In one embodiment, when the process management machine is in the initial positioning state, the liftable camera 20 is controlled to acquire the first image.
[0076] Step S102: Perform face detection on the first image to determine the face center in the first image and obtain a preset reference point.
[0077] The preset reference point represents the target position of the face center in the image required for the composition of the ID photo. The method of obtaining it is as described above and will not be repeated here.
[0078] In one embodiment, before acquiring the first image, it is determined whether a face is detected in the image; if no face is detected, at least one of the liftable camera 20 and the liftable photo chair 10 is driven to move by a preset step size and repeatedly acquire the image until a face is detected, and then the first image is acquired.
[0079] Step S103: Determine the corresponding actual height difference based on the pixel deviation between the face center and the preset reference point in the vertical direction.
[0080] Continuing with the previous example, the vertical coordinate of the face center is 340 pixels, the vertical coordinate of the preset reference point is 240 pixels, the pixel deviation is 100 pixels, and the actual height difference is calculated to be 60 millimeters.
[0081] In step S104, based on the actual height difference, drive the liftable camera 20 and the liftable photo chair 10 to move in opposite directions in the vertical direction so that the center of the face in the first image approaches the preset reference point.
[0082] The sum of the height adjustment range of the liftable camera 20 and the height adjustment range of the liftable photo chair 10 equals the actual height difference.
[0083] In one implementation, the process management machine is in a height adjustment state when performing step S104. Continuing with the previous example, the liftable camera 20 moves downward by 30 mm, and the liftable photo chair 10 moves upward by 30 mm. The sum of their height adjustments is 60 mm, which is equal to the actual height difference. After the movement is completed, the optical axis of the liftable camera 20 is collinear with the horizontal line where the user's eyes are located.
[0084] Step S105: Determine whether the residual pixel deviation meets the standard.
[0085] When the residual pixel deviation does not fall within the preset threshold range, return to step S104 and continue to drive the liftable camera 20 and liftable photo chair 10 to move; when the residual pixel deviation falls within the preset threshold range, proceed to step S106.
[0086] Step S106: Obtain the second image of the user after height adjustment, estimate the head pose of the second image, and obtain the Euler angles of the head pose.
[0087] The second image refers to the image obtained after height adjustment is completed, in order to estimate head pose and guide the user to adjust their pose.
[0088] Head pose estimation refers to determining the orientation of a user's head in three-dimensional space based on an image. Euler angles of head pose are the pitch, yaw, and roll angles that characterize the user's head orientation in three-dimensional space. The pitch angle represents the degree to which the head tilts up or down, the yaw angle represents the degree to which the head turns left or right, and the roll angle represents the degree to which the head tilts left or right.
[0089] In one implementation, the head pose estimation model used for head pose estimation can be a lightweight Hope-Net model to estimate head pose locally in real time on the imaging device 100.
[0090] Step S107: Merge Euler angles with the pixel coordinates of facial key points in the second image to generate pose guidance instructions, and output the pose guidance instructions to guide the user to adjust the pose.
[0091] Facial key points refer to the distinctive features of a user's face. In one implementation, facial key points include five key points: the pupils of both eyes, the tip of the nose, and the corners of the mouth.
[0092] A posture guidance instruction is an instruction used to guide the user to adjust their posture. In one implementation, the process management machine is in a posture guidance state when executing step S107.
[0093] Step S108: Determine whether the user's posture meets the standard.
[0094] If the user's posture does not meet the standard, return to step S107, continue to generate and output posture guidance instructions to guide the user to adjust the posture; when the user's posture meets the standard, proceed to step S109.
[0095] Step S109: Acquire the target image and output it.
[0096] The target image refers to the image acquired after the user's posture meets the requirements and is used as an ID photo output. In one implementation, the process management machine is in the final shooting state when executing step S109.
[0097] In one implementation, after outputting the target image, the process management unit switches from the final shooting state back to the standby state to wait for the next user.
[0098] It should be noted that the sequence numbers of the above steps are only used to distinguish different steps and do not constitute a restriction on the order of execution of the steps; the order of execution of the above steps can be adjusted when there is no conflict between them.
[0099] In some embodiments of this application, after performing head pose estimation on the second image to obtain Euler angles and determining the pixel coordinates of facial key points in the second image, referring to Figure 5, step S107 generates pose guidance instructions, which specifically includes the following steps: Step S201: Determine the deviation of the overall head orientation based on Euler angles.
[0100] The deviation of the overall head orientation refers to the deviation between the orientation of the user's head in three-dimensional space and the orientation when looking straight ahead. In one implementation, when at least one of the pitch angle, yaw angle, and roll angle exceeds its respective preset range, it is determined that there is a deviation in the overall head orientation; when the pitch angle, yaw angle, and roll angle all fall within their respective preset ranges, it is determined that the overall head orientation does not exceed the preset range.
[0101] Step S202: Determine the deviation of the gaze direction based on the pixel coordinates of the eye key points in the facial key points.
[0102] The deviation in eye gaze direction refers to the difference between the user's eye gaze direction and the eye gaze direction when looking straight ahead. In one implementation, the orientation of the user's eyes relative to the head is determined based on the pixel coordinates of key points in both eyes, thereby determining whether the eye gaze direction exceeds a preset range.
[0103] Step S203: Determine whether the deviation of the overall head orientation does not exceed the preset range, but the deviation of the eye orientation exceeds the preset range; if yes, proceed to step S204; if no, proceed to step S205.
[0104] Step S204: Generate posture guidance instructions to guide the user to adjust the direction of their gaze.
[0105] Step S205: Generate attitude guidance commands corresponding to the deviation.
[0106] Step S206: Output attitude guidance instructions to guide the user to adjust the attitude.
[0107] When judging posture solely based on Euler angles, if a user's head is generally facing forward but their gaze is deviating from it, the Euler angles cannot reflect this deviation because the head's orientation hasn't changed, thus failing to guide the user to adjust their gaze. However, the pixel coordinates of key points in both eyes can reflect the user's gaze direction.
[0108] Therefore, by fusing Euler angles with the pixel coordinates of key points of both eyes, when the deviation of the overall head orientation does not exceed the preset range, but the deviation of the eye orientation exceeds the preset range, that is, when the user's head is facing forward but the eyes are deviating from forward, it is still possible to generate a posture guidance instruction to guide the user to adjust the eye orientation, guide the user to adjust the eyes to look forward, and thus more accurately guide the user to complete the posture adjustment.
[0109] In one implementation, a multimodal guidance engine outputs posture guidance instructions. The multimodal guidance engine outputs posture guidance instructions in at least one of the following ways: speech, animation, and text. When outputting in speech mode, speech synthesis is used to play a voice message guiding the user to adjust their posture; when outputting in animation mode, an animation demonstrating how the user should adjust their posture is played; when outputting in text mode, text guiding the user to adjust their posture is displayed.
[0110] In one implementation, the multimodal bootstrapping engine supports bilingual bootstrapping in Mandarin and English to accommodate the needs of different users.
[0111] In one implementation, the posture guidance instruction is an instruction that guides the user to perform specific adjustment actions. For example, when the pitch angle indicates that the user is looking down and exceeds a preset range, a posture guidance instruction to guide the user to look up is generated and output; when it is determined in step S204 that the user's head is facing forward but their eyes are deviating from forward, a posture guidance instruction to guide the user to look forward is generated and output.
[0112] After generating and outputting the posture guidance instructions, the user adjusts the posture according to the posture guidance instructions, reacquires the image, and re-evaluates and judges the head posture until the user's posture meets the standard. Then, step S109 in block four is entered to acquire the target image and output it.
[0113] In some embodiments of this application, after acquiring the target image, it is also necessary to perform quality and compliance judgment on the target image, and to perform classification and rollback when the target image is unqualified.
[0114] Referring to Figure 6, the specific steps include the following: Step S301: The target image is judged for exposure, sharpness, closed eyes, facial occlusion, and composition compliance.
[0115] Quality and compliance assessment refers to the judgment made on the imaging quality of the target image and whether it conforms to the specifications for ID photos. In one implementation, quality and compliance assessment includes at least one of the following: exposure, sharpness, closed eyes, facial expression, facial occlusion, and composition compliance.
[0116] Among them, exposure is used to determine whether the brightness of the target image is appropriate, sharpness is used to determine whether the target image is clear, closed eyes is used to determine whether the user's eyes are closed in the target image, expression is used to determine whether the user's expression in the target image is appropriate, face occlusion is used to determine whether the user's face is occluded in the target image, and composition compliance is used to determine whether the composition of the target image conforms to the ID photo specifications.
[0117] In one implementation, the quality and compliance assessment also includes assessing the color balance of the target image.
[0118] Step S302: Determine whether the target image is qualified; if the target image is qualified, proceed to step S306 and output the target image; if the target image is unqualified, classify and backtrack according to the reason for the unqualification.
[0119] The following provides further explanation regarding category rollback.
[0120] Step S303: Determine if the target image is not in good pose; if so, return to step S107 in block four to generate pose guidance instructions and re-guide the user to adjust the pose; if not, proceed to step S304.
[0121] Step S304: Determine if the target image is unqualified due to composition or height. If yes, return to step S104 in block four to drive the liftable camera 20 and the liftable photo chair 10 to move in opposite directions and readjust the height. If no, proceed to step S305.
[0122] Step S305: Reacquire the target image.
[0123] Specifically, when the target image is unqualified due to posture, it indicates that the user's posture changed when the target image was acquired. Therefore, the process returns to step S107 to guide the user to adjust their posture. When the target image is unqualified due to composition or height, it indicates that the user's face position in the target image has deviated. Therefore, the process returns to step S104 to readjust the height. When the target image is unqualified due to closed eyes, inappropriate expression, or facial occlusion, it indicates that the user closed their eyes, made an inappropriate expression, or had their face occluded at the moment the target image was acquired, but their posture and height did not change. Therefore, only the target image needs to be acquired again, without needing to readjust the posture guidance or height.
[0124] By using the above-described classification and rollback method, the corresponding steps are returned according to the reason why the target image is unqualified. Compared with the method of returning to reacquire the image and re-adjust the height and posture guidance when the target image is unqualified, unnecessary repeated adjustments are reduced and the efficiency of acquiring qualified target images is improved.
[0125] Step S306: Output the target image.
[0126] In one implementation, outputting the target image includes at least one of photo printing, electronic download, and email sending.
[0127] In one embodiment, the imaging device 100 is further equipped with a payment module that supports at least one payment method, namely QR code payment and card payment. Payment is completed through the payment module during the output of the target image.
[0128] After outputting the target image, the imaging device 100 resets to standby mode to await the next user.
[0129] In the foregoing embodiments, the liftable camera 20 and the liftable photo chair 10 are each driven by a motor via a transmission mechanism to rise and fall. In another embodiment, the liftable camera 20 and the liftable photo chair 10 can also be driven by other drive mechanisms capable of vertical lifting, such as lead screw mechanisms, rack and pinion mechanisms, or hydraulic mechanisms, as long as they can enable the liftable camera 20 and the liftable photo chair 10 to rise and fall vertically. This application does not limit the specific form of the drive mechanism.
[0130] In the aforementioned embodiment, the example given was that the height adjustment range of the liftable camera 20 and the height adjustment range of the liftable photo chair 10 were both equal to half of the actual height difference. In another embodiment, the height adjustment range of the liftable camera 20 and the height adjustment range of the liftable photo chair 10 may not be equal, as long as their sum equals the actual height difference, the center of the face can be brought closer to a preset reference point. When the height adjustment range of the liftable camera 20 and the height adjustment range of the liftable photo chair 10 are both equal to half of the actual height difference, the optical axis of the liftable camera 20 is collinear with the horizontal line where the user's eyes are located after the movement is completed, which can avoid perspective distortion in the photo and result in better image quality.
[0131] In the foregoing embodiments, the center of the face detection bounding box is used as the face center. In another embodiment, other points that can characterize the position of the face in the image can also be used as the face center. This application does not limit the specific method for determining the face center.
[0132] In the foregoing embodiments, the pixel width of the face detection box is used as a measure of the distance between the user's face and the retractable camera 20. In another embodiment, other measures that can reflect this distance may be used to correct the preset scaling factor, and this application does not limit this to any particular measure.
[0133] In the foregoing embodiments, the shooting device 100 and shooting method were described using the shooting of ID photos as an example. In another embodiment, the shooting device 100 and shooting method provided in this application can also be applied to other shooting scenarios that require aligning a face with a specific position in an image; this application does not limit the specific application scenarios.
[0134] It should be noted that the technical features in the foregoing embodiments can be combined with each other as long as they do not conflict with each other. For example, the aforementioned implementation method of correcting the preset proportional coefficient based on the pixel width of the face detection box can be combined with the aforementioned implementation method of real-time feedback correction to further improve the accuracy of height adjustment.
[0135] In one embodiment, the functions of the aforementioned visual perception module 30 and control module 40 can be implemented by a processor executing a computer program stored in a memory. Accordingly, embodiments of this application also provide an imaging device, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the aforementioned imaging method.
[0136] In one embodiment, this application also provides a computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the aforementioned shooting method.
[0137] Those skilled in the art will understand that all or part of the aforementioned shooting method can be implemented using hardware related to computer program instructions. This computer program can be stored in a computer-readable storage medium, and when executed, it can include the aforementioned shooting method. The aforementioned storage medium can be a magnetic disk, optical disk, read-only memory, or random access memory, etc.
[0138] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application, and not to limit them; under the concept of this application, the technical features of the above embodiments or different embodiments can also be combined, the steps can be implemented in any order, and there are many other variations of different aspects of this application as described above, which are not provided in detail for the sake of brevity; although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or make equivalent substitutions for some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of this application.
Claims
1. A shooting device, characterized in that, include: A height-adjustable photo chair for supporting users; A retractable camera is used to capture images of the user. A visual perception module, connected to the liftable camera, is used to perform face detection on the image to determine the center of the face in the image and obtain a preset reference point; the preset reference point represents the target position of the face center required for the composition of the ID photo in the image; A control module, connected to the visual perception module and to the liftable camera and the liftable photo chair respectively, is used to determine the actual height difference based on the pixel deviation between the center of the face and the preset reference point in the vertical direction, and drive the liftable camera and the liftable photo chair to move in opposite directions in the vertical direction according to the actual height difference, so that the center of the face in the image approaches the preset reference point. Wherein, the sum of the height adjustment amount of the liftable camera and the height adjustment amount of the liftable photo chair is equal to the actual height difference.
2. The shooting device according to claim 1, characterized in that, The height adjustment range of the liftable camera and the height adjustment range of the liftable photo chair are both equal to half of the actual height difference, so that the optical axis of the liftable camera is collinear with the horizontal line where the user's eyes are located after the coordinated opposite movement.
3. The shooting device according to claim 1, characterized in that, The control module is also used to control the liftable camera to continuously acquire images during the coordinated opposite movement, determine the residual pixel deviation in real time based on the continuously acquired images, and iteratively drive the liftable camera and the liftable photo chair to move until the residual pixel deviation falls within a preset threshold range.
4. The shooting device according to claim 1, characterized in that, The control module determines the actual height difference based on the pixel deviation, including: The actual height difference is obtained by multiplying the pixel deviation by a preset scaling factor, which is obtained by calibrating the liftable camera.
5. The shooting device according to claim 4, characterized in that, The visual perception module is also used to determine the pixel width of the face detection box in the image; The control module is also used to correct the preset scaling factor based on the ratio of the pixel width to the width of the reference face detection box during calibration, and to determine the actual height difference based on the corrected scaling factor.
6. The shooting device according to claim 1, characterized in that, It also includes pressure sensors and a database; The pressure sensor is installed on the height-adjustable photo chair and connected to the visual perception module; the database stores the correspondence between user characteristics and composition parameters. The visual perception module acquires the preset reference point, including: Obtain the user's user characteristics, and match the preset reference point from the database based on the user characteristics; The user characteristics include at least one of the following: weight measured by the pressure sensor, gender and age range estimated based on the face region in the image, and sitting height range estimated based on the size and position of the face region in the image.
7. The shooting device according to claim 1, characterized in that, The control module is also used to drive the liftable camera to a preset initial height and acquire an initial image with a preset field of view before acquiring an image for determining the center of the face; When the visual perception module fails to detect a face from the initial image, the control module drives at least one of the liftable camera and the liftable photo chair to move by a preset step size and repeatedly acquire data until a face is detected.
8. A shooting method, characterized in that, include: Get the first image of the seated user; Face detection is performed on the first image to determine the face center in the first image, and a preset reference point is obtained; the preset reference point represents the target position of the face center required for the composition of the ID photo in the image; The actual height difference is determined based on the pixel deviation between the face center and the preset reference point in the vertical direction, and the liftable camera and the liftable photo chair are driven to move in opposite directions in the vertical direction according to the actual height difference, so that the face center in the first image approaches the preset reference point; wherein, the sum of the height adjustment amount of the liftable camera and the height adjustment amount of the liftable photo chair is equal to the actual height difference; A second image of the user after height adjustment is acquired. The head pose is estimated in the second image to obtain the Euler angles of the head pose. The Euler angles are fused with the pixel coordinates of facial key points in the second image to generate a pose guidance instruction. The pose guidance instruction is then output to guide the user to adjust their pose. Once the user's posture meets the requirements, the target image is acquired and output.
9. The method according to claim 8, characterized in that, The process of fusing the Euler angles with the pixel coordinates of the facial key points to generate pose guidance instructions includes: The deviation of the overall head orientation is determined based on the Euler angles, and the deviation of the gaze orientation is determined based on the pixel coordinates of the eye key points among the facial key points. When the deviation of the overall head orientation does not exceed the preset range, but the deviation of the eye orientation exceeds the preset range, a posture guidance instruction is generated to guide the user to adjust the eye orientation.
10. The method according to claim 8, characterized in that, Also includes: The target image is subjected to quality and compliance assessment, which includes at least one of exposure, sharpness, closed eyes, facial occlusion, and composition compliance. When the target image is unqualified, a classification and rollback are performed based on the reason for the unqualification, including: If the posture is unacceptable, return to the step of generating posture guidance instructions; if the composition or height is unacceptable, return to the step of driving the liftable camera and the liftable photography chair to move in opposite directions in coordination; if the target image is unacceptable due to closed eyes or face occlusion, reacquire the target image.