Display device, autostereoscopic display method, and eye positioning method
Through the combination of binocular distance measurement algorithm and face deviation state, positioning the spatial position of the eyes is solved, and the problems of inaccurate eye positioning and large jitter in the prior art are achieved, and a more efficient naked-eye 3D display is achieved.
Patent Information
- Application Number
- PCT/CN2023/135434
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-11-30
- Publication Date
- 2025-06-05
AI Technical Summary
Existing naked-eye 3D technology has problems of low accuracy and great jitter when positioning and tracking eyes, especially when faces deflect or are far away from the center of the camera.
A binocular distance measurement algorithm is used to combine the deviation state of the face relative to the center position of the binocular camera to locate the spatial position of the eyes, and the images are arranged based on this position to display the naked 3D picture.
Improve the accuracy and stability of eye positioning, reduce the problem of jitter when face deflection or distance from the center of the camera, and achieve better naked-eye 3D display effect.
Smart Images

Figure CN2023135434_05062025_PF_FP_ABST
Abstract
Description
Display device, naked-eye 3D display method, and eye positioning method Technical Field
[0001] The present disclosure relates to the field of machine vision technology, and in particular to a display device, a naked-eye 3D display method, and an eye positioning method. Background Art
[0002] It's well known that 3D images deliver a powerful visual feast. To experience the 3D effect, users can wear polarized glasses that use polarized light. This means the left eye only sees the image projected by the left camera, and the right eye only sees the image projected by the right camera, resulting in a three-dimensional image. Autostereoscopy is a general term for technologies that achieve stereoscopic visual effects without the aid of external tools such as polarized glasses. Autostereoscopy allows users to directly experience the 3D effect without the need for external tools. Autostereoscopy with eye tracking offers an unrestricted viewing angle and can adjust the display based on a user's location. This allows users to view the image while moving within a certain range. However, it requires accurate and stable eye positioning and tracking methods.
[0003] Eye tracking uses image processing technology to locate the pupil, obtain the coordinates of the pupil center, and calculate the eye's gaze point through a certain algorithm. In the current eye tracking solution using a monocular algorithm, the monocular algorithm sets a fixed pupil distance, but the pupil distance of each person's eyes is different, so the accuracy varies from person to person.
[0004] Summary of the Invention
[0005] The present disclosure provides a display device, a naked-eye 3D display method, and an eye positioning method, which are used to locate the spatial position of the eyes through a binocular ranging algorithm and the deviation state of the face relative to the center position of the binocular camera, arrange images based on the located spatial position, and display naked-eye 3D images.
[0006] In a first aspect, an embodiment of the present disclosure provides a display device, the display device including a display unit and a control circuit, wherein:
[0007] The display unit is configured to display a stereoscopic image based on parallax;
[0008] The control circuit includes a processor and a memory, wherein the memory is used to store a program executable by the processor, and the processor is used to read the program in the memory and perform the following steps:
[0009] Acquire an image pair, where the image pair includes a first view and a second view, the first view and the second view include a human face, and the image pair is captured by a binocular camera;
[0010] Determining the depth of the eyes of the human face using a binocular ranging algorithm based on the first view and the second view; wherein the depth of the eyes represents the distance from the eyes to the display unit;
[0011] Determining a deviation state of the face relative to a center position, wherein the center position is determined based on the center of a binocular camera capturing the face;
[0012] The spatial position of the eyes of the face is determined according to the depth and deviation state of the eyes, and a three-dimensional image is displayed based on the spatial position of the eyes.
[0013] In a second aspect, an embodiment of the present disclosure provides a display device, comprising: a display unit and a control circuit, wherein:
[0014] The display unit is configured to perform display of content;
[0015] The control circuit includes a processor and a memory, wherein the memory is used to store a program executable by the processor, and the processor is used to read the program in the memory and perform the following steps:
[0016] Obtaining an image pair, and determining whether a first view and a second view in the image pair contain the same face, wherein the image pair is captured by a binocular camera;
[0017] When it is determined that the first view and the second view contain the same human face, determining the depth of the eyes of the human face using a binocular ranging algorithm, wherein the depth of the eyes represents the distance from the eyes to the display unit;
[0018] The spatial position of the eyes of the human face is determined based on the depth of the eyes and the image pair.
[0019] In a third aspect, the present disclosure further provides a naked-eye 3D display method, the method comprising:
[0020] Acquire an image pair, where the image pair includes a first view and a second view, the first view and the second view include a human face, and the image pair is captured by a binocular camera;
[0021] Determining the depth of the eyes of the human face using a binocular ranging algorithm based on the first view and the second view; wherein the depth of the eyes represents the distance from the eyes to the display unit;
[0022] Determining a deviation state of the face relative to a center position, wherein the center position is determined based on the center of a binocular camera capturing the face;
[0023] The spatial position of the eyes of the face is determined according to the depth and deviation state of the eyes, and a three-dimensional image is displayed based on the spatial position of the eyes.
[0024] In a fourth aspect, an embodiment of the present disclosure further provides an eye positioning method, the method comprising:
[0025] Obtaining an image pair, and determining whether a first view and a second view in the image pair contain the same face, wherein the image pair is captured by a binocular camera;
[0026] When it is determined that the first view and the second view contain the same human face, determining the depth of the eyes of the human face using a binocular ranging algorithm, wherein the depth of the eyes represents the distance from the eyes to the display unit;
[0027] The spatial position of the eyes of the human face is determined based on the depth of the eyes and the image pair.
[0028] In a fifth aspect, an embodiment of the present disclosure further provides a naked-eye 3D display device, the device comprising:
[0029] An image acquisition module is used to acquire an image pair, wherein the image pair includes a first view and a second view, the first view and the second view include a human face, and the image pair is obtained by taking a binocular camera;
[0030] a depth determination module, configured to determine the depth of the eyes of the human face using a binocular ranging algorithm based on the first view and the second view; wherein the depth of the eyes represents the distance from the eyes to the display unit;
[0031] a deviation determination module, configured to determine a deviation state of a human face relative to a center position, wherein the center position is determined based on the center of a binocular camera that captures the human face;
[0032] The position determination module is used to determine the spatial position of the eyes of a face according to the depth and deviation state of the eyes, and to display a three-dimensional image based on the spatial position of the eyes.
[0033] In a sixth aspect, an embodiment of the present disclosure further provides an eye positioning device, the device comprising:
[0034] an image determination module, configured to obtain an image pair and determine whether a first view and a second view in the image pair contain the same face, the image pair being captured by a binocular camera;
[0035] a depth determination module, configured to determine the depth of the eyes of the face using a binocular ranging algorithm when it is determined that the first view and the second view contain the same face, wherein the depth of the eyes represents the distance from the eyes to the display unit;
[0036] A position determination module is used to determine the spatial position of the eyes of the face according to the depth of the eyes and the image pair.
[0037] In a seventh aspect, an embodiment of the present disclosure further provides an electronic device, comprising a processor and a memory, wherein the memory is configured to store a program executable by the processor, and the processor is configured to read the program in the memory and perform the following steps:
[0038] Acquire an image pair, where the image pair includes a first view and a second view, the first view and the second view include a human face, and the image pair is captured by a binocular camera;
[0039] Determining the depth of the eyes of the human face using a binocular ranging algorithm based on the first view and the second view; wherein the depth of the eyes represents the distance from the eyes to the display unit;
[0040] Determining a deviation state of the face relative to a center position, wherein the center position is determined based on the center of a binocular camera capturing the face;
[0041] Determine the spatial position of the eyes of a face based on their depth and deviation.
[0042] In an eighth aspect, an embodiment of the present disclosure further provides an electronic device comprising a processor and a memory, wherein the memory is used to store programs executable by the processor, and the processor is used to read the programs in the memory and execute the steps of the method described in any one of the third or fourth aspects above.
[0043] In a ninth aspect, an embodiment of the present disclosure further provides a computer storage medium on which a computer program is stored, which, when executed by a processor, is used to implement the steps of the method described in any one of the third or fourth aspects above.
[0044] In a tenth aspect, the present disclosure provides a computer program product, comprising: a computer program code, which, when executed on a computer, enables the computer to execute the method described in any one of the third aspect or the fourth aspect.
[0045] These and other aspects of the present disclosure will become more readily apparent from the following description of the embodiments. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] In order to more clearly illustrate the technical solutions in the embodiments of the present disclosure, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present disclosure. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0047] FIG1 is a schematic diagram of a display device provided by an embodiment of the present disclosure;
[0048] FIG2 is a schematic diagram of a binocular camera photographing a human face according to an embodiment of the present disclosure;
[0049] FIG3 is a parallax principle diagram of a binocular ranging algorithm provided by an embodiment of the present disclosure;
[0050] FIG4 is a schematic diagram of head posture estimation provided by an embodiment of the present disclosure;
[0051] FIG5 is a schematic diagram of a 2D to 3D conversion relationship provided by an embodiment of the present disclosure;
[0052] FIG6 is a schematic diagram of a binocular camera field of view estimation method according to an embodiment of the present disclosure;
[0053] FIG7 is a flowchart of eye positioning with reduced jitter provided by an embodiment of the present disclosure;
[0054] FIG8 is a flowchart of an implementation of an eye positioning tracking algorithm provided by an embodiment of the present disclosure;
[0055] FIG9 is a schematic diagram of a testing device provided in an embodiment of the present disclosure;
[0056] FIG10 is a schematic diagram of a spatial coordinate system provided by an embodiment of the present disclosure;
[0057] FIG11 is a schematic diagram of a display device provided by an embodiment of the present disclosure;
[0058] FIG12 is a flowchart of an implementation method of a naked-eye 3D display method provided by an embodiment of the present disclosure;
[0059] FIG13 is a flowchart of an eye positioning method according to an embodiment of the present disclosure;
[0060] FIG14 is a schematic diagram of an electronic device provided by an embodiment of the present disclosure;
[0061] FIG15 is a schematic diagram of a naked-eye 3D display device provided by an embodiment of the present disclosure;
[0062] FIG16 is a schematic diagram of an eye positioning device provided by an embodiment of the present disclosure;
[0063] FIG17 is a schematic diagram of an electronic device provided in an embodiment of the present disclosure. DETAILED DESCRIPTION
[0064] To make the objectives, technical solutions, and advantages of the present disclosure more clear, the present disclosure will be further described in detail below with reference to the accompanying drawings. It is apparent that the embodiments described are only a portion of the embodiments of the present disclosure, rather than all of them. All other embodiments derived by persons of ordinary skill in the art based on the embodiments of the present disclosure without creative effort are intended to fall within the scope of protection of the present disclosure.
[0065] In the embodiments of the present disclosure, the term "and / or" describes the association relationship between associated objects, indicating that three relationships can exist. For example, A and / or B can represent three situations: A exists alone, A and B exist simultaneously, and B exists alone. The character " / " generally indicates that the associated objects are in an "or" relationship.
[0066] The application scenarios described in the embodiments of the present disclosure are intended to more clearly illustrate the technical solutions of the embodiments of the present disclosure and do not constitute a limitation on the technical solutions provided by the embodiments of the present disclosure. Persons skilled in the art will appreciate that, as new application scenarios emerge, the technical solutions provided by the embodiments of the present disclosure will also be applicable to similar technical problems. In the description of the present disclosure, unless otherwise specified, "multiple" means two or more.
[0067] Before introducing the display device, naked-eye 3D display method, and eye positioning method provided by the embodiments of the present disclosure, for ease of understanding, the technical background of the embodiments of the present disclosure is first introduced in detail below.
[0068] It's well known that 3D images deliver a powerful visual feast. To experience the 3D effect, users can wear polarized glasses that use polarized light. This means the left eye only sees the image projected by the left camera, and the right eye only sees the image projected by the right camera, resulting in a three-dimensional image. Autostereoscopy is a general term for technologies that achieve stereoscopic visual effects without the aid of external tools such as polarized glasses. Autostereoscopy allows users to directly experience the 3D effect without the need for external tools. Autostereoscopy with eye tracking offers an unrestricted viewing angle and can adjust the display based on a user's location. This allows users to view the image while moving within a certain range. However, it requires accurate and stable eye positioning and tracking methods.
[0069] Eye tracking uses image processing technology to locate the pupil, obtain the coordinates of the pupil center, and calculate the eye's gaze point using an algorithm. Monocular visual positioning is based on camera imaging principles and image processing algorithms. Camera imaging involves focusing light from a scene onto an image sensor through an optical lens, forming a digital image. Image processing algorithms process digital images to extract feature information, such as edges, corners, and color, to locate and track the target object. However, monocular ranging algorithms cannot accurately measure the true interpupillary distance of the human eye. Current eye tracking solutions using monocular ranging algorithms use a fixed interpupillary distance, but this varies from person to person, resulting in individual accuracy. When a face is tilted or far from the camera center, the spatial coordinates of key points can be inaccurately calculated, resulting in significant jitter. Furthermore, binocular ranging algorithms can capture a wider range of facial movement when acquiring image pairs than monocular cameras, allowing for a wider viewing angle.
[0070] To address the above issues, the present disclosure provides a display device, a glasses-free 3D display method, and an eye positioning method. These utilize a binocular ranging algorithm to avoid the issues of relying on pupil distance and a small viewing angle when using only one camera. Furthermore, the spatial position of the eyes is determined by the deviation of the face from the center of the binocular camera, thereby resolving the issue of significant jitter caused by face deflection or distance from the camera center. The display device and method provided by the present disclosure optimize the eye positioning algorithm and, based on the spatial position of the located eyes, arrange images and display glasses-free 3D images.
[0071] Among them, the principle of binocular visual positioning is to observe the same object through two eyes at the same time. Since the positions of the two eyes are different, the images they see are also different. This difference is called parallax and is the basis of binocular visual positioning. When the brain receives different images from the two eyes, it will compare and analyze this information to determine the position and distance of the object. The present disclosure estimates the depth of the eyes based on parallax, and then determines the distance from the human eye to the display unit based on the estimated depth of the eyes, and then determines the spatial position of the human eye, arranges the image based on the different spatial positions of the human eye, and displays the arranged image on the display unit, thereby presenting a naked-eye 3D display effect.
[0072] It should be noted that the display device in this embodiment includes, but is not limited to, a glasses-free 3D display. It consists of four components: a 3D stereoscopic reality terminal, playback software, production software, and application technology. It is a cross-stereoscopic reality system that integrates modern high-tech technologies such as optics, photography, computers, automatic control, software, and 3D animation production. The glasses-free 3D display in this embodiment provides a display effect where objects in the image appear to be both prominent and hidden within the frame. Furthermore, the image is vibrantly colored, with distinct layers, vivid, and lifelike, creating a truly three-dimensional image.
[0073] The technical principles of glasses-free 3D displays include but are not limited to: light barrier technology, lenticular lens technology, and light field display lamps. Light barrier 3D technology utilizes a switchable LCD screen, a polarizing film, and a polymer liquid crystal layer. The liquid crystal layer and polarizing film create a series of vertical stripes oriented at 90°. These stripes are tens of microns wide, and light passing through them forms a vertical, fine-stripe pattern known as a "parallax barrier." This technology utilizes a parallax barrier placed between the backlight module and the LCD panel. In stereoscopic display mode, when the image intended for the left eye is displayed on the LCD screen, the opaque stripes block the view for the right eye. Similarly, when the image intended for the right eye is displayed on the LCD screen, the opaque stripes block the view for the left eye. By separating the left and right eye's visible images, the viewer perceives a 3D image. Lenticular lens technology, also known as micro-lenticular 3D technology, places the LCD screen's image plane at the focal plane of the lens. This allows each pixel in the image beneath each lenticular lens to be divided into several sub-pixels, allowing the lens to project each sub-pixel in a different direction. So when each eye looks at the display from different angles, it sees different sub-pixels. Because lenticular lens technology doesn't affect screen brightness like light barriers, it produces better display quality. Light field displays use dense fields of light to produce full-color, real-time 3D video, eliminating the need for glasses. This method of creating a 3D display allows several people to simultaneously view a virtual scene that resembles a real 3D object.
[0074] 1 , the display device in this embodiment is introduced.
[0075] The display device 100 in this embodiment includes a display unit 1040, a processor 1080 and a memory 1020, wherein the display unit 1040 includes a display panel 1041, which is used to display information input by the user or information provided to the user and various operation interfaces of the application, etc. In the embodiment of the present disclosure, it is mainly used to display the interface, shortcut window, three-dimensional menu model, menu information of menu items, etc. of the client installed in the display device 100.
[0076] Optionally, the display panel 1041 may be configured in the form of a liquid crystal display (LCD) or an organic light-emitting diode (OLED).
[0077] The processor 1080 is configured to read a computer program and then execute the method defined by the computer program. For example, the processor 1080 reads an application, thereby running the application on the display device 100 and displaying the application interface on the display unit 1040. The processor 1080 may include one or more general-purpose processors and may also include one or more DSPs (Digital Signal Processors) to perform related operations to implement the technical solutions provided by the embodiments of the present disclosure.
[0078] The memory 1020 generally includes internal memory and external memory. The internal memory can be a random access memory (RAM), a read-only memory (ROM), and a cache (CACHE), etc. The external memory can be a hard disk, an optical disk, a USB disk, a floppy disk or a tape drive, etc. The memory 1020 is used to store computer programs and other data. The computer program includes an application corresponding to the client, etc. Other data may include data generated after the operating system or application is run, and the data includes system data (such as configuration parameters of the operating system) and user data. In the embodiment of the present disclosure, program instructions are stored in the memory 1020, and the processor 1080 executes the program instructions in the memory 1020 to implement any one of the three-dimensional menu display methods provided by the present disclosure.
[0079] In addition, the display device 100 may further include a touch unit 1100 for receiving input digital information, word information, contact touch operations, or contactless gestures, and generating signal inputs related to user settings and function control of the display device 100. The touch unit 1100 includes, but is not limited to, an infrared touch unit, a capacitive touch unit, an electromagnetic touch unit, a camera acquisition unit, etc., wherein the camera acquisition unit is used to capture gestures of the user without touching the display screen. When the touch unit 1100 includes an infrared touch unit or an electromagnetic touch unit, the touch unit 1100 and the display unit 1040 may be stacked. For example, when a user performs a touch operation on the touch screen, the touch unit 1100 may collect the user's touch operation on or near it (such as an operation performed by the user using a finger, a stylus, or any other suitable object or accessory on the display panel 1041) and drive the corresponding connection device according to a pre-set program.
[0080] Optionally, the touch control unit 1100 may include two parts: a touch detection device and a touch processor. The touch detection device detects the user's touch direction, detects the signal caused by the touch operation, and transmits the signal to the touch processor; the touch processor receives the touch information from the touch detection device, converts it into touch point coordinates, and then sends it to the processor 1080. It can also receive commands sent by the processor 1080 and execute them. In the embodiment of the present disclosure, if the user clicks on the application, the touch detection device in the touch control unit 1100 detects a touch operation, and sends the signal corresponding to the detected touch operation to the touch processor. The touch processor converts the signal into touch point coordinates and sends them to the processor 1080. The processor 1080 determines the operation that the user needs to perform based on the received touch point coordinates.
[0081] The display panel 1041 can be implemented using various types such as resistive, capacitive, infrared, and surface acoustic wave. In addition to the display unit 1040 and the touch unit 1100, the display device 100 can also include an input unit 1030. The input unit 1030 can include an image input device 1031 and other input devices 1032. The other input devices 1032 can be, but are not limited to, one or more of a physical keyboard, function keys (such as volume control keys, power keys, etc.), a trackball, a mouse, a joystick, and the like.
[0082] In addition to the above, the display device 100 may also include a power supply 1090 for powering other modules, an audio circuit 1060, a near-field communication module 1070, and an RF circuit 1010. The display device 100 may also include one or more sensors 1050, such as an accelerometer, a light sensor, a pressure sensor, etc. The audio circuit 1060 specifically includes a speaker 1061 and a microphone 1062. For example, the display device 100 can collect the user's voice through the microphone 1062 to perform corresponding operations.
[0083] As an embodiment, the number of the processors 1080 may be one or more, and the processor 1080 and the memory 1020 may be coupled or relatively independently configured.
[0084] As an embodiment, the processor 1080 is configured to perform the following steps:
[0085] An image pair is obtained, the image pair including a first view and a second view, the first view and the second view including a human face, the image pair being captured by a binocular camera; a binocular ranging algorithm is used to determine the depth of the eyes of the human face based on the first view and the second view, wherein the depth of the eyes represents the distance from the eyes to a display unit; a deviation state of the human face relative to a center position is determined, wherein the center position is determined based on the center of the binocular camera capturing the face; a spatial position of the eyes of the human face is determined based on the depth and deviation state of the eyes, and a stereoscopic image is displayed based on the spatial position of the eyes.
[0086] The specific execution process of the processor in this embodiment is described below:
[0087] Step 1: Acquire an image pair, wherein the image pair includes a first view and a second view, wherein the first view and the second view include a human face, and the image pair is captured by a binocular camera; and determine the depth of the eyes of the human face using a binocular ranging algorithm based on the first view and the second view;
[0088] During implementation, it can be determined whether the first view and the second view in the image pair contain the same face. When the first view and the second view in the image pair contain the same face, a binocular ranging algorithm can be used to determine the depth of the eyes of the face. The facial images in the first view and the second view have different content, for example, the first view and the second view have different presentation angles of the face, and the face is captured from different angles. The depth of the eyes represents the distance from the eyes to the display unit.
[0089] During implementation, an image pair is first captured by the binocular camera on the display device. The captured image pair includes two views, namely a first view and a second view. Optionally, the first view is obtained by photographing a face with the first camera of the binocular camera, and the second view is obtained by photographing the face with the second camera of the binocular camera. The first camera and the second camera of the binocular camera are located on the same horizontal axis. As shown in Figure 2, this embodiment provides a schematic diagram of a binocular camera photographing a face. The first camera and the second camera of the binocular camera simultaneously photograph the face, obtaining a first view corresponding to the first camera and a second view corresponding to the second camera. The binocular camera then synchronously transmits the captured first view and second view to the display device, thereby obtaining an image pair consisting of the first view and the second view.
[0090] During implementation, the binocular camera in this embodiment is configured on the display device and faces the user who is watching the image displayed on the display unit, so as to locate the spatial position of the user's eyes by capturing the user's face, and arrange the image based on the spatial position, thereby achieving a naked-eye 3D display effect.
[0091] It should be noted that the premise of using the binocular ranging algorithm to calculate the depth of the human eye is that the same face must be detected in both views captured by the binocular camera. Therefore, after obtaining the image pair, it is also necessary to determine whether the first view and the second view contained in the image pair contain the same face. When it is determined that the first view and the second view in the image pair contain the same face, the binocular ranging algorithm can be used to determine the depth of the face's eyes.
[0092] In some embodiments, the depth of the eyes of a face can be determined using a binocular ranging algorithm by the following steps:
[0093] A binocular ranging algorithm is used to determine the depth of the eyes based on parallax, where the parallax represents the horizontal position deviation of the eyes of the same face in a first view and a second view.
[0094] It should be noted that the origin positions of the coordinate systems in the first view and the second view are the same, for example, both are the upper left corner of the view, and the directions of the X-axis and Y-axis are also the same, for example, the X-axis of the coordinate systems in the first view and the second view are both to the right, and the Y-axis are both downward.
[0095] The binocular ranging algorithm provided in this embodiment is simpler to calculate than the monocular ranging algorithm, and only needs to use parallax to calculate the depth of the two eyes.
[0096] As shown in FIG3 , this embodiment provides a parallax principle diagram of a binocular ranging algorithm. In the diagram, f represents the focal length of the camera, D represents the distance to the optical center, P represents the eye, and O represents the distance to the eye. L represents the first camera of the binocular camera, O R represents the second camera of the binocular camera, and Z represents the depth of the eye.
[0097] The principle formula of parallax ranging is: Where b represents the parallax, f represents the focal length of the camera, D represents the distance to the optical center, and Z represents the depth of the eye.
[0098] Optionally, in this embodiment, the parallax of the eye changes in inverse proportion to the depth of the eye.
[0099] In some embodiments, the depth of the eye is determined based on the parallax of the eye in the first view and the second view by the following steps:
[0100] Step 1a: Get the plane coordinates.
[0101] Obtain the plane coordinates of the left eye and the right eye in the first view, and the plane coordinates of the left eye and the right eye in the second view;
[0102] During implementation, a facial key point recognition algorithm can be used to perform facial key point recognition on the first view to obtain the plane coordinates of the left eye and the right eye in the first view, where the plane coordinates of the left eye in this embodiment include the plane coordinates of the left eye key point, and the left eye key point refers to the coordinates of the center position of the left eye; the plane coordinates of the right eye include the plane coordinates of the right eye key point, and the right eye key point refers to the coordinates of the center position of the right eye.
[0103] It should be noted that the origin positions of the plane coordinate systems in the first view and the second view are the same, for example, both are the upper left corner of the view, and the directions of the X-axis and Y-axis are also the same, for example, the X-axis of the coordinate systems in the first view and the second view are both to the right, and the Y-axis are both downward.
[0104] The facial key point recognition algorithm in this embodiment includes but is not limited to a method based on geometric features, a method based on templates, and a method based on models. Among them, the method based on geometric features is the earliest and most traditional method, and usually needs to be combined with other algorithms to achieve better results; the template-based method can be divided into a method based on correlation matching, a feature face method, a linear discriminant analysis method, a singular value decomposition method, a neural network method, a dynamic connection matching method, etc. The model-based method includes methods based on hidden Markov models, active shape models, and active appearance models, etc. The facial key points can also be predicted based on a method based on a deep learning model. This embodiment does not impose too many restrictions on the selection of facial key point recognition algorithms (i.e., face recognition algorithms).
[0105] For example, the plane coordinates of the left eye key point in the first view are (u Ll ,v Ll ), the plane coordinates of the right eye key point are (u Lr ,v Lr ); the plane coordinates of the left eye key point in the second view are (u Rl ,v Rl ), the plane coordinates of the right eye key point are (u Rr ,v Rr ).
[0106] Step 1b: Calculate depth based on disparity.
[0107] Determine the left eye disparity based on the plane coordinates of the left eye in the first view and the second view, and determine the left eye depth based on the left eye disparity, focal length, and optical center distance; the left eye disparity represents the horizontal position deviation of the left eye of the same face in the first view and the second view;
[0108] Determine the right eye disparity based on the plane coordinates of the right eye in the first view and the second view, and determine the right eye depth based on the right eye disparity, focal length, and optical center distance; the right eye disparity represents the horizontal position deviation of the right eye of the same face in the first view and the second view;
[0109] Wherein, the optical center distance represents the distance between the centers of the two camera lenses of the binocular camera;
[0110] From the structure of the binocular camera, we can see that the two cameras of the binocular camera are located on the same horizontal axis and are set parallel to each other. Therefore, the binocular camera only has an offset on the X-axis. Therefore, the parallax can be calculated by the eye coordinate component on the X-axis to estimate the depth of the eye.
[0111] In implementation, for example, according to the left eye key point (u Ll ,v Ll ) in the X-axis component and the left eye key point (u Rl ,v Rl ) on the X-axis to determine the parallax of the left eye, and the right eye key point (u Lr ,v Lr ) in the X-axis component and the right eye key point (u Rr ,v Rr )The difference in the components on the X-axis determines the parallax of the right eye.
[0112] The depth of the left and right eyes is calculated based on the principle formula of parallax ranging, as follows:
[0113] The left eye depth is calculated using the following formula:
[0114] In formula (1), u Ll Indicates the x coordinate of the left eye key point in the first view, u Rl Indicates the x coordinate of the left eye key point in the second view, u Ll -u Rl represents the parallax of the left eye, f represents the focal length of the camera, D represents the distance from the optical center, and Z l Indicates the depth of the left eye.
[0115] The right eye depth is calculated using the following formula:
[0116] In formula (1), u Lr Indicates the x coordinate of the right eye key point in the first view, u Rr Indicates the x coordinate of the right eye key point in the second view, u Lr -u Rr represents the parallax of the right eye, f represents the focal length of the camera, D represents the distance from the optical center, and Z r Indicates the depth of the right eye.
[0117] Step 1c: Determine the depth of the eye based on the left eye depth and the right eye depth.
[0118] In practice, the depths of the two eyes can be calculated using the above formula.
[0119] In some embodiments, when it is determined that the first view and the second view do not contain the same face, the depth of the eyes is determined using a monocular ranging algorithm;
[0120] The spatial position of the eyes of a human face is determined based on the depth of the eyes determined by the monocular ranging algorithm.
[0121] Optionally, determine the depth of the eye using a monocular ranging algorithm by following these steps:
[0122] Step 2a: Using a facial key point recognition algorithm, obtain the plane coordinates of facial key points in the selected view;
[0123] Step 2b: Determine the head posture angle based on the plane coordinates of the facial key points, the spatial coordinates of the preset facial key points, and the intrinsic parameter matrix of the binocular camera, and determine the rotation matrix based on the head posture angle;
[0124] During implementation, the spatial coordinates of the preset facial key points are determined based on the 3D coordinates of the facial key points of the preset 3D face model, and the number of facial key points includes but is not limited to any one of 68, 98, 106 and more than 1000 key points.
[0125] Head pose estimation refers to obtaining the head pose angles from a facial image. Figure 4 shows a schematic diagram of head pose estimation. In 3D space, the rotation of an object can be represented by three Euler angles: pitch (rotation around the X-axis), yaw (rotation around the Y-axis), and roll (rotation around the Z-axis).
[0126] When performing head posture estimation, as shown in FIG5 , a schematic diagram of a 2D to 3D conversion relationship is provided, wherein the 2D facial key point detection is first performed on the face image (50) in the first view or the second view to obtain the plane coordinates (510) of the facial key point, i.e., the 2D coordinates. Then, based on the 3D coordinates (530) of the 3D facial key point corresponding to the preset 3D face model (52), 3D face model matching is performed to solve the conversion relationship between the 3D key point and the corresponding 2D key point to obtain a rotation matrix, and finally the Euler angle is solved through the rotation matrix.
[0127] Optionally, based on the plane coordinates of the facial key points, the head posture angle (θ x ,θ y ,θz ). In practice, the posture of an object relative to the camera can be represented by a rotation matrix and a translation matrix. The translation matrix refers to the spatial position relationship matrix of the object relative to the camera, and the rotation matrix refers to the spatial posture relationship matrix of the object relative to the camera. Therefore, when solving the head posture angle using a two-dimensional image (i.e., the first view or the second view), it is necessary to convert between different coordinate systems. The coordinate systems include the world coordinate system (UVW), the camera coordinate system (XYZ), the image center coordinate system (uv), and the pixel coordinate system (xy).
[0128] (1)PnP algorithm.
[0129] OpenCV has provided a function solvePnp() to solve the PnP problem. Its output includes translation vectors and rotation vectors, so we only need to use the rotation vectors to calculate the Euler angles. The algorithm principle is to transform the 3D facial key points from the world coordinate system to the camera coordinate system, completing the mapping conversion and calibration between the world coordinate system (3D), the input 2D image (first view or second view), and the camera coordinate system. The Pitch (θ around the X axis) x ), Yaw (θ around the Y axis y ), Roll(θ around the Z axis z Given the plane coordinates of the facial key points, the 3D coordinates of the facial key points, and the intrinsic parameter matrix of the binocular camera, call the OpenCV solvePnp function to obtain the rotation matrix and then calculate the Euler angles.
[0130] (2) Gold standard algorithm.
[0131] 2a) Normalization processing;
[0132] In the implementation, the plane coordinates of the 2D face key points in the first view or the second view and the spatial coordinates of the preset 3D face key points are translated so that the centroids of the 2D face key points and the 3D face key points are located at the origin, and the 2D face key points and the 3D face key points after the translation processing are scaled so that the average value of the distance from the corresponding centroid to the origin is
[0133] 2b) The correspondence between 2D face key points and 3D face key points;
[0134] Calculate the correspondence between 2D face key points and 3D face key points, where each set of 2D face key points and 3D face key points has a relationship of the form AX=b, where A represents the transformation matrix, X represents the plane coordinates of the 2D face key points, and b represents the spatial coordinates of the 3D face key points.
[0135] 2c) Compute the affine matrix.
[0136] Solve the pseudo-inverse matrix of A, perform denormalization, and obtain the affine matrix. After angle conversion through the affine matrix, the head posture angle is obtained.
[0137] Step 2c: Determine the depth of the eye center point based on the eye's plane coordinates, the rotation matrix, and the intrinsic parameter matrix;
[0138] The eye center coordinates are preset as (x c ,y c ,z c ), from (x c ,y c ,z c ) is moved d / 2 to the left and right along the line connecting the two eyes to obtain the coordinates of the left eye and the right eye, where d represents the fixed pupil distance.
[0139] Using the head posture angle (θ x ,θ y ,θ z ) Construct the rotation matrix R x 、R y 、R z , as shown in the following formula:
[0140] Using the intrinsic parameter matrix of the binocular camera The plane coordinates of the eyes obtained by using facial key points recognition in the first view or the second view, that is, the plane coordinates of the left eye key points (u l ,v l ) Plane coordinates of the right eye key point (u r ,v r ), construct a set of equations to solve the depth z of the center of the eye c , as shown below:
[0141] In formula (4), d represents the fixed pupil distance, z c Indicates the depth of the eye center. That is, the depth of the eye center can be calculated using the plane coordinates, posture angles, intrinsic parameter matrix, and fixed pupil distance of the left and right eyes.
[0142] Step 2d: Determine the left eye depth and the right eye depth based on the depth of the eye center point, the pupil distance, and the rotation matrix.
[0143] Preset the left eye space coordinates (x l ,y l ,z l ), right eye space coordinate (x r ,y r,z r ), the coordinates of the eye center are (x c ,y c ,z c ), from (x c ,y c ,z c ) point is moved d / 2 to the left and right along the line connecting the two eyes to obtain the spatial coordinates of the left eye and the right eye, where d represents the fixed pupil distance.
[0144] First, according to the rotation matrix R x 、R y 、R z , fixed pupil distance d, and the depth of the eye center z c Estimate the left eye depth and right eye depth. The estimation formula is as follows:
[0145] In formula (5), * represents unknown number, z l Indicates the left eye depth, z r Indicates the right eye depth.
[0146] In some embodiments, after obtaining the depth of the eye through the above monocular ranging method, the depth of the eye can be further used to calculate the spatial coordinates of the center point of the eye. The specific calculation process is as follows:
[0147] In formula (6), is the internal parameter matrix of the binocular camera, (u l ,v l ) represents the plane coordinates of the left eye key point, (u r ,v r ) represents the plane coordinates of the right eye key point, z l Indicates the left eye depth, z r Indicates the right eye depth, x′ l , y′ l , z′ l represents the intermediate variable, x′ r , y′ r , z′ r Indicates an intermediate variable. (x c ,y c ,z c ) represents the coordinates of the eye center, (x l ,y l ,z l ) represents the left eye space coordinate, (x r ,y r ,z r ) represents the right eye space coordinate.
[0148] During implementation, when calculating the spatial coordinates of the eye center using a monocular ranging algorithm, the left and right eye depths are first calculated based on the depth of the eye center. Then, the spatial coordinates of the eye center are calculated by combining the intrinsic parameter matrix, the left and right eye plane coordinates, and the rotation matrix. As can be seen, using a monocular ranging algorithm to calculate the depth of the eye and the spatial coordinates of the eye center requires constructing a large number of formulas, resulting in a large amount of computation, complexity, and time consumption. The binocular ranging algorithm provided in this embodiment, however, calculates the depth of the eye using only parallax, focal length, and optical center distance. This simplifies the calculation, conserves processing resources, and reduces processing time.
[0149] After obtaining the left eye depth and the right eye depth using the binocular ranging algorithm provided in this embodiment, an optional implementation method is to use the above formula (6) to calculate the spatial coordinates of the eye center based on the plane coordinates of the eyes in the first view or the second view. Optionally, the first view or the second view is selected randomly, or a view corresponding to the deviation state is selected from the first view and the second view.
[0150] In some embodiments, before using a binocular ranging algorithm to determine the depth of the eyes, this embodiment needs to determine whether the first view and the second view contain the same face. When the first view and the second view contain the same face, the binocular ranging algorithm is used to determine the depth of the eyes of the face based on the first view and the second view.
[0151] This embodiment provides multiple ways to determine whether the first view and the second view contain the same face, as follows:
[0152] The first judgment method is based on similarity matching.
[0153] A similarity match is performed on the faces in the first view and the second view, and a determination is made based on the matching result whether the first view and the second view contain the same face.
[0154] During implementation, without considering the computing power limitations of the computer system, face detection can be performed separately on the first view and the second view. When multiple faces are detected, the closer the face is to the screen, the larger the area of the face captured by the binocular camera in the view. Therefore, in order to ensure the accuracy of detection, the face detected by the face detection frame with the largest area is used as the reference, and the face image in the detection frame with the largest area is cropped out, and key points of the face image are recognized to locate the eyes in the face.
[0155] During implementation, when the persons in the first view and the second view match, it is determined that the first view and the second view contain the same face; otherwise, it is determined that the first view and the second view do not contain the same face. That is, when the persons in the first view and the second view do not match, at this time, only one camera has captured the face, and the face only exists in one view, that is, the face only exists in the first view or the second view. A monocular ranging algorithm is required to calculate the depth of the human eye and then locate the spatial coordinates of the human eye.
[0156] The second judgment method is based on position mapping.
[0157] It should be noted that when judging whether the first view and the second view contain the same face based on position mapping, it is only necessary to perform face recognition processing on one view, and crop the face image in the other view through position mapping. By comparing whether the cropped face image contains a face, it is determined whether the first view and the second view contain the same face.
[0158] Whether the first view and the second view contain the same face is determined based on whether the faces in the first view and the second view have a position mapping relationship.
[0159] During implementation, based on the characteristic that a binocular camera captures faces through two cameras, the same facial key points will have a position mapping relationship in the two views. When the faces in the first view and the second view have a position mapping relationship, it is determined that the first view and the second view contain the same face, and the binocular ranging algorithm can be used to calculate the depth of the eyes; when the faces in the first view and the second view do not have a position mapping relationship, it means that the first view and the second view do not contain the same face, and a monocular ranging algorithm needs to be used to calculate the depth of the eyes.
[0160] The third judgment method is based on any of the following coordinate estimation methods.
[0161] Method a1,
[0162] 1) Obtaining the first spatial coordinates of the eye center point corresponding to the face in the first view, and estimating the facial region of the face in the second view based on the first spatial coordinates, the optical center distance of the binocular camera, and the intrinsic parameter matrix; wherein the eye center point is located at the midpoint of the line connecting the left eye and the right eye, and the optical center distance represents the distance between the centers of the two camera lenses of the binocular camera.
[0163] During implementation, the first spatial coordinates of the center point of the eye in the first view are determined based on the depth of the eye in the first view; the second spatial coordinates of the center point of the eye in the second view are estimated based on the first spatial coordinates and the optical center distance; and the facial area of the face in the second view is estimated based on the second spatial coordinates and the intrinsic parameter matrix.
[0164] Specifically, assuming there are M preset 3D face key points, M can be 68, 98, 106, 468, etc., calculate the 3D eye center key point (x gt-c ,y gt-c ,z gt-c ) and other 3D facial key points; specifically expressed as follows: Landmark gt =[(x i ,y i ,z i )], i=1,2,…,M; Landmark offset =[(x gt-c -x i ,y gt-c -y i ,z gt-c -z i )], i=1,2,…,M.
[0165] Among them, (x gt-c ,y gt-c ,z gt-c ) is the 3D eye center key point, M is the number of facial key points, Landmark gt Landmark is the 3D face key point other than the 3D eye center key point. offset It is the deviation between the 3D eye center key point and other 3D face key points in the preset 3D face key points.
[0166] In implementation, a preset face and its spatial coordinates are known. Using the eye center of the facial landmark as a reference, the deviations between the other facial landmarks and the eye center are calculated. The second spatial coordinate of the eye center in the second view is estimated using the first spatial coordinate of the eye center and the optical center distance D through the following steps:
[0167] The first spatial coordinate (x Lc ,y Lc ,z Lc ), the second spatial coordinate (x Rc ,y Rc ,z Rc ) to estimate:
[0168] In formula (7), Landmark R The spatial coordinates of other facial key points except the eye center in the second view; Landmark offset is the deviation, D is the optical center distance, and M is the number of facial key points.
[0169] After obtaining the spatial coordinates of the estimated facial key points in the second view, the plane coordinates of the corresponding facial key points in the second view can be obtained based on the intrinsic parameter matrix of the binocular camera. Then, the face area can be determined based on the plane coordinates of the facial key points in the second view. The estimation formula is as follows:
[0170] In formula (8), Landmark RUV represents the plane coordinates of the facial key points in the second view, is the internal parameter matrix, Landmark R Represents the spatial coordinates of the facial key points in the second view.
[0171] The face image is cropped according to the positions of the facial key points in the second view to obtain the face region.
[0172] 2) Determine whether the faces in the first view and the second view include the same face based on whether a face is detected in the face area in the second view.
[0173] During implementation, when a face is detected in the face area in the second view, it is determined that the first view and the second view contain the same face. At this time, the binocular ranging algorithm can be used to calculate the depth of the eyes; when no face is detected in the face area in the second view, it is determined that the first view and the second view do not contain the same face. At this time, it is necessary to use a monocular ranging algorithm to calculate the depth of the eyes, and then calculate the single-view spatial coordinates of the eyes.
[0174] Method a2,
[0175] 1) Obtain the second spatial coordinates of the eye center point corresponding to the face in the second view, and estimate the facial area of the face in the first view based on the second spatial coordinates, the optical center distance of the binocular camera, and the intrinsic parameter matrix;
[0176] The eye center point is located at the midpoint of the line connecting the left eye and the right eye, and the optical center distance represents the distance between the centers of the two camera lenses of the binocular camera.
[0177] During implementation, the second spatial coordinates of the center point of the eye in the second view are determined based on the depth of the eye in the second view; the first spatial coordinates of the center point of the eye in the first view are estimated based on the second spatial coordinates and the optical center distance; and the facial area of the face in the first view is estimated based on the first spatial coordinates and the intrinsic parameter matrix.
[0178] Specifically, assuming there are M preset 3D face key points, M can be 68, 98, 106, 468, etc., calculate the 3D eye center key point (x gt-c ,y gt-c ,z gt-c) and other 3D facial key points; specifically expressed as follows: Landmark gt =[(x i ,y i ,z i )], i=1,2,…,M; Landmark offset =[(x gt-c -x i ,y gt-c -y i ,z gt-c -z i )], i=1,2,…,M.
[0179] Among them, (x gt-c ,y gt-c ,z gt-c ) is the 3D eye center key point, M is the number of facial key points, Landmark gt Landmark is the 3D face key point other than the 3D eye center key point. offset It is the deviation between the 3D eye center key point and other 3D face key points in the preset 3D face key points.
[0180] In implementation, given a preset face and its spatial coordinates, the eye center of the facial landmarks is used as a reference to calculate the deviation between the other facial landmarks and the eye center. The first spatial coordinate of the eye center in the first view is estimated using the second spatial coordinate of the eye center and the optical center distance D through the following steps:
[0181] The second spatial coordinates (x Rc ,y Rc ,z Rc ), the first spatial coordinate (x Lc ,y Lc ,z Lc ) to estimate:
[0182] In formula (9), Landmark L Landmark is the spatial coordinate of other facial key points except the eye center in the first view; offset is the deviation, D is the optical center distance, and M is the number of facial key points.
[0183] After obtaining the spatial coordinates of the estimated facial key points in the first view, the plane coordinates of the corresponding facial key points in the first view can be obtained based on the intrinsic parameter matrix of the binocular camera. Then, the face area can be determined based on the plane coordinates of the facial key points in the first view. The estimation formula is as follows:
[0184] In formula (8), Landmark LUV represents the plane coordinates of the facial key points in the first view, is the internal parameter matrix, Landmark L Represents the spatial coordinates of the facial key points in the first view.
[0185] According to the positions of the facial key points in the first view, the facial image is cropped to obtain the facial region.
[0186] 2) Determine whether the faces in the first view and the second view include the same face based on whether a face is detected in the face area in the first view.
[0187] During implementation, when a face is detected in the face area in the second view, it is determined that the first view and the second view contain the same face. At this time, it is necessary to use a binocular ranging algorithm to calculate the depth of the eyes; when no face is detected in the face area in the second view, it is determined that the first view and the second view do not contain the same face. At this time, it is necessary to use a monocular ranging algorithm to calculate the depth of the eyes, and then calculate the single-view spatial coordinates of the eyes.
[0188] The fourth judgment method is based on the field of view.
[0189] It should be noted that the field of view refers to the field of view of the binocular camera, which refers to the size of the area actually captured by the binocular camera. When determining whether the first and second views contain the same face based on the field of view, only one view needs to be processed. For example, the first view can be selected for face recognition, key point detection, eye center depth calculation, eye location, and other processing steps. Whether the first and second views contain the same face is determined based on whether the horizontal position of the located eyes is within the field of view.
[0190] In some embodiments, the processor is specifically configured to determine the field of view area by:
[0191] Determine the depth of the eye center in the first view according to a monocular ranging algorithm, and determine the field of view area according to the depth of the eye center in the first view, the field of view angle, and the optical center distance; or
[0192] determining the depth of the eye center in the second view according to a monocular ranging algorithm, and determining the field of view area according to the depth of the eye center in the second view, the field of view angle, and the optical center distance;
[0193] The eye center point is located at the midpoint of the line connecting the left eye and the right eye; the optical center distance represents the distance between the centers of the two camera lenses of the binocular camera, and the field of view angle represents the horizontal field of view angle captured by a single camera in the binocular camera.
[0194] As shown in Figure 6, a schematic diagram of binocular camera field of view estimation is provided. In the figure, C1 and C2 represent two cameras, D is the optical center distance between the centers of the two camera lenses, α is the horizontal field of view angle, Z1+Z2=Z, and Z is the depth of the eye center point. The X-axis coordinate range at this depth Z can be estimated.
[0195] Method b1: determine the visual field area and the single-view spatial coordinates of the eyes in the first view, and determine whether the first view and the second view contain the same face based on the relationship between the single-view spatial coordinates of the eyes in the first view and the visual field area.
[0196] During implementation, the single-view spatial coordinates of the eyes in the first view are obtained, the field of view area is determined based on the depth of the center point of the eye in the first view, and whether the first view and the second view contain the same face is determined based on the relationship between the single-view spatial coordinates of the eyes in the first view and the field of view area.
[0197] In some embodiments, the field of view is determined as follows:
[0198] Determine the depth of the eye center in the first view according to a monocular ranging algorithm, and determine the field of view area according to the depth of the eye center in the first view, the field of view angle, and the optical center distance; or
[0199] determining the depth of the eye center in the second view according to a monocular ranging algorithm, and determining the field of view area according to the depth of the eye center in the second view, the field of view angle, and the optical center distance;
[0200] The eye center point is located at the midpoint of the line connecting the left eye and the right eye; the optical center distance represents the distance between the centers of the two camera lenses of the binocular camera, and the field of view angle represents the horizontal field of view angle captured by a single camera in the binocular camera.
[0201] During implementation, a monocular ranging algorithm is used to determine the single-view spatial coordinates of the eyes in the first view. First, face recognition is used to obtain the plane coordinates of the eyes in the first view. Based on the plane coordinates of the eyes, the intrinsic parameter matrix of the binocular camera, and the spatial coordinates of the preset 3D face key points, the depth of the center point of the eye in the first view is calculated. The single-view spatial coordinates of the eyes are calculated based on the depth of the center point of the eyes in the first view, and then the spatial coordinates of the left eye and the right eye in the first view are obtained.
[0202] The depth Z of the eye center can be estimated by the monocular ranging algorithm. Given the horizontal field of view angle α and the optical center distance D, the field of view area in the X-axis direction can be determined by the following formula:
[0203] In formula (11), D is the distance from the optical center, Z l is the depth of the eye center in the first view, X c For the field of view area.
[0204] During implementation, when the component of the spatial coordinate of the left eye in the first view in the X-axis direction and the component of the spatial coordinate of the right eye in the X-axis direction are both within the field of view, it is determined that the first view and the second view contain the same face. In this case, the binocular ranging algorithm is used to calculate the depth of the eyes. Otherwise, it is determined that the first view and the second view do not contain the same face. In this case, the monocular ranging algorithm is used to calculate the depth of the eyes, thereby obtaining the spatial position of the eyes.
[0205] Method b2: determine the visual field area and the single-view spatial coordinates of the eyes in the second view, and determine whether the first view and the second view contain the same face based on the relationship between the single-view spatial coordinates of the eyes in the second view and the visual field area.
[0206] During implementation, the single-view spatial coordinates of the eye in the second view are obtained, the field of view area is determined based on the depth of the eye center point in the second view, and whether the first view and the second view contain the same face is determined based on the relationship between the single-view spatial coordinates of the eye in the second view and the field of view area; wherein the eye center point is located at the midpoint of the line connecting the left eye and the right eye.
[0207] In some embodiments, the field of view is determined as follows:
[0208] Method to determine the field of view:
[0209] Determine the depth of the eye center in the first view according to a monocular ranging algorithm, and determine the field of view area according to the depth of the eye center in the first view, the field of view angle, and the optical center distance; or
[0210] determining the depth of the eye center in the second view according to a monocular ranging algorithm, and determining the field of view area according to the depth of the eye center in the second view, the field of view angle, and the optical center distance;
[0211] The eye center point is located at the midpoint of the line connecting the left eye and the right eye; the optical center distance represents the distance between the centers of the two camera lenses of the binocular camera, and the field of view angle represents the horizontal field of view angle captured by a single camera in the binocular camera.
[0212] During implementation, a monocular ranging algorithm is used to determine the single-view spatial coordinates of the eyes in the second view. First, face recognition is used to obtain the plane coordinates of the eyes in the second view. Based on the plane coordinates of the eyes, the intrinsic parameter matrix of the binocular camera, and the spatial coordinates of the preset 3D face key points, the depth of the center point of the eye in the second view is calculated. The single-view spatial coordinates of the eyes are calculated based on the depth of the center point of the eyes in the second view, and then the spatial coordinates of the left eye and the right eye in the second view are obtained.
[0213] The depth Z of the eye center can be estimated by the monocular ranging algorithm. Given the horizontal field of view angle α and the optical center distance D, the field of view area in the X-axis direction can be determined by the following formula:
[0214] In formula (12), D is the distance from the optical center, Z r is the depth of the eye center in the second view, X c For the field of view area.
[0215] During implementation, when the component of the spatial coordinate of the left eye in the X-axis direction and the component of the spatial coordinate of the right eye in the X-axis direction in the second view are both within the field of view, it is determined that the first view and the second view contain the same face. In this case, the binocular ranging algorithm is used to calculate the depth of the eyes. Otherwise, it is determined that the first view and the second view do not contain the same face. In this case, the monocular ranging algorithm is used to calculate the depth of the eyes, thereby obtaining the spatial position of the eyes.
[0216] In some embodiments, the processor is further configured to execute:
[0217] When it is determined that the first view and the second view do not contain the same face, determining the depth of the eyes using a monocular ranging algorithm;
[0218] The spatial position of the eyes of a human face is determined based on the depth of the eyes determined by the monocular ranging algorithm.
[0219] Step 2: Determine the deviation state of the face relative to the center position, determine the spatial position of the eyes of the face according to the depth and deviation state of the eyes, and display a three-dimensional image based on the spatial position of the eyes.
[0220] Optionally, this embodiment can select a corresponding view from the first and second views based on the deviation state of the eyes, and use this view and the depth of the eyes to determine the spatial position of the eyes. Alternatively, a weight ratio of the first and second views can be determined based on the deviation state of the eyes, and the spatial position of the eyes can be determined based on the weight ratio and the depth of the eyes. The center position is determined based on the center of the binocular camera capturing the face.
[0221] In some embodiments, the processor is specifically configured to: determine the deviation state of the human face relative to the center position, select a view corresponding to the deviation state from the first view and the second view; and determine the spatial position of the eyes of the human face based on the depth of the eyes and the selected view.
[0222] Optionally, the processor is specifically configured to determine the view corresponding to the deviation state by any of the following methods:
[0223] (1) When the deviation state is that the face is deviated to the left relative to the center position, the view corresponding to the deviation state is determined to be the first view, and the first view is taken by the left camera of the binocular camera;
[0224] (2) When the deviation state is that the face is biased to the right relative to the center position, the view corresponding to the deviation state is determined to be the second view, and the second view is taken by the right camera of the binocular camera.
[0225] It should be noted that the distinction between the left and right cameras of the binocular camera is determined based on the face of the person facing the binocular camera as the reference system. That is, for a person facing the binocular camera, the left camera seen by the human eye is the left camera of the binocular camera, and the right camera seen by the human eye is the right camera of the binocular camera.
[0226] In practice, the center position defined in this embodiment is the midpoint of the line connecting the centers of the two cameras of the binocular camera. When the face deviates from the center position, the view closer to the center position can be selected from the first and second views to calculate the spatial position of the eyes.
[0227] In some embodiments, the deviation state of the face relative to the center position is determined by the following steps, including:
[0228] Step c1, using the depth of the eyes to determine the single view spatial coordinates of the eyes in the first view and the second view respectively; wherein the spatial coordinates of the left eye and the right eye in the first view and the spatial coordinates of the left eye and the right eye in the second view are determined.
[0229] In implementation, the single view space coordinates of the eye in the first view and the second view can be determined using the depth of the eye in the following manner:
[0230] Assume that the left eye space coordinate in the first view is (x Ll ,y Ll ,z Ll ), the right eye space coordinate is (x Lr ,y Lr ,z Lr ), the spatial coordinates of the eye center are (x Lc ,y Lc ,z Lc), where the conversion relationship between the eye center coordinates and the left and right eye space coordinates is as follows:
[0231] First, according to the rotation matrix R x 、R y 、R z , fixed pupil distance d, and the depth of the eye center z Lc Estimate the left eye depth and right eye depth. The estimation formula is as follows:
[0232] In formula (13), * represents an unknown number, z Ll Indicates the left eye depth in the first view, z Lr Indicates the right eye depth in the first view.
[0233] The depth of the eye is used to continue calculating the first spatial coordinate of the eye center point. The specific calculation process is as follows:
[0234] In formula (14), is the internal parameter matrix of the binocular camera, (u Ll ,v Ll ) represents the plane coordinates of the left eye key point in the first view, (u Lr ,v Lr ) represents the plane coordinates of the right eye key point in the first view, z Ll Indicates the left eye depth, z Lr Indicates the right eye depth, x′ Ll , y′ Ll , z′ Ll represents the intermediate variable, x′ Lr , y′ Lr , z′ Lr Indicates an intermediate variable. (x Lc ,y Lc ,z Lc ) represents the first spatial coordinate of the eye center point, (x Ll ,y Ll ,z Ll ) represents the left eye space coordinate in the first view, (x Lr ,y Lr ,z Lr ) represents the right eye space coordinate in the first view.
[0235] The spatial coordinates of the left eye and the right eye in the first view and the first spatial coordinates of the eye center point in the first view are calculated through the above steps.
[0236] Similarly, assuming that the left eye space coordinate in the second view is (xRl ,y Rl ,z Rl ), the right eye space coordinate is (x Rr ,y Rr ,z Rr ), the spatial coordinates of the eye center are (x Rc ,y Rc ,z Rc ), where the conversion relationship between the eye center coordinates and the left and right eye space coordinates is as follows:
[0237] First, according to the rotation matrix R x 、R y 、R z , fixed pupil distance d, and the depth of the eye center z Rc Estimate the left eye depth and right eye depth. The estimation formula is as follows:
[0238] In formula (15), * represents an unknown number, z Rl Indicates the left eye depth in the second view, z Rr Indicates the right eye depth in the second view.
[0239] The depth of the eye is used to continue calculating the second spatial coordinates of the eye center point. The specific calculation process is as follows:
[0240] In formula (16), is the internal parameter matrix of the binocular camera, (u Rl ,v Rl ) represents the plane coordinates of the left eye key point in the second view, (u Rr ,v Rr ) represents the plane coordinates of the right eye key point in the second view, z Rl Indicates the left eye depth, z Rr Indicates the right eye depth, x′ Rl , y′ Rl , z′ Rl represents the intermediate variable, x′ Rr , y′ Rr , z′ Rr Indicates an intermediate variable. (x Rc ,y Rc ,z Rc ) represents the second space coordinate of the eye center point (x Rl ,y Rl ,z Rl ) represents the left eye space coordinate in the second view, (x Rr ,y Rr ,zRr ) represents the right eye space coordinates in the second view.
[0241] The spatial coordinates of the left eye and the right eye in the second view, as well as the first spatial coordinates of the eye center point in the second view, are calculated through the above steps. The eye center point is located at the midpoint of the line connecting the left eye and the right eye. Optionally, the eye center point is located at the midpoint between the left eye center key point and the right eye center key point.
[0242] Step c2: Determine the deviation state of the face relative to the center position based on the single-view space coordinates of the eyes in the first view and the second view.
[0243] In some embodiments, the deviation state of the face relative to the center position is determined based on the single-view space coordinates of the eyes in the first view and the second view by the following steps:
[0244] Step c21: Determine the first spatial coordinates of the eye center point in the first view according to the spatial coordinates of the left eye and the right eye in the first view;
[0245] Step c22: determining the second spatial coordinates of the eye center point in the second view according to the spatial coordinates of the left eye and the right eye in the second view;
[0246] The eye center point is located at the midpoint of the line connecting the left eye and the right eye.
[0247] Step c23: Determine the deviation state of the face relative to the center position based on the average of the first spatial coordinate and the second spatial coordinate.
[0248] After obtaining the spatial coordinates of the eye center points in the first and second views through the above steps, the mean of the spatial coordinates of the two eye center points is determined using the following formula, and the deviation state is determined based on the size of the mean:
[0249] In formula (17), (x Lc ,y Lc ,z Lc ) represents the first spatial coordinate of the eye center point, (x Rc ,y Rc ,z Rc ) represents the second space coordinate of the eye center point of the second view, (x c ,y c ,z c ) represents the mean.
[0250] Optionally, the deviation state of the face relative to the center position and the view corresponding to the deviation state are determined according to the average of the first spatial coordinate and the second spatial coordinate in any one or more of the following ways:
[0251] Method d1) When the component of the mean in the X-axis direction is greater than the reference value, the left deviation of the face relative to the center position is determined; optionally, the reference value is a value within a preset range including zero, or the reference value is zero. The selection of the reference value can be defined according to the accuracy requirements, and this embodiment does not impose too many restrictions on this.
[0252] During implementation, when the deviation state is biased to the left, the view corresponding to the deviation state is determined to be the first view, and the first view is taken by the left camera of the binocular camera; it should be noted that the first view is the image taken by the first camera of the binocular camera, and the first camera is located on the left side of the binocular camera (the first camera is located on the left side relative to the photographed face). The first view is understood as the left view. Since the face is biased to the left relative to the center position at this time, it means that the face is closer to the left camera. Therefore, the face taken by the left camera is clearer and the face occupies a larger area in the first view. Determining the spatial position of the eyes through a clearer face can improve the positioning accuracy to a certain extent and improve the problems of inaccurate spatial coordinate calculation and large jitter.
[0253] Mode d2) when the component of the mean in the X-axis direction is less than a reference value, determining that the face is deviated to the right relative to the center position;
[0254] During implementation, when the deviation state is biased to the right, the view corresponding to the deviation state is determined to be the second view, and the second view is taken by the right camera of the binocular camera; it should be noted that the second view is the image taken by the second camera of the binocular camera, and the second camera is located on the right side of the binocular camera (the second camera is located on the right side relative to the photographed face). The second view can also be understood as the right view. Since the face is biased to the right relative to the center position at this time, it means that the face is closer to the right camera. Therefore, the face taken by the right camera is clearer and the face occupies a larger area in the first view. Determining the spatial position of the eyes through a clearer face can improve the positioning accuracy to a certain extent and improve the problems of inaccurate spatial coordinate calculation and large jitter.
[0255] Mode d3) When the component of the mean in the X-axis direction is equal to the reference value, it is determined that the face is facing the center position.
[0256] During implementation, when the deviation state is facing the center position, the view corresponding to the deviation state is determined to be the first view or the second view. It should be noted that the first view is the image captured by the first camera of the binocular camera, which is located to the left of the binocular camera. The first view is understood as the left view. The second view is the image captured by the second camera of the binocular camera. The second camera is located to the right of the binocular camera. The second view can also be understood as the right view. Since the face is facing the center position at this time, the distance from the face to the left camera and the right camera is the same. Therefore, the quality of the face captured by the left camera or the right camera is almost the same. Therefore, either the first view or the second view can be selected for eye location calculation.
[0257] In some embodiments, the depth of the eyes can also be used to determine the single-view spatial coordinates of the eyes in the first view and the second view respectively, and the deviation state of the face relative to the center position can be determined based on the single-view spatial coordinates of one eye in the first view and the second view.
[0258] During implementation, the deviation state of the face relative to the center position is determined based on the deviation state of the single view space coordinates of a single eye relative to the center position.
[0259] In some embodiments, the processor is specifically configured to determine the single view space coordinates of the other eye by:
[0260] When the selected view is the first view, determining the single view space coordinates of the one eye in the first view according to the depth of the one eye, and estimating the single view space coordinates of the other eye using the single view space coordinates of the one eye; or
[0261] When the selected view is the second view, the single view space coordinates of the one eye in the second view are determined according to the depth of the one eye, and the single view space coordinates of the other eye are estimated using the single view space coordinates of the one eye.
[0262] During implementation, when the selected view is the first view, the spatial coordinates of the left eye in the first view are determined according to the left eye depth, and the spatial coordinates of the right eye are estimated using the spatial coordinates of the left eye, or, the spatial coordinates of the right eye in the first view are determined according to the right eye depth, and the spatial coordinates of the left eye are estimated using the spatial coordinates of the right eye; or, when the selected view is the second view, the spatial coordinates of the right eye in the second view are determined according to the right eye depth, and the spatial coordinates of the left eye are estimated using the spatial coordinates of the right eye, or, the spatial coordinates of the left eye in the second view are determined according to the left eye depth, and the spatial coordinates of the right eye are estimated using the spatial coordinates of the left eye.
[0263] Optionally, the spatial position of the eyes of the face is determined according to the depth of the eyes and the selected view in the following manner:
[0264] According to the depth of the eyes, the single-view spatial coordinates of an eye in the selected view are determined, where the one eye includes a left eye or a right eye; the single-view spatial coordinates of one eye are used to estimate the single-view spatial coordinates of the other eye, and the spatial positions of the eyes of the face are determined according to the spatial coordinates of the left eye and the right eye.
[0265] Optionally, in this embodiment, the depth of the eyes includes the left eye depth and the right eye depth; the single view space coordinates of the other eye are determined by the following method:
[0266] Mode e1: When the selected view is the first view, the spatial coordinates of the left eye in the first view are determined according to the left eye depth, and the spatial coordinates of the right eye are estimated using the spatial coordinates of the left eye.
[0267] In practice, when the mean x of the first spatial coordinate and the second spatial coordinate c When ∂>0, it means that the left eye is closer to the center position. Therefore, the view corresponding to the deviation state is determined to be the first view (i.e., the left view). The spatial coordinates of the left eye in the first view are used to estimate the spatial coordinates of the right eye. The estimation formula is as follows:
[0268] In formula (18), (x Ll ,y Ll ,z Ll ) represents the spatial coordinates of the left eye in the first view, d is the fixed pupil distance, R z 、R y 、R x Both represent rotation matrices, (x r ,y r , z r ) represents the estimated spatial coordinates of the right eye.
[0269] Mode e2: when the selected view is the second view, the spatial coordinates of the right eye in the second view are determined according to the right eye depth, and the spatial coordinates of the left eye are estimated using the spatial coordinates of the right eye.
[0270] In practice, when the mean x of the first spatial coordinate and the second spatial coordinate c When ∂<0, it means that the face is deviated to the right relative to the center position, and the right eye is closer to the center position. Therefore, the view corresponding to this deviation state is determined to be the second view (i.e., the right view). The spatial coordinates of the right eye in the second view are used to estimate the spatial coordinates of the left eye. The estimation formula is as follows:
[0271] In formula (18), (x Rr ,y Rr ,z Rr ) represents the right eye spatial coordinate in the second view, d is the fixed pupil distance, Rz 、R y 、R x Both represent rotation matrices, (x l ,y l , z l ) represents the estimated spatial coordinates of the left eye.
[0272] In some embodiments, in addition to the above-mentioned method for estimating the single-view spatial coordinates of an eye using one view and depth, this embodiment further provides a method for estimating the single-view spatial coordinates of an eye using a first view and a second view combined with depth. The specific implementation process is as follows:
[0273] Determining the single view space coordinates of the eye in the first view and the second view respectively using the depth of the eye; determining weights corresponding to the first view and the second view according to the selected view, wherein the weights are determined based on the deviation state;
[0274] The weights corresponding to the first view and the second view are used to perform weighted averaging on the single-view spatial coordinates of the eyes in the first view and the second view, and the spatial positions of the eyes of the face are determined according to the weighted average.
[0275] During implementation, the depth of the eyes is obtained by parallax calculation, and the first spatial coordinates of the left eye and the right eye in the first view, as well as the second spatial coordinates of the left eye and the right eye in the second view, are calculated using the depth of the eyes; according to the respective weights of the first view and the second view, the first spatial coordinate of the left eye in the first view and the second spatial coordinate of the left eye in the second view are weightedly summed to obtain the spatial position of the left eye; and the first spatial coordinate of the right eye in the first view and the second spatial coordinate of the right eye in the second view are weightedly summed to obtain the spatial position of the right eye.
[0276] In some examples, the spatial position of the eye is obtained by weighted summing the following formula:
[0277] In formula (20), th represents the threshold, for example, th = 0.3; θ represents the angle of the face deviation from the center position in the X direction determined based on the deviation state; λ represents the weight ratio of the first view and the second view; eyecenter l [2] represents the depth of the eye center in the first view, eyecenter l [0] represents the horizontal coordinate of the eye center point in the first view, eyecenter l [1] represents the vertical coordinate of the eye center point in the first view, eyecenter r[2] represents the depth of the eye center in the second view, eyecenter r [0] represents the horizontal coordinate of the eye center point in the second view, eyecenter r [1] represents the vertical coordinate of the eye center in the second view; eyecenter[0] represents the horizontal coordinate of the eye center, eyecenter[1] represents the vertical coordinate of the eye center, and eyecenter[2] represents the depth of the eye center.
[0278] The depth of the eye center is calculated using the above formula (20), and the spatial position of the eye is determined using the depth of the eye center. In implementation, the weight ratio of the first view and the second view is determined using the deviation state. The weight ratio of the first view and the second view is used to perform a weighted sum of the depths of the eye center in the first view and the second view to obtain the final depth of the eye center. The horizontal coordinates of the eye center in the first view and the second view are averaged to obtain the final horizontal coordinate of the eye center. Similarly, the vertical coordinates of the eye center in the first view and the second view are averaged to obtain the final vertical coordinate of the eye center. The spatial position of the eye is determined based on the final spatial coordinates of the eye center.
[0279] Optionally, the weight corresponding to the first view can be one or two, and the weight corresponding to the second view can also be one or two. When the weight corresponding to the first view is one, the weight corresponding to the left eye and the weight corresponding to the right eye in the first view are the same. Similarly, when the weight corresponding to the second view is one, the weight corresponding to the left eye and the weight corresponding to the right eye in the second view are the same. When the weight corresponding to the first view is two, the weight corresponding to the left eye and the weight corresponding to the right eye in the first view are different. Similarly, when the weight corresponding to the second view is two, the weight corresponding to the left eye and the weight corresponding to the right eye in the second view are different.
[0280] Optionally, when the corresponding view is determined to be the first view based on the deviation state, indicating that the left eye in the first view is closer to the center position, the weight of the left eye in the first view can be set to be greater than the weight of the right eye, and the weight of the left eye can be set to be greater than any weight corresponding to the second view. Similarly, when the corresponding view is determined to be the second view based on the deviation state, indicating that the right eye in the second view is closer to the center position, the weight of the right eye in the second view can be set to be greater than the weight of the left eye, and the weight of the right eye can be set to be greater than any weight corresponding to the first view.
[0281] Optionally, when the deviation state is directly facing the center position, the weights corresponding to the first view and the second view are the same.
[0282] In practice, after calculating the depth of the eye and the selected view through the above steps, the spatial position of the eye can also be determined based on the depth of the eye and the selected view. The specific calculation process is as follows:
[0283] In formula (21), is the internal parameter matrix of the binocular camera, (u l ,v l ) represents the plane coordinates of the left eye key point in the selected view, (u r ,v r ) represents the plane coordinates of the right eye key point in the selected view, z l Indicates the left eye depth, z r Indicates the right eye depth, x′ l , y′ l , z′ l represents the intermediate variable, x′ r , y′ r , z′ r Indicates an intermediate variable. (x c ,y c ,z c ) represents the coordinates of the eye center, (x l ,y l ,z l ) represents the left eye space coordinate, (x r ,y r ,z r ) represents the right eye space coordinate.
[0284] In implementation, the left eye spatial coordinates are determined using the left eye depth and the plane coordinates of the left eye in the selected first view (i.e., the plane coordinates of the left eye keypoints) and the intrinsic parameter matrix, and the right eye spatial coordinates are estimated using the left eye spatial coordinates in the selected first view, the fixed pupil distance, and the rotation matrix to obtain the spatial position of the eye. Alternatively, the right eye spatial coordinates are determined using the right eye depth and the plane coordinates of the right eye in the selected second view (i.e., the plane coordinates of the right eye keypoints) and the intrinsic parameter matrix, and the left eye spatial coordinates are estimated using the right eye spatial coordinates in the selected second view, the fixed pupil distance, and the rotation matrix to obtain the spatial position of the eye.
[0285] Optionally, when the view corresponding to the deviation state is the first view, the left eye spatial coordinates can be determined using the left eye depth, the plane coordinates of the left eye in the selected first view (i.e., the plane coordinates of the left eye key points), and the intrinsic parameter matrix; the right eye spatial coordinates can be determined using the right eye depth, the plane coordinates of the right eye in the first view (i.e., the plane coordinates of the right eye key points), and the intrinsic parameter matrix; thereby obtaining the spatial position of the eye. Alternatively, the left eye spatial coordinates can be determined using the left eye depth, the plane coordinates of the left eye in the selected first view (i.e., the plane coordinates of the left eye key points), and the intrinsic parameter matrix; the right eye spatial coordinates can be determined using the right eye depth, the plane coordinates of the right eye in the second view (i.e., the plane coordinates of the right eye key points), and the intrinsic parameter matrix; thereby obtaining the spatial position of the eye.
[0286] Optionally, when the view corresponding to the deviation state is the second view, the right eye spatial coordinates can be determined using the right eye depth, the plane coordinates of the right eye in the selected second view (i.e., the plane coordinates of the right eye key point), and the intrinsic parameter matrix; the left eye spatial coordinates can be determined using the left eye depth, the plane coordinates of the left eye in the second view (i.e., the plane coordinates of the left eye key point), and the intrinsic parameter matrix; thereby obtaining the spatial position of the eye. Alternatively, the right eye spatial coordinates can be determined using the right eye depth, the plane coordinates of the right eye in the selected second view (i.e., the plane coordinates of the right eye key point), and the intrinsic parameter matrix; the left eye spatial coordinates can be determined using the left eye depth, the plane coordinates of the left eye in the first view (i.e., the plane coordinates of the left eye key point), and the intrinsic parameter matrix; thereby obtaining the spatial position of the eye.
[0287] In practice, the parallax principle of a binocular ranging algorithm can be used to estimate eye depth, and then the spatial position of the eye can be calculated based on the depth estimated by the binocular ranging algorithm. However, this method can result in significant jitter in the spatial coordinates of the eye. For example, the calculated key points of the eye may be offset by one pixel, resulting in a 5-6mm deviation in eye depth. Therefore, this embodiment also provides a method for correcting the fixed pupil distance using the parallax principle, thereby obtaining a pupil distance close to the true pupil distance of the human eye, and using the corrected pupil distance to calculate eye depth.
[0288] In some embodiments, the processor of this embodiment is specifically configured to execute:
[0289] (1) correcting the fixed pupil distance to obtain a corrected pupil distance according to the depth of the eyes determined by the binocular ranging algorithm;
[0290] During implementation, correction is performed in the following ways:
[0291] Determining a first depth of the eye using a monocular ranging algorithm; wherein the first depth of the eye is determined using the monocular ranging algorithm based on a fixed pupil distance and an intrinsic parameter matrix of a binocular camera;
[0292] Determine the depth of the eyes using binocular ranging algorithms;
[0293] The fixed pupil distance is corrected according to the first depth and the depth of the eyes determined by the binocular distance measurement algorithm to obtain a corrected pupil distance.
[0294] Optionally, the corrected pupil distance and the depth of the eye determined by the binocular distance measurement algorithm change in direct proportion, and the corrected pupil distance and the first depth change in inverse proportion.
[0295] Specifically, the corrected pupil distance is obtained by the following formula:
[0296] In formula (22), D represents the fixed pupil distance, z represents the depth of the eye determined by the binocular ranging algorithm, and z single Indicates the first depth, D real Indicates corrected interpupillary distance.
[0297] (2) using a monocular ranging algorithm to determine the corrected depth of the eyes of the face according to the corrected pupil distance;
[0298] In practice, the corrected depth is obtained as follows:
[0299] Determining, using a monocular ranging algorithm, a first corrected depth of the eyes of the human face in the first view and a second corrected depth of the eyes of the human face in the second view based on the corrected pupil distance;
[0300] The corrected depth of the eyes of the human face is determined according to an average value of the first corrected depth and the second corrected depth.
[0301] (3) Determine the spatial position of the eyes of the human face based on the corrected depth and deviation state of the eyes of the human face.
[0302] Optionally, a view corresponding to the deviation state is selected from the first view and the second view; and the spatial position of the eyes of the human face is determined according to the corrected depth of the eyes and the selected view.
[0303] Optionally, based on the corrected depth of the eye, the single-view spatial coordinates of an eye in the selected view are determined, and the one eye includes the left eye or the right eye; the single-view spatial coordinates of one eye are used to estimate the single-view spatial coordinates of the other eye, and the spatial position of the eyes of the face is determined based on the spatial coordinates of the left eye and the right eye.
[0304] Optionally, determine the single view space coordinates of the other eye as follows:
[0305] When the selected view is the first view, determining the single view space coordinates of the one eye in the first view according to the corrected depth of the one eye, and estimating the single view space coordinates of the other eye using the single view space coordinates of the one eye; or
[0306] When the selected view is the second view, the single view space coordinates of the one eye in the second view are determined according to the corrected depth of the one eye, and the single view space coordinates of the other eye are estimated using the single view space coordinates of the one eye.
[0307] Optionally, the processor is further configured to execute:
[0308] The method comprises the steps of: using the corrected depth of the eye to respectively determine the single-view spatial coordinates of the eye in the first view and the second view; determining weights corresponding to the first view and the second view according to the selected view, wherein the weights are determined based on the deviation state; and using the weights corresponding to the first view and the second view to perform weighted averaging on the single-view spatial coordinates of the eye in the first view and the second view, and determining the spatial position of the eye of the face according to the weighted average.
[0309] In the implementation, the pupil distance d and the head pose angle θ are used to estimate the 3D pupil position using the camera intrinsic parameter M. The head pose angle θ is obtained using the facial key points using the PnP algorithm or the gold standard algorithm. The head pose angle is relative to the camera global coordinate system and consists of the pitch angle, yaw angle, and roll angle rotation, denoted as θ respectively. x ,θ y ,θ z , which correspond to the rotation angles around the X, Y, and Z axes respectively. The camera intrinsic parameter matrix of the binocular camera is expressed as, By θ x ,θ y ,θ z The constructed rotation matrix is expressed as:
[0310] Among them, C z =cos(θ z ), S z = sin(θ z ), C y =cos(θ y ), S y = sin(θ y ), C x =cos(θ x ), S x = sin(θ x );
[0311] Taking a single view as an example (i.e., either the first view or the second view), the spatial coordinates of the left eye center position and the right eye center position are marked as X l and X r , the center of the eye is marked as X c The left eye center position X l (spatial coordinates) and the right eye center position Xr (Spatial coordinates) can be expressed by the eye center point as: Therefore, the number of unknown variables is reduced from 6 to 3, with X c As a variable, the equation is established according to the projection formula of the pinhole camera, where x l and x r They are the plane coordinates of the left eye and the right eye, i.e. the 2D image positions.
[0312] In formula (23), M represents the internal parameter matrix, R represents the rotation matrix, d is the pupil distance, X c is the spatial coordinate of the eye center, k l 、k l is the intermediate variable, x l [0],x l [1] represents the horizontal and vertical coordinates of the left eye, x r [0],x r [1] represents the horizontal and vertical coordinates of the right eye. The number of independent variables in the formula is 3, and the irrelevant k is removed. l and k r Variables, the actual number of effective equations is 4, which is more than the number of independent variables. This embodiment proposes a robust solution method to estimate the Z-axis coordinate of the center point of the eye (i.e., the depth of the center point of the eye), as shown below:
[0313] Where Δx = x r -x l ,
[0314] In formula (24), X c [2] represents the depth of the eye center, x r is the plane coordinate of the right eye, x l is the plane coordinate of the left eye, x c Represents the plane coordinates of the center point of the eye, Δx[0] represents the X-axis component of the difference between the plane coordinates of the right eye and the left eye, Δx[1] represents the Y-axis component of the difference between the plane coordinates of the right eye and the left eye, x c [0] represents the X-axis coordinate of the eye center point, x c [1] represents the Y-axis coordinate of the eye center point, f x 、f y , x0, y0 represent the parameters in the internal parameter matrix, C z =cos(θ z ), C y =cos(θ y ), S y = sin(θ y ),θ yrepresents the yaw angle, θ z Indicates the roll angle.
[0315] The left eye depth and right eye depth, i.e. the Z-axis coordinates, are estimated using the following formula:
[0316] In formula (25), R is the rotation matrix constructed by the attitude angle, d represents the pupil distance, X c [2] represents the depth of the eye center, X l [2] represents the left eye depth, X r [2] indicates the right eye depth.
[0317] Then the spatial coordinates of the left and right eyes are calculated as follows:
[0318] In formula (26), represents the intermediate variable, x l [0],x l [1] represents the horizontal and vertical coordinates of the left eye, respectively. r [0],x r [1] represents the horizontal and vertical coordinates of the right eye, M represents the internal parameter matrix, X l [2] represents the left eye depth, X r [2] represents the right eye depth, X l Indicates the spatial coordinates of the left eye, X r Indicates the spatial coordinates of the right eye.
[0319] In some embodiments, to solve the problem of large jitter in the calculated eye depth or large jitter in the calculated spatial position, this embodiment provides any of the following jitter optimization solutions, as shown below:
[0320] Solution 1: Reduce the impact of jitter by filtering the plane coordinates of the eyes.
[0321] 1a) obtaining the plane coordinates of the eyes corresponding to the current image pair and the consecutive multiple frames of historical image pairs before the current image pair;
[0322] During implementation, each time a historical image pair is acquired, facial key point recognition is performed on the historical image pair to determine the plane coordinates of the eyes corresponding to each of the consecutive historical image pairs. Optionally, the information of the facial key points in the historical image pairs (the plane coordinates of the facial key points) can be saved. The facial key point information of the consecutive historical image pairs before the current image pair is acquired.
[0323] 1b) performing smoothing filtering on the plane coordinates of the eyes corresponding to each of the consecutive frames of historical image pairs, and determining the plane coordinates of the eyes in the current image pair based on the smoothed filtered coordinates;
[0324] 1c) determining the spatial position of the eye in the current image pair based on the plane coordinates of the eye in the current image pair and the depth of the eye.
[0325] Optionally, the jitter effect can be reduced by filtering the spatial coordinates. The details are as follows:
[0326] Obtain the spatial coordinates of the eyes corresponding to the current image pair and the multiple consecutive frames of historical image pairs before the current image pair; perform smoothing filtering on the spatial coordinates of the eyes corresponding to the multiple consecutive frames of historical image pairs, and determine the spatial position of the eyes in the current image pair based on the coordinates after smoothing filtering.
[0327] During implementation, the depths of the eyes corresponding to the consecutive multiple frames of historical images are determined, and the spatial positions of the eyes corresponding to the consecutive multiple frames of historical images are determined based on the depths of the eyes corresponding to the consecutive multiple frames of historical images.
[0328] It should be noted that the output of the key point regression model has jitter due to camera image noise, which is the main error source of the eye tracking algorithm. Ordinary key point denoising may be practical to average N consecutive frames, such as N frames of face key point data (L x (n,i),L y (n,i)), where i = 0, 1, ... 4, n = 0, 1, ... N, i represents the facial key point, and n represents the number of frames. One de-jittering principle is to directly calculate the average value of the key points; another method is to use the SG polynomial smoothing algorithm Landmark*W for point multiplication, where W = 0, 1, ... N is the correlation coefficient. The closer to the current point, the larger the coefficient. However, although this processing method stabilizes the key points, it introduces response delay. Therefore, to eliminate this delay, this embodiment uses collaborative filtering key point denoising. The specific steps are as follows:
[0329] (a) First save N frames of facial key point data (L x (n,i),L y (n,i)), calculate the average value of key points
[0330] Where i = 0, 1,…4, n = 0, 1,…N.
[0331] (b) Calculate the key point data of the current frame ((L x (n,N-1),L y (n,N-1)) and mean The deviation between x ,S y );
[0332] In formula (27), λ represents the ratio of the left and right eyes, L x_l (n,N-1) and L x_r (n, N-1) are the X coordinates of the left eye key point and the right eye key point respectively, L y_l (n,N-1) and L y_r (n, N-1) are the Y coordinates of the left eye key point and the right eye key point respectively. x Indicates the deviation in the X-axis direction, S y Indicates the deviation in the Y-axis direction.
[0333] (c) The facial key point data of the current frame after filtering is:
[0334] After the above processing and correction, the delayed response caused by the average value is eliminated, the filtering effect can be achieved, and feedback can be given in time according to the person's movement.
[0335] Solution 2: Reduce the impact of jitter by filtering the depth of the eyes.
[0336] 2a) obtaining a current image pair and a plurality of consecutive historical image pairs before the current image pair;
[0337] 2b) when it is determined that the first view and the second view in the continuous multi-frame historical image pair contain the same human face, determining the depth of each corresponding eye in the continuous multi-frame historical image pair using a binocular ranging algorithm;
[0338] 2c) performing smoothing filtering on the depths of the eyes corresponding to the respective frames of the historical image pairs, and determining the depth of the eyes of the face in the current image pair based on the depths of the eyes after the smoothing filtering.
[0339] Solution 3: Reduce the impact of jitter by filtering the depth and plane coordinates of the eyes. As shown in Figure 7, this embodiment provides an eye positioning process for reducing jitter. The specific implementation steps are as follows:
[0340] Step 700: Obtain the plane coordinates of the eyes corresponding to the current image pair and the consecutive multiple frames of historical image pairs before the current image pair;
[0341] In the implementation, a historical image pair is a video frame, which is an image pair containing two images. Facial key point recognition is performed on multiple consecutive historical image pairs to determine the plane coordinates of the eyes corresponding to each of the multiple consecutive historical image pairs.
[0342] Step 701: Perform smoothing filtering on the plane coordinates of the eyes corresponding to each of the consecutive frames of historical image pairs, and determine the plane coordinates of the eyes in the current image pair based on the smoothed filtered coordinates;
[0343] Step 702: When it is determined that the first view and the second view in the continuous multi-frame historical image pair contain the same face, a binocular ranging algorithm is used to determine the depth of each corresponding eye in the continuous multi-frame historical image pair;
[0344] Step 703: Smoothing filter the depths of the eyes corresponding to the consecutive frames of historical image pairs, and determining the depths of the eyes of the face in the current image pair based on the depths of the eyes after smoothing filter;
[0345] Step 704: Determine the spatial position of the eye in the current image pair according to the plane coordinates of the eye in the current image pair and the depth of the eye.
[0346] Optionally, the smoothing filtering method in this embodiment includes but is not limited to mean filtering, median filtering, Gaussian filtering, etc. For example, mean filtering can be used to save the depth z of each corresponding eye of N frames of historical images. li , z ri , calculate the average depth of the eyes corresponding to each of the N frames of historical images, and the average value is as follows:
[0347] in, represents the average left eye depth, Indicates the average depth of the right eye, z li Indicates the depth of the left eye in the i-th frame of the historical image pair, z ri represents the depth of the right eye in the i-th historical image pair. N represents the number of image pairs / frames.
[0348] As shown in FIG8 , this embodiment also provides a flowchart of an implementation of an eye positioning and tracking algorithm, which is specifically as follows:
[0349] Step 800: Acquire a video frame, the video frame including an image pair;
[0350] Step 801: Determine whether the first view in the image pair contains a human face. If so, proceed to step 802; otherwise, proceed to step 812.
[0351] Optionally, it may be determined first whether there is a human face in the first view, or it may be determined first whether there is a human face in the second view, and this embodiment does not impose too many restrictions on this.
[0352] Step 802: extract key points of the face in the first view and perform smoothing filtering;
[0353] During implementation, the extracted facial key points are plane coordinates, and smoothing filtering is performed on the plane coordinates of the extracted facial key points, wherein the plane coordinates of the facial key points include but are not limited to the plane coordinates of the eyes.
[0354] Step 803: Calculate the first spatial coordinates of the eye center point in the first view using a monocular ranging algorithm;
[0355] Step 804: Estimate the positions of facial key points in the second view using the optical center distance, and crop the facial region from the second view using the estimated positions of the facial key points.
[0356] Step 805: extracting facial key points in the face area;
[0357] Step 806: Determine whether it is a face based on the extracted facial key points. If yes, proceed to step 807; otherwise, proceed to step 811.
[0358] Step 807: Perform smoothing filtering on the facial key points in the extracted facial region, calculate the depth of the eyes based on the parallax, and perform smoothing filtering on the calculated depth of the eyes;
[0359] Step 808: Determine whether the deviation state is to the left. If so, execute step 809; otherwise, execute step 810.
[0360] Step 809: Calculate the left eye spatial coordinates using the left eye depth in the first view, and estimate the right eye spatial coordinates based on the left eye spatial coordinates, the fixed pupil distance, and the head posture angle;
[0361] Step 810: Calculate the right eye plane coordinates using the right eye depth in the second view, and estimate the left eye space coordinates based on the right eye space coordinates, the fixed pupil distance, and the head posture angle.
[0362] Step 811: Calculate the single view space coordinates of the eye in the first view using a monocular ranging algorithm;
[0363] Step 812: Determine whether the second view in the image pair contains a human face. If so, proceed to step 813; otherwise, proceed to step 800.
[0364] Step 813: extracting facial key points in the second view and performing smoothing filtering;
[0365] Step 814: Calculate the single-view space coordinates of the eye in the second view using a monocular ranging algorithm.
[0366] This embodiment uses a binocular ranging algorithm and the deviation of the face from the center of the binocular camera to locate the spatial position of the eyes, arranges images based on the located spatial position, and displays naked-eye 3D images. The binocular ranging algorithm avoids the problems of relying on pupil distance and small viewing angle when using only one camera. The spatial position of the eyes is located by the deviation of the face from the center of the binocular camera, thus solving the problem of large jitter caused by face deflection or distance from the camera center. Smoothing filtering can also be used to solve the problem of large jitter when calculating spatial coordinates or depth.
[0367] As shown in Figure 9, this embodiment also provides a schematic diagram of a test setup, which experimentally verifies the accuracy of the eye-space position determination algorithm proposed in this embodiment. In the test environment, a binocular camera was placed above the 3D display screen. The camera had a resolution of 640×480, a frame rate of 60 fps, a field of view of 80 degrees, and a baseline of 120 mm. The camera's external parameters were calibrated to ensure that the camera was completely parallel to the screen. A head model was placed in front of the screen, and its movement was controlled using an electromechanical device with exceptionally high precision, achieving a repeatability of less than 0.01 mm. During the experimental verification, the head model was placed at 25 test points at four different distances from the camera (0.5 m, 0.7 m, 1.0 m, and 1.5 m), five different orientations (-25°, -12°, 0°, 12°, and 25°), and at different heights from the camera (0.06, 0.14, 0.22, 0.30, and 0.38). In order to verify the effect of binoculars on people with different pupil distances, this embodiment uses head models with different pupil distances between 55-75 mm for testing.
[0368] As shown in Figure 10, the pupil position of each test point (i.e., the spatial position of the eye) is represented by the horizontal angle, vertical angle, and Z coordinate (depth). The accuracy of the algorithm of this embodiment is calculated by comparing the predicted pupil position with the actual position obtained by the head model through mechanical equipment. The error difference between the monocular ranging algorithm and the binocular ranging algorithm in eye tracking positioning at an interpupillary distance of 65mm is not large. The horizontal and vertical angles of the pupil are both less than 0.2°, and the maximum error of the Z coordinate is less than 4%, meeting the requirements of naked-eye 3D display devices for eye tracking.
[0369] To verify the extent of the impact of jitter on the viewer's experience, an experiment was conducted. Fixed eye coordinate data was used, random noise was added to the data, and the resulting 3D effect was observed. The experimental results showed that the viewer's experience was significantly degraded when the horizontal angle offset exceeded 0.02%, the y-axis offset exceeded 0.1%, and the z-axis offset exceeded 0.05%. Binocular jitter was 0.01° horizontal, 0.016° vertical, and 0.025% Z-coordinate. Monocular jitter was 0.0154° horizontal, 0.023° vertical, and 0.031% Z-coordinate.
[0370] Comparing the Z-axis accuracy of eye tracking for binocular and monocular vision at different interpupillary distances (IPDs) between 55 and 75 mm, it's clear that binocular vision is unaffected by interpupillary distance discrepancies. However, for the monocular ranging algorithm, the maximum interpupillary distance deviation can reach as high as 20%. For those with an interpupillary distance deviation of 65 mm, the monocular ranging algorithm becomes unfriendly, affecting stereoscopic perception. Different interpupillary distances only affect the Z coordinate value, leaving the vertical and horizontal angles unaffected. Regardless of interpupillary distance, the binocular ranging algorithm maintains Z-axis accuracy of less than 4%. However, for interpupillary distance deviations of 65 mm, the monocular ranging algorithm's Z-axis accuracy exceeds 4%, and the greater the deviation, the lower the accuracy.
[0371] Based on the same inventive concept, this embodiment further provides a display device, as shown in FIG11 , which includes a display unit 1100 and a control circuit 1101, wherein:
[0372] The display unit 1100 is configured to display content;
[0373] The control circuit 1101 includes a processor and a memory, wherein the memory is used to store a program executable by the processor, and the processor is used to read the program in the memory and perform the following steps:
[0374] Obtaining an image pair, and determining whether a first view and a second view in the image pair contain the same face, wherein the image pair is captured by a binocular camera;
[0375] When it is determined that the first view and the second view contain the same human face, determining the depth of the eyes of the human face using a binocular ranging algorithm, wherein the depth of the eyes represents the distance from the eyes to the display unit;
[0376] The spatial position of the eyes of the human face is determined based on the depth of the eyes and the image pair.
[0377] Based on the same inventive concept, as shown in FIG12 , the embodiment of the present disclosure further provides a naked-eye 3D display method, the specific implementation process of which is as follows:
[0378] Step 1200: Acquire an image pair, where the image pair includes a first view and a second view, the first view and the second view include a face, and the image pair is captured by a binocular camera;
[0379] Step 1201: Determine the depth of the eyes of the face using a binocular ranging algorithm based on the first view and the second view; wherein the depth of the eyes represents the distance from the eyes to a display unit;
[0380] Step 1202: Determine the deviation of the face from the center position, where the center position is determined based on the center of the binocular camera that captures the face.
[0381] Step 1203: Determine the spatial position of the eyes of the human face according to the depth and deviation state of the eyes, and display a stereoscopic image based on the spatial position of the eyes.
[0382] Based on the same inventive concept, as shown in FIG13 , the embodiment of the present disclosure further provides an eye positioning method, the specific implementation process of which is as follows:
[0383] Step 1300: Acquire an image pair, and determine whether a first view and a second view in the image pair contain the same face, wherein the image pair is captured by a binocular camera;
[0384] Step 1301: When it is determined that the first view and the second view contain the same human face, a binocular ranging algorithm is used to determine the depth of the eyes of the human face, where the depth of the eyes represents the distance from the eyes to the display unit.
[0385] Step 1302: Determine the spatial position of the eyes of the face based on the depth of the eyes and the image pair.
[0386] Based on the same inventive concept, an embodiment of the present disclosure provides an electronic device that can implement the functions of the naked-eye 3D display method or eye positioning method discussed above. Please refer to Figure 14. The device includes a processor 1401 and a memory 1402. The memory 1402 is used to store program instructions; the processor 1401 is used to call the program instructions stored in the memory 1402, and execute the steps included in any document generation method in the above embodiments according to the obtained program instructions.
[0387] The embodiment of the present disclosure does not limit the specific connection medium between the memory 1402 and the processor 1401. For example, the memory 1402 and the processor 1401 are connected via a bus, which can be divided into an address bus, a data bus, a control bus, and the like.
[0388] The memory 1402 may include read-only memory (ROM) and random access memory (RAM), and may also include non-volatile memory (NVM), such as at least one disk storage. Alternatively, the memory may be at least one storage device located remotely from the processor.
[0389] The processor 1401 may be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it may also be a digital signal processing (DSP), an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic device, a discrete gate or transistor logic device, a discrete hardware component, etc.
[0390] Based on the same inventive concept, as shown in FIG15 , an embodiment of the present disclosure provides a naked-eye 3D display device, which includes:
[0391] An image acquisition module 1500 is configured to acquire an image pair, wherein the image pair includes a first view and a second view, wherein the first view and the second view include a human face, and the image pair is captured by a binocular camera;
[0392] A depth determination module 1501 is configured to determine the depth of the eyes of a human face using a binocular ranging algorithm based on the first view and the second view; wherein the depth of the eyes represents the distance from the eyes to a display unit;
[0393] a deviation determination module 1502 for determining a deviation state of a face relative to a center position, where the center position is determined based on the center of a binocular camera capturing the face;
[0394] The position determination module 1503 is used to determine the spatial position of the eyes of the face according to the depth and deviation state of the eyes, and display a three-dimensional image based on the spatial position of the eyes.
[0395] Based on the same inventive concept, as shown in FIG16 , an embodiment of the present disclosure further provides an eye positioning device, which includes:
[0396] An image determination module 1600 is configured to obtain an image pair and determine whether a first view and a second view in the image pair contain the same face, wherein the image pair is captured by a binocular camera;
[0397] a depth determination module 1601 configured to determine the depth of the eyes of the face using a binocular ranging algorithm when it is determined that the first view and the second view contain the same face, wherein the depth of the eyes represents the distance from the eyes to the display unit;
[0398] The position determination module 1602 is configured to determine the spatial position of the eyes of the face according to the depth of the eyes and the image pair.
[0399] Based on the same inventive concept, an embodiment of the present disclosure provides an electronic device, as shown in FIG17 . The electronic device includes a processor 1700 and a memory 1701. The memory 1701 is configured to store programs executable by the processor 1700. The processor 1700 is configured to read the programs in the memory 1701 and execute the following steps:
[0400] Acquire an image pair, where the image pair includes a first view and a second view, the first view and the second view include a human face, and the image pair is captured by a binocular camera;
[0401] Determining the depth of the eyes of the human face using a binocular ranging algorithm based on the first view and the second view; wherein the depth of the eyes represents the distance from the eyes to the display unit;
[0402] Determining a deviation state of the face relative to a center position, wherein the center position is determined based on the center of a binocular camera capturing the face;
[0403] Determine the spatial position of the eyes of a face based on their depth and deviation.
[0404] Based on the same inventive concept, embodiments of the present disclosure provide a computer storage medium comprising computer program code. When the computer program code is executed on a computer, the computer executes any of the aforementioned naked-eye 3D display methods or eye positioning methods. Because the principles underlying the problems solved by the aforementioned computer storage medium are similar to those of the naked-eye 3D display method or eye positioning method, the implementation of the aforementioned computer storage medium can be referenced to the implementation of the method, and any repetitions will not be repeated.
[0405] In a specific implementation process, computer storage media may include: Universal Serial Bus Flash Drive (USB), mobile hard disk, Read-Only Memory (ROM), Random Access Memory (RAM), magnetic disk or optical disk, and other storage media that can store program code.
[0406] Based on the same inventive concept, embodiments of the present disclosure further provide a computer program product, comprising: computer program code, which, when executed on a computer, causes the computer to execute any of the above-described naked-eye 3D display methods or eye positioning methods. Because the principles underlying the problems solved by the above-described computer program products are similar to those of the naked-eye 3D display methods or eye positioning methods, the implementation of the above-described computer program products can be referred to as the implementation of the methods, and any repetitions will not be repeated.
[0407] The computer program product can employ any combination of one or more readable media. The readable medium can be a readable signal medium or a readable storage medium. The readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or component, or any combination thereof. More specific examples (non-exhaustive list) of readable storage media include: an electrical connection with one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof.
[0408] Those skilled in the art will appreciate that the embodiments of the present disclosure may be provided as methods, systems, or computer program products. Therefore, the present disclosure may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Furthermore, the present disclosure may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage and optical storage, etc.) containing computer-usable program code.
[0409] The present disclosure is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present disclosure. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device produce a device for implementing the functions specified in one or more processes in the flowchart and / or one or more boxes in the block diagram.
[0410] These computer program instructions may also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce a manufactured product including an instruction device that implements the functions specified in one or more processes in the flowchart and / or one or more boxes in the block diagram.
[0411] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, so that the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one or more processes in the flowchart and / or one or more boxes in the block diagram.
[0412] Obviously, those skilled in the art may make various changes and modifications to the present disclosure without departing from the spirit and scope of the present disclosure. Thus, if these modifications and variations of the present disclosure fall within the scope of the claims of the present disclosure and their equivalents, the present disclosure is intended to include these modifications and variations.
Claims
1. A display device, wherein, the display device includes a display unit and a control circuit, wherein: the display unit is configured to display a stereoscopic image based on parallax; the control circuit includes a processor and a memory, the memory is used to store programs executable by the processor, and the processor is used to read the programs in the memory and execute the following steps: obtain an image pair, the image pair includes a first view and a second view, the first view and the second view include a human face, and the image pair is obtained by shooting with a binocular camera; based on the first view and the second view, use the binocular ranging algorithm to determine the depth of the eyes of the human face; wherein the depth of the eyes represents the distance from the eyes to the display unit; determine the deviation state of the human face relative to the central position, and the central position is determined according to the center of the binocular camera that shoots the human face; according to the depth of the eyes and the deviation state, determine the spatial position of the eyes of the human face, and display a stereoscopic image based on the spatial position of the eyes.
2. The display device according to claim 1, wherein, the processor is specifically configured to execute: select a view corresponding to the deviation state from the first view and the second view; according to the depth of the eyes and the selected view, determine the spatial position of the eyes of the human face.
3. The display device according to claim 2, wherein, the processor is specifically configured to determine the view corresponding to the deviation state in the following manner: when the deviation state is that the human face is deviated to the left relative to the central position, determine that the view corresponding to the deviation state is the first view, and the first view is taken by the left camera of the binocular camera; when the deviation state is that the human face is deviated to the right relative to the central position, determine that the view corresponding to the deviation state is the second view, and the second view is taken by the right camera of the binocular camera.
4. The display device according to claim 2, wherein, the processor is specifically configured to execute: according to the depth of the eyes, determine the single-view spatial coordinates of one eye in the selected view, wherein the one eye includes the left eye or the right eye; estimate the single-view spatial coordinates of the other eye by using the single-view spatial coordinates of one eye, and determine the spatial position of the eyes of the human face according to the single-view spatial coordinates of the left eye and the right eye.
5. The display device according to claim 1, wherein, the processor is specifically further configured to execute: judge whether the first view and the second view include the same human face; the processor is specifically configured to: when the first view and the second view include the same human face, based on the first view and the second view, use the binocular ranging algorithm to determine the depth of the eyes of the human face.
6. The display device according to claim 5, wherein, the processor is specifically configured to execute: perform similarity matching on the human faces in the first view and the second view, and judge whether the first view and the second view include the same human face according to the matching result.
7. The display device according to claim 5, wherein, the processor is specifically configured to execute: judge whether the first view and the second view include the same human face according to whether the human faces in the first view and the second view have a position mapping relationship.
8. The display device according to claim 5, wherein, the processor is specifically configured to perform: acquire the first spatial coordinates of the eye center point corresponding to the face in the first view, and estimate the face area of the face in the second view according to the first spatial coordinates, the optical center distance of the binocular camera, and the internal parameter matrix; determine whether the first view and the second view contain the same face according to whether a face is detected within the face area in the second view; or, acquire the second spatial coordinates of the eye center point corresponding to the face in the second view, and estimate the face area of the face in the first view according to the second spatial coordinates, the optical center distance of the binocular camera, and the internal parameter matrix; determine whether the first view and the second view contain the same face according to whether a face is detected within the face area in the first view; wherein the eye center point is located at the midpoint of the line connecting the left eye and the right eye, and the optical center distance represents the distance between the centers of the two camera lenses of the binocular camera.
9. The display device according to claim 5, wherein, the processor is specifically configured to perform: determine the field of view area and the single-view spatial coordinates of the eyes in the first view, and determine whether the first view and the second view contain the same face according to the relationship between the single-view spatial coordinates of the eyes in the first view and the field of view area; or, determine the field of view area and the single-view spatial coordinates of the eyes in the second view, and determine whether the first view and the second view contain the same face according to the relationship between the single-view spatial coordinates of the eyes in the second view and the field of view area.
10. The display device according to claim 9, wherein, the processor is specifically configured to determine the field of view area in the following manner: determine the depth of the eye center point in the first view according to the monocular ranging algorithm, and determine the field of view area according to the depth of the eye center point in the first view, the field of view angle, and the optical center distance; or, determine the depth of the eye center point in the second view according to the monocular ranging algorithm, and determine the field of view area according to the depth of the eye center point in the second view, the field of view angle, and the optical center distance; wherein the eye center point is located at the midpoint of the line connecting the left eye and the right eye; the optical center distance represents the distance between the centers of the two camera lenses of the binocular camera, and the field of view angle represents the horizontal field of view angle captured by a single camera in the binocular camera.
11. The display device according to claim 5, wherein, the processor is further specifically configured to perform: when it is determined that the first view and the second view do not contain the same face, use the monocular ranging algorithm to determine the depth of the eyes; determine the spatial position of the eyes of the face according to the depth of the eyes determined by the monocular ranging algorithm.
12. The display device according to claim 1, wherein, the processor is specifically configured to perform: use the binocular ranging algorithm to determine the depth of the eyes according to the parallax, where the parallax represents the position deviation in the horizontal direction of the eyes of the same person's face in the first view and the second view.
13. The display device according to claim 12, wherein, the parallax and the depth of the eyes vary inversely.
14. The display device according to claim 12, wherein, the processor is specifically configured to perform: Obtain the planar coordinates of the left and right eyes in the first view, and the planar coordinates of the left and right eyes in the second view; Determine the left-eye parallax based on the planar coordinates of the left eye in the first view and the second view, and determine the left-eye depth based on the left-eye parallax, focal length, and optical center distance; the left-eye parallax represents the horizontal position deviation of the left eye of the same human face in the first view and the second view; the optical center distance represents the distance between the centers of the two camera lenses of the binocular camera; Determine the right-eye parallax based on the planar coordinates of the right eye in the first view and the second view, and determine the right-eye depth based on the right-eye parallax, focal length, and optical center distance; The right-eye parallax represents the horizontal position deviation of the right eye of the same human face in the first view and the second view; Determine the depth of the eyes based on the left-eye depth and the right-eye depth.
15. The display device according to claim 1, wherein, the processor is specifically configured to execute: Use the depth of the eyes to respectively determine the single-view spatial coordinates of the eyes in the first view and the second view; Determine the deviation state of the human face relative to the central position according to the single-view spatial coordinates of the eyes in the first view and the second view.
16. The display device according to claim 15, wherein, the processor is specifically configured to execute: Determine the first spatial coordinates of the eye center point in the first view according to the spatial coordinates of the left and right eyes in the first view, and determine the second spatial coordinates of the eye center point in the second view according to the spatial coordinates of the left and right eyes in the second view; wherein the eye center point is located at the midpoint of the line connecting the left and right eyes; Determine the deviation state of the human face relative to the central position according to the average value of the first spatial coordinates and the second spatial coordinates.
17. The display device according to claim 1, wherein, the processor is specifically further configured to execute: Use the depth of the eyes to respectively determine the single-view spatial coordinates of the eyes in the first view and the second view; determine the weights corresponding to the first view and the second view according to the selected view, wherein the weights are determined based on the deviation state; Use the weights corresponding to the first view and the second view to perform weighted averaging on the single-view spatial coordinates of the eyes in the first view and the second view, and determine the spatial position of the eyes of the human face according to the weighted average value.
18. The display device according to claim 1, wherein, the processor is specifically configured to execute: Correct the fixed pupil distance to obtain a corrected pupil distance according to the depth of the eyes determined by the binocular ranging algorithm; Use the monocular ranging algorithm to determine the corrected depth of the eyes of the human face according to the corrected pupil distance; Determine the spatial position of the eyes of the human face according to the corrected depth of the eyes of the human face and the deviation state.
19. The display device according to claim 18, wherein, the processor is specifically configured to execute: Use the monocular ranging algorithm to determine the first depth of the eyes; Use the binocular ranging algorithm to determine the depth of the eyes; Correct the fixed pupil distance to obtain a corrected pupil distance according to the first depth and the depth of the eyes determined by the binocular ranging algorithm.
20. The display device according to claim 19, wherein, The corrected interpupillary distance varies directly with the depth of the eyes determined by the binocular ranging algorithm, and the corrected interpupillary distance varies inversely with the first depth.
21. The display device according to claim 18, wherein, the processor is specifically configured to execute: using a monocular ranging algorithm, determining a first corrected depth of the eyes of the face in the first view and a second corrected depth of the eyes of the face in the second view according to the corrected interpupillary distance; determining a corrected depth of the eyes of the face according to the average value of the first corrected depth and the second corrected depth.
22. The display device according to any one of claims 1 to 21, wherein, the processor is specifically further configured to determine the spatial position of the eyes of the face in the following manner: acquiring a current image pair and the planar coordinates of the eyes corresponding to each of a plurality of consecutive historical image pairs before the current image pair; performing smoothing filtering on the planar coordinates of the eyes corresponding to each of the plurality of consecutive historical image pairs, and determining the planar coordinates of the eyes in the current image pair according to the coordinates after smoothing filtering; determining the spatial position of the eyes in the current image pair according to the planar coordinates of the eyes in the current image pair, the depth of the eyes, and the deviation state.
23. The display device according to any one of claims 1 to 21, wherein, the processor is specifically further configured to determine the depth of the eyes of the face in the following manner: acquiring a current image pair and a plurality of consecutive historical image pairs before the current image pair; when it is determined that the first view and the second view in the plurality of consecutive historical image pairs contain the same face, using a binocular ranging algorithm to determine the depth of the eyes corresponding to each of the plurality of consecutive historical image pairs; performing smoothing filtering on the depth of the eyes corresponding to each of the plurality of consecutive historical image pairs, and determining the depth of the eyes of the face in the current image pair according to the depth of the eyes after smoothing filtering.
24. A display device, wherein, the display device includes a display unit and a control circuit, wherein: the display unit is configured to perform the display of content; the control circuit includes a processor and a memory, the memory is used to store programs executable by the processor, and the processor is used to read the programs in the memory and execute the following steps: acquiring an image pair, and determining whether the first view and the second view in the image pair contain the same face, the image pair being obtained by shooting with a binocular camera; when it is determined that the first view and the second view contain the same face, using a binocular ranging algorithm to determine the depth of the eyes of the face, wherein the depth of the eyes represents the distance from the eyes to the display unit; determining the spatial position of the eyes of the face according to the depth of the eyes and the image pair.
25. A naked-eye 3D display method, wherein, the method includes: acquiring an image pair, the image pair including a first view and a second view, the first view and the second view including a face, the image pair being obtained by shooting with a binocular camera; based on the first view and the second view, using a binocular ranging algorithm to determine the depth of the eyes of the face; wherein the depth of the eyes represents the distance from the eyes to the display unit; Determine the deviation state of the human face relative to the central position, where the central position is determined according to the center of the binocular camera that captures the human face; Determine the spatial position of the eyes of the human face based on the depth and deviation state of the eyes, and display a stereoscopic image based on the spatial position of the eyes.
26. An eye positioning method, wherein, the method includes: Obtain an image pair, and determine whether the first view and the second view in the image pair contain the same human face. The image pair is obtained by a binocular camera; When it is determined that the first view and the second view contain the same human face, use the binocular ranging algorithm to determine the depth of the eyes of the human face, where the depth of the eyes represents the distance from the eyes to the display unit; Determine the spatial position of the eyes of the human face according to the depth of the eyes and the image pair.
27. An electronic device, wherein, the electronic device includes a processor and a memory. The memory is used to store a program executable by the processor, and the processor is used to read the program in the memory and execute the following steps: Obtain an image pair, where the image pair includes a first view and a second view, and the first view and the second view contain a human face. The image pair is obtained by a binocular camera; Based on the first view and the second view, use the binocular ranging algorithm to determine the depth of the eyes of the human face; where the depth of the eyes represents the distance from the eyes to the display unit; Determine the deviation state of the human face relative to the central position, where the central position is determined according to the center of the binocular camera that captures the human face; Determine the spatial position of the eyes of the human face according to the depth of the eyes and the deviation state.
28. An electronic device, wherein, the electronic device includes a processor and a memory. The memory is used to store a program executable by the processor, and the processor is used to read the program in the memory and execute the steps of the method according to claim 25 or 26.
29. A computer storage medium, on which a computer program is stored, wherein, when the program is executed by a processor, it implements the steps of the method according to claim 25 or 26.
Citation Information
Patent Citations
Method, device, device and storage medium for calibrating spatial position of human eye
CN109040736A
Naked eye 3D crosstalk removing method and system, storage medium and electronic device
CN110381305A
Human eye tracking method for naked eye 3D display
CN115733967A
Image interaction system, method for detecting finger position, stereo display system and control method of stereo display
US20140176676A1