Camera extrinsic parameter calibration method and system, and electronic device
By calibrating the stereo model to simulate the face and automatically calibrate the camera external parameters in the naked eye 3D technology, the problems of low accuracy and dependence on the calibration plate in the existing technology are solved, and high-precision camera external parameters calibration and stable eye positioning are achieved.
Patent Information
- Application Number
- PCT/CN2025/070780
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-01-22
- Filing Date
- 2025-01-06
- Publication Date
- 2025-07-31
AI Technical Summary
In the prior art, the accuracy of camera external parameter calibration in naked-eye 3D technology is not high, and it is necessary to use the calibration plate, so fully automated and high-precision external parameter calibration cannot be achieved.
By simulating the face with a calibrated stereo model, the predicted spatial coordinates of the eye of the stereo model are automatically calibrated, including controlling the calibrated stereo model to move to the zero point, collecting images, and adjusting the spatial coordinates to determine the camera external parameters.
It realizes fully automated camera external parameter calibration without calibration plates, improves calibration accuracy and stability, and ensures the accuracy and stability of eyeball positioning in naked-eye 3D technology.
Smart Images

Figure CN2025070780_31072025_PF_FP_ABST
Abstract
Description
Camera extrinsic parameter calibration method, system and electronic equipment
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS
[0002] This application claims priority to the Chinese patent application filed with the State Intellectual Property Office of the People's Republic of China on January 22, 2024, with application number 202410089138.7 and invention name “A camera extrinsic parameter calibration method, system and electronic device”, the entire contents of which are incorporated by reference into this application. Technical Field
[0003] The present disclosure relates to the field of eye positioning technology, and in particular to a camera extrinsic parameter calibration method, system and electronic equipment. Background Art
[0004] It's well known that 3D images deliver a powerful visual feast. To experience the 3D effect, users can wear polarized glasses that use polarized light. This means their left eye only sees the image projected by the left camera, and their right eye only sees the image projected by the right camera, resulting in a three-dimensional image. Autostereoscopy is a general term for technologies that achieve stereoscopic visual effects without the aid of external tools such as polarized glasses. Autostereoscopy allows users to directly experience the 3D effect without the need for other external tools. Autostereoscopy uses eye tracking, offering an unlimited viewing angle and adjusting the display based on a user's location. Users can watch while moving within a certain range, but this requires both accurate and stable eye positioning and tracking methods.
[0005] The accuracy of eye positioning is not only related to the eye positioning algorithm, but the accuracy of camera extrinsic calibration is also a key factor. Summary of the Invention
[0006] The present disclosure provides a camera extrinsic parameter calibration method, system and electronic device, which automatically calibrate the camera extrinsic parameters without the aid of a calibration board.
[0007] In a first aspect, an embodiment of the present disclosure provides a camera extrinsic parameter calibration method, the method comprising:
[0008] Controlling the calibration three-dimensional model to move to a zero position, wherein the calibration three-dimensional model at the zero position faces the display screen, and a line connecting the first center point of the calibration three-dimensional model and the center point of the display screen is perpendicular to the plane of the display screen;
[0009] Capturing a calibration stereo model image through a camera of the display screen, and determining predicted spatial coordinates of the eye of the calibration stereo model according to the calibration stereo model image;
[0010] The camera extrinsic parameters of the display screen's camera are determined based on the predicted spatial coordinates of the eye of the calibrated stereo model.
[0011] The camera extrinsic calibration method provided by this disclosure automatically calibrates camera extrinsics without the need for a calibration plate. This method simulates a human face using a calibrated stereo model and calibrates the camera extrinsics by predicting the spatial coordinates of the eyes in the calibrated stereo model.
[0012] In a second aspect, an embodiment of the present disclosure provides a camera extrinsic parameter calibration system, the system comprising a display screen, a calibration fixture, and a calibration stereo model, wherein the display screen and the calibration fixture establish a communication connection, wherein:
[0013] The calibration fixture is configured to control the calibration stereo model to move to a zero position, wherein the calibration stereo model at the zero position faces the display screen, and a line connecting the first center point of the calibration stereo model and the center point of the display screen is perpendicular to the plane of the display screen;
[0014] The display screen is configured to use a camera to capture a calibration stereo model image, determine the predicted spatial coordinates of the calibration stereo model eyes based on the calibration stereo model image; and determine the camera extrinsic parameters of the camera of the display screen based on the predicted spatial coordinates of the calibration stereo model eyes.
[0015] In a third aspect, an embodiment of the present disclosure further provides an electronic device, comprising a processor and a memory, wherein the memory is configured to store a program executable by the processor, and the processor is configured to read the program in the memory and perform the following steps:
[0016] Controlling the calibration three-dimensional model to move to a zero position, wherein the calibration three-dimensional model at the zero position faces the display screen, and a line connecting the first center point of the calibration three-dimensional model and the center point of the display screen is perpendicular to the plane of the display screen;
[0017] Capturing a calibration stereo model image through a camera of the display screen, and determining predicted spatial coordinates of the eye of the calibration stereo model according to the calibration stereo model image;
[0018] The camera extrinsic parameters of the display screen's camera are determined based on the predicted spatial coordinates of the eye of the calibrated stereo model.
[0019] In a fourth aspect, an embodiment of the present disclosure further provides a camera extrinsic parameter calibration device, comprising:
[0020] a zero setting module, configured to control the calibration three-dimensional model to move to a zero position, wherein the calibration three-dimensional model at the zero position faces the display screen, and a line connecting the first center point of the calibration three-dimensional model and the center point of the display screen is perpendicular to the plane of the display screen;
[0021] A prediction module, configured to capture a calibration stereo model image through a camera of the display screen, and determine predicted spatial coordinates of the eyes of the calibration stereo model based on the calibration stereo model image;
[0022] The calibration module is used to determine the camera extrinsic parameters of the display screen camera according to the predicted spatial coordinates of the eyes of the calibration stereo model.
[0023] In a fifth aspect, an embodiment of the present disclosure further provides a computer storage medium on which a computer program is stored, which, when executed by a processor, is used to implement the steps of the method described in any one of the above-mentioned first aspects.
[0024] In a sixth aspect, the present disclosure provides a computer program product, comprising: a computer program code, which, when executed on a computer, enables the computer to execute any one of the methods described in the first aspect.
[0025] These and other aspects of the present disclosure will become more readily apparent from the following description of the embodiments. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] In order to more clearly illustrate the technical solutions in the embodiments of the present disclosure, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present disclosure. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0027] FIG1 is a flowchart of an implementation method of a camera extrinsic parameter calibration method provided by an embodiment of the present disclosure;
[0028] FIG2 is a schematic diagram of a display screen image during zero point calibration according to an embodiment of the present disclosure;
[0029] FIG3 is a schematic diagram of a zero point position of a calibrated stereo model provided by an embodiment of the present disclosure;
[0030] 4A-4D are schematic diagrams of images displayed on a display screen according to an embodiment of the present disclosure;
[0031] FIG5 is a flowchart of a method for implementing camera extrinsic calibration according to an embodiment of the present disclosure;
[0032] FIG6 is a schematic diagram of a conversion relationship between data sets provided by an embodiment of the present disclosure;
[0033] FIG7 is a schematic diagram of a method for determining a rotation relationship provided by an embodiment of the present disclosure;
[0034] FIG8 is a schematic diagram of the structure of an external parameter calibration system provided by an embodiment of the present disclosure;
[0035] FIG9 is a flowchart of an implementation of camera extrinsic calibration according to an embodiment of the present disclosure;
[0036] FIG10 is a schematic diagram of a camera extrinsic calibration system provided by an embodiment of the present disclosure;
[0037] FIG11 is a schematic diagram of an electronic device provided by an embodiment of the present disclosure;
[0038] FIG12 is a schematic diagram of a camera extrinsic parameter calibration device provided in an embodiment of the present disclosure. DETAILED DESCRIPTION
[0039] To make the objectives, technical solutions, and advantages of the present disclosure more clear, the present disclosure will be further described in detail below with reference to the accompanying drawings. It is apparent that the embodiments described are only a portion of the embodiments of the present disclosure, rather than all of them. All other embodiments derived by persons of ordinary skill in the art based on the embodiments of the present disclosure without creative effort are intended to fall within the scope of protection of the present disclosure.
[0040] In the embodiments of the present disclosure, the term "and / or" describes the association relationship between associated objects, indicating that three relationships can exist. For example, A and / or B can represent three situations: A exists alone, A and B exist simultaneously, and B exists alone. The character " / " generally indicates that the associated objects are in an "or" relationship.
[0041] The application scenarios described in the embodiments of the present disclosure are intended to more clearly illustrate the technical solutions of the embodiments of the present disclosure and do not constitute a limitation on the technical solutions provided by the embodiments of the present disclosure. Persons skilled in the art will appreciate that, as new application scenarios emerge, the technical solutions provided by the embodiments of the present disclosure will also be applicable to similar technical problems. In the description of the present disclosure, unless otherwise specified, "multiple" means two or more.
[0042] Before introducing the camera extrinsic parameter calibration method provided by the embodiment of the present disclosure, in order to facilitate understanding, the technical background of the embodiment of the present disclosure is first introduced in detail below.
[0043] It's well known that 3D images deliver a powerful visual feast. To experience the 3D effect, users can wear polarized glasses that use polarized light. This means their left eye only sees the image projected by the left camera, and their right eye only sees the image projected by the right camera, resulting in a three-dimensional image. Autostereoscopy is a general term for technologies that achieve stereoscopic visual effects without the aid of external tools such as polarized glasses. Autostereoscopy allows users to directly experience the 3D effect without the need for other external tools. Autostereoscopy uses eye tracking, offering an unlimited viewing angle and adjusting the display based on a user's location. Users can watch while moving within a certain range, but this requires both accurate and stable eye positioning and tracking methods.
[0044] In addition to being related to the eye positioning algorithm, the accuracy of camera extrinsic calibration is also a key factor in the accuracy of eye positioning. Camera extrinsic calibration refers to determining the camera's position, including rotation angle and translation. It maps the camera from 3D space to 2D image space. In order to perform camera extrinsic calibration, a calibration plate is currently required. Generally speaking, the steps of camera extrinsic calibration include: (1) Installing the calibration plate in 3D space and measuring the accurate coordinates of feature points. (2) Using the camera to capture the image of the calibration plate and calculate the camera intrinsic parameters. (3) Calculating the camera extrinsic parameters, including the rotation matrix and translation vector. (4) Evaluating the calibration results to ensure accuracy. Since it cannot be guaranteed that the calibration plate and the display are completely parallel, the accuracy of camera extrinsic calibration is currently not high and it is not fully automatic.
[0045] Based on this, the present disclosure provides a camera extrinsic parameter calibration method that automatically calibrates camera extrinsics without the need for a calibration plate. This method uses a calibrated stereo model to simulate a human face and calibrates the camera extrinsics by calibrating the predicted spatial coordinates of the stereo model's eyes.
[0046] It should be noted that the display screen in this embodiment includes, but is not limited to, a glasses-free 3D display screen. It consists of four components: a 3D stereoscopic reality terminal, playback software, production software, and application technology. This system integrates modern high-tech technologies, including optics, photography, computers, automatic control, software, and 3D animation production, creating a cross-stereoscopic reality system. The glasses-free 3D display screen in this embodiment provides a display effect where objects appear to be both prominent and hidden within the image. Furthermore, the images are vibrantly colored, layered, vivid, and lifelike, creating a truly three-dimensional image.
[0047] The technical principles of glasses-free 3D displays include but are not limited to: light barrier technology, lenticular lens technology, and light field display lamps. Light barrier 3D technology utilizes a switchable LCD screen, a polarizing film, and a polymer liquid crystal layer. The liquid crystal layer and polarizing film create a series of vertical stripes oriented at 90 degrees. These stripes are tens of microns wide, and light passing through them forms a vertical, fine-stripe pattern known as a "parallax barrier." This technology utilizes a parallax barrier placed between the backlight module and the LCD panel. In stereoscopic display mode, when the image intended for the left eye is displayed on the LCD screen, the opaque stripes block the view for the right eye. Similarly, when the image intended for the right eye is displayed on the LCD screen, the opaque stripes block the view for the left eye. By separating the left and right eye's visible images, the viewer perceives a 3D image. Lenticular lens technology, also known as micro-lenticular 3D technology, places the LCD screen's image plane at the focal plane of the lens. This allows each pixel in the image beneath each lenticular lens to be divided into several sub-pixels, allowing the lens to project each sub-pixel in different directions. So when each eye looks at the display from different angles, it sees different sub-pixels. Because lenticular lens technology doesn't affect screen brightness like light barriers, it produces better display quality. Light field displays use dense fields of light to produce full-color, real-time 3D video, eliminating the need for glasses. This method of creating a 3D display allows several people to simultaneously view a virtual scene that resembles a real 3D object.
[0048] As shown in FIG1 , the implementation process of a camera extrinsic parameter calibration method provided in this embodiment is as follows:
[0049] Step 100: Control the calibration 3D model to move to a zero position, wherein the calibration 3D model at the zero position faces the display screen, and a line connecting a first center point of the calibration 3D model and a center point of the display screen is perpendicular to a plane where the display screen is located;
[0050] Optionally, the calibration three-dimensional model in this embodiment includes but is not limited to a head model, and the first center point of the calibration three-dimensional model includes but is not limited to the center point of the eyes of the head model.
[0051] During implementation, the calibration stereo model is used to simulate the movement of the human face. First, the calibration stereo model is controlled to face the display screen and moved to the zero point so that the line connecting the first center point of the calibration stereo model and the center point of the display screen is perpendicular to the plane of the display screen.
[0052] During implementation, the calibration stereo model is first roughly moved to a position where its eyes are directly opposite the center of the display screen. To ensure that the calibration stereo model moves with the center of the display screen as its origin, a hole can be opened at the center of the eye of the calibration stereo model. A laser is positioned behind the calibration stereo model so that the laser passes through the hole at the center of the eye of the calibration stereo model, forming a laser spot on the display screen. As shown in FIG2 , the display screen image provided by this embodiment is at the zero point, where the center of the image displayed on the display screen is a white area, and the origin is the laser projection point. The offset between the laser projection point and the center of the white area is calculated, and the calibration stereo model (head model) is controlled to move to the zero point based on this offset. The calibration stereo model can move in any direction along the X, Y, and Z axes in a defined world coordinate system. The world coordinate system is created according to the right-hand rule with the center of the display screen as its origin. FIG3 shows a schematic diagram of the zero point of the calibration stereo model provided by this embodiment, where the laser passes through the calibration stereo model and appears at the center of the display screen. The laser, the first center point of the calibration stereo model (the center of the eye), and the center of the display screen are collinear.
[0053] Step 101: Capture a calibration stereo model image through a camera on a display screen, and determine predicted spatial coordinates of the eyes of the calibration stereo model based on the calibration stereo model image;
[0054] During implementation, after the calibrated stereo model stops moving, the display screen uses its own camera to shoot the calibrated stereo model to obtain a calibrated stereo model image, and uses eye tracking technology to locate the position of the calibrated stereo model's eyes, thereby obtaining the predicted spatial coordinates of the calibrated stereo model's eyes.
[0055] Eye tracking uses image processing technology to locate the pupil, obtain the coordinates of the pupil center, and calculate the eye's gaze point using an algorithm. Monocular visual positioning is based on camera imaging principles and image processing algorithms. Camera imaging involves focusing light from a scene onto an image sensor through an optical lens, forming a digital image. Image processing algorithms process digital images to extract feature information, such as edges, corners, and color, to locate and track the target object. For example, a monocular ranging algorithm can be used to process images of a calibration stereo model to determine the predicted spatial coordinates of the model's eyes.
[0056] Step 102: Determine the camera extrinsic parameters of the display screen camera based on the predicted spatial coordinates of the eyes of the calibrated stereo model.
[0057] This embodiment provides at least two calibration methods when calibrating camera extrinsic parameters:
[0058] In the first calibration method, the calibration stereo model is fixed at the zero point and does not move.
[0059] In some embodiments, in this manner, the camera extrinsic parameters are determined as follows:
[0060] Process a) adjusting the predicted spatial coordinates multiple times, rearranging pixels of an image to be displayed on a display screen using the adjusted spatial coordinates, and determining the image quality after the pixel rearrangement;
[0061] During implementation, after the calibration stereo model is moved to the zero position, the calibration stereo model is fixed and does not move. First, the predicted spatial coordinates of the eyes of the current calibration stereo model are determined based on the calibration stereo model image. On this basis, the predicted spatial coordinates are adjusted multiple times. The image quality corresponding to each adjusted spatial coordinate obtained after adjustment is selected, and the adjusted spatial coordinate corresponding to the best image quality is used as the optimal position for viewing the calibration stereo model. The translation amount of the camera extrinsic parameter is determined based on the adjusted spatial coordinates and the predicted spatial coordinates of the optimal position. The specific adjustment process is as follows:
[0062] According to a preset step size, the coordinate components of the predicted space coordinates on the X-axis and the Z-axis are adjusted multiple times, wherein the preset step size can be one or more, and this embodiment does not impose too many restrictions on this.
[0063] During implementation, the current predicted space coordinates are adjusted once or multiple times according to one or more preset step sizes. The current predicted space coordinates can also be iteratively adjusted, that is, after adjusting the current predicted space coordinates to obtain the adjusted space coordinates, the adjusted space coordinates are continued to be adjusted multiple times. This embodiment does not impose too many restrictions on the specific method of adjusting the predicted space coordinates.
[0064] Optionally, in this embodiment, the image quality may be determined by the light leakage rate of the image, or the image quality may be determined based on other index parameters, which is not limited in this embodiment.
[0065] Process b) screening out adjustment space coordinates corresponding to image qualities that meet preset requirements, where one adjustment space coordinate corresponds to one image quality;
[0066] First, the predicted space coordinates are adjusted to obtain an adjusted space coordinate. Then, pixels are rearranged based on the adjusted space coordinates to obtain a corresponding pixel-rearranged image, and the image quality of the image is calculated. That is, each adjustment has the following corresponding relationship:
[0067] Adjustment times → adjustment space coordinates → pixel rearrangement image → image quality, so the corresponding adjustment space coordinates can be screened out according to the image quality that meets the preset requirements. When the image quality meets the preset requirements, the adjustment space coordinates corresponding to the image quality are screened.
[0068] In implementation, the image quality that meets the preset requirements includes but is not limited to the image quality index parameter being lower than the first threshold or greater than the second threshold, the image quality index parameter being the minimum or maximum, etc. This embodiment does not impose too many restrictions on the screening conditions for image quality.
[0069] In some embodiments, the image quality obtained by rearranging pixels of the image to be displayed is determined by:
[0070] The method involves capturing a pixel-rearranged image displayed on a display screen using a binocular camera calibrated at the eyes of a 3D model (head model), calculating a light leakage rate for the pixel-rearranged image, and determining the image quality of the pixel-rearranged image displayed on the display screen based on the light leakage rate. A lower light leakage rate indicates better image quality.
[0071] During implementation, the calibrated stereo model is moved to the zero position and then fixed. The predicted spatial coordinates of the calibrated stereo model at the zero position are used as the initial values for pixel rearrangement. Then, the coordinate components of the predicted spatial coordinates on the X and Z axes are adjusted. The binocular camera at the eyes of the calibrated stereo model is used to capture the display content of the display screen, and the light leakage rate corresponding to different predicted spatial coordinates of the calibrated stereo model is calculated. During the adjustment of the predicted spatial coordinates, the calibrated stereo model is fixed.
[0072] In this embodiment, a binocular view (i.e., an image displayed on a display screen) captured by a binocular camera is used instead of the human eye to capture an image. As shown in FIG4A-4D , this embodiment provides an image displayed on a display screen. Specifically, the quality of the image captured by the binocular camera is determined by the following process:
[0073] In process b11), a completely white image (such as FIG. 4A ) is first displayed using a soft interleaving algorithm to locate the display area. A threshold is then applied to binarize the image to obtain the display boundary I1, where pixels at the display location are set to 1 and pixels at other locations are set to 0. The thresholded binarized image (display) serves as a mask image. The soft interleaving algorithm is used to rearrange pixels.
[0074] In process b12), a soft interleaving algorithm is performed on a completely black image (such as FIG4B ) to obtain an image I2 displayed on the display screen, and the pixel mean Th in the newly obtained image I2 is calculated, that is, the brightness of the displayed completely black image is calculated.
[0075] In process b13, a soft interleaving algorithm is performed on an image with white on the left and black on the right (as shown in FIG4C ). At this time, the right view captured by the binocular camera should show a completely black display image I3. However, due to light leakage, a gray image is displayed. The proportion of pixels in I3 that are greater than the mean value Th in image I2 is calculated and used as the light leakage rate r3. The calculation formula for the light leakage rate is as follows:
[0076] In process b14, a soft interleaving algorithm is performed on the image with black on the left and white on the right (as shown in FIG4D ). At this time, the left view captured by the binocular camera should show a completely black display image I4, but due to light leakage, a gray image is displayed. The proportion of pixels in I4 that are greater than the mean value Th in image I2 is calculated and used as the light leakage rate r4. The calculation formula for the light leakage rate is as follows:
[0077] Different light leakage rates are obtained by inputting different adjustment space coordinates into the soft interleaving algorithm. The sum of the light leakage rates corresponding to the right and left views captured by the binocular camera is calculated, that is, r3+r4 is calculated. When r3+r4 reaches the minimum value, the corresponding adjustment space coordinate is determined as the optimal position for calibrating the stereo model. The optimal position and the predicted space coordinate are used to determine the translation of the camera extrinsic parameters.
[0078] During implementation, due to the small installation error, the rotation angle deviation of the camera of the display screen is close to zero, and the offset of the camera in the Y direction is almost zero, and the deviation in the Y direction has little effect on the final display effect. Therefore, when calibrating the camera extrinsic parameters, only the coordinate components on the X and Z axes are adjusted. By adjusting the predicted spatial coordinates based on the image quality as the reference standard, the camera extrinsic parameters can be accurately calibrated without moving the fixed calibration stereo model.
[0079] Process c) determining the camera extrinsic parameters based on the screened adjustment space coordinates and the predicted space coordinates.
[0080] Optionally, the image quality meeting the preset requirements includes minimizing the light leakage rate of the image, that is, when the light leakage rate of the displayed image is minimized, the camera extrinsic parameters are determined according to the adjusted space coordinates and the predicted space coordinates.
[0081] During implementation, the predicted spatial coordinates are adjusted to obtain the adjusted spatial coordinates. When the pixels of the image to be displayed on the display screen are rearranged according to the adjusted spatial coordinates, and the calculated light leakage rate is minimized, the adjusted spatial coordinates are used as the optimal adjustment position, and the camera extrinsic parameters are determined based on the optimal adjusted spatial coordinates and the predicted spatial coordinates. Alternatively, light leakage rates less than a threshold value may be screened out, and the camera extrinsic parameters may be determined based on the adjusted spatial coordinates corresponding to the screened light leakage rates and the predicted spatial coordinates. For example, the average of the adjusted spatial coordinates corresponding to the screened light leakage rates may be calculated, and the camera extrinsic parameters may be determined based on this average and the predicted spatial coordinates.
[0082] In some embodiments, the camera extrinsic parameters are determined as follows:
[0083] The camera extrinsic parameters are determined according to a translation relationship between the adjusted space coordinates corresponding to the image quality that meets the preset requirements and the predicted space coordinates.
[0084] Optionally, the coordinate components of the adjusted space coordinates and the predicted space coordinates in the Y-axis direction are the same, so it is only necessary to calculate the difference between the coordinate components of the adjusted space coordinates and the predicted space coordinates in the X-axis direction and the Z-axis direction to obtain the translation amount as the extrinsic parameter of the calibrated camera.
[0085] The translation relationship in this embodiment includes but is not limited to the translation amount.
[0086] In some embodiments, the translation relationship is determined as follows:
[0087] Determining difference components between the adjusted spatial coordinates and the predicted spatial coordinates in the X-axis and Z-axis directions;
[0088] The translation relationship is determined based on the difference components in the X-axis and Z-axis directions, and the offset component of the camera of the display screen relative to the center point of the display screen in the Y-axis direction.
[0089] During implementation, eye tracking technology is activated to obtain stable predicted spatial coordinates (X T , Y T , Z T ), where the predicted spatial coordinates are the eye positions of the calibrated stereo model in the camera coordinate system calculated by the SoC end of the display screen. T , Y T , Z T ) is the initial value for pixel rearrangement. First, fine-tune the image quality of the corresponding display in the X-axis direction to achieve the optimal effect. Then, fine-tune the image quality of the corresponding display in the Z-axis direction to achieve the optimal effect. Record the adjustment space coordinates (X G , Y G , Z G ), where Y G =Y T =D Y , thus determining the translation of the camera extrinsic parameter according to the adjusted space coordinate value and the predicted space coordinate, that is, since the rotation angle of the camera extrinsic parameter is assumed to be close to zero, the camera extrinsic parameter is (X G -X T , D Y , Z G -Z T ).
[0090] As shown in FIG5 , this embodiment further provides a method for calibrating camera extrinsic parameters, and the specific implementation process of the method is as follows:
[0091] Step 500: Control the calibration stereo model to move to the zero position;
[0092] During implementation, a red cross is displayed in the center area of the display screen, the calibration stereo model is controlled to move, and a hole is opened in the center of the eye of the calibration stereo model so that the laser behind the calibration stereo model passes through the center of the eye of the calibration stereo model and shines on the red cross. The distance (depth) from the calibration stereo model to the display screen is set to z, such as 650 mm.
[0093] Step 501: Obtain display parameters of the display screen;
[0094] Display parameters are used to represent the parameters used by the Offset lenticular lens grating to rearrange the 3D image. Optional display parameters include but are not limited to: Pitch (horizontal spacing of the lens), h (lens placement height), Θ (lens tilt angle), Offset (lens attachment offset). The default values of the camera external parameters that can be set for the display are: rotation angle (0, 0, 0), translation (0, D Y , 0), where D Y Indicates the distance from the camera center to the display center of the display.
[0095] Step 502: Capture a calibration stereo model image through the camera of the display screen, and determine the predicted spatial coordinates of the eyes of the calibration stereo model based on the calibration stereo model image;
[0096] Step 503: Using the predicted spatial coordinates as initial values, rearrange the pixels of the image to be displayed on the display screen, capture the rearranged image using a binocular camera calibrated to the eyes of the stereo model, and calculate the light leakage rate of the image.
[0097] Step 504: Adjust the predicted spatial coordinates multiple times to obtain adjusted spatial coordinates, rearrange the pixels of the image to be displayed using the adjusted spatial coordinates, capture the image of the display screen again using the binocular camera calibrated at the eyes of the stereo model, and calculate the light leakage rate of the image corresponding to each adjusted spatial coordinate.
[0098] Step 505: Filter out the adjustment space coordinate corresponding to the minimum light leakage rate from the light leakage rates of the images corresponding to the adjustment space coordinates;
[0099] Step 506: Determine the camera extrinsic parameters based on the difference components between the adjusted space coordinates corresponding to the minimum light leakage rate and the predicted space coordinates in the X-axis and Z-axis directions, and the offset component of the camera of the display screen relative to the center point of the display screen in the Y-axis direction.
[0100] The second calibration method controls the movement of the calibration stereo model to record the standard space coordinates of the calibration stereo model.
[0101] This embodiment can also obtain the standard space coordinates of the calibration stereo model in the world coordinate system by controlling the movement of the calibration stereo model, and determine the camera extrinsic parameters through the relationship between the standard space coordinates and the predicted space coordinates of the calibration stereo model.
[0102] In some embodiments, the camera extrinsic parameters are determined as follows:
[0103] a) Control the calibration stereo model to move to different positions except the zero point;
[0104] b) determining the predicted spatial coordinates and the standard spatial coordinates of the eye of the calibration stereo model corresponding to each position;
[0105] The predicted spatial coordinates are determined based on the position of the eye of the calibrated stereo model in a camera coordinate system, and the standard spatial coordinates are determined based on the position of the eye of the calibrated stereo model in a world coordinate system, wherein the world coordinate system is established according to the right-hand rule with the center point of the display screen as the origin;
[0106] During implementation, each time the calibrated stereo model moves, the predicted spatial coordinates of the calibrated stereo model's eyes in the camera coordinate system and the standard spatial coordinates of the calibrated stereo model's eyes in the world coordinate system are calculated using the image of the calibrated stereo model captured by the display screen. Coordinate pairs for the calibrated stereo model at each position in the camera coordinate system and the world coordinate system are obtained. The calibrated stereo model can move in any one or more directions along the X, Y, and Z axes at a time. This is not specifically limited in this embodiment. The step size of the movement of the calibrated stereo model can be fixed, random, or multiple step sizes can be set, with movement along the X, Y, and Z axes at different step sizes. This embodiment does not specifically limit the step size. Optionally, the range of movement of the calibrated stereo model is based on the field of view of the display camera, and the movement of the calibrated stereo model is controlled within the camera's field of view. Alternatively, the positions of the calibrated stereo model along the X, Y, and Z axes can be controlled to be as evenly distributed as possible.
[0107] It should be noted that since it takes about 2 seconds to stabilize the predicted spatial coordinates, it is necessary to ensure that the calibrated stereo model is controlled to move next time after the predicted spatial coordinates corresponding to the moving position of the calibrated stereo model are calculated.
[0108] c) Determine the camera extrinsic parameters of the display screen's camera based on the predicted spatial coordinates and the standard spatial coordinates of the eye of the calibrated stereo model corresponding to each position.
[0109] In implementation, since the calibration stereo model corresponds to a predicted space coordinate and a standard space coordinate at the same position, the camera extrinsic parameters can be calibrated as long as the conversion relationship between the predicted space coordinate and the standard space coordinate is found.
[0110] In some embodiments, the camera extrinsic parameters of the camera of the display screen are determined according to the predicted spatial coordinates and the standard spatial coordinates of the eye of the calibrated stereo model corresponding to each position in the following manner:
[0111] According to the predicted space coordinates and standard space coordinates of the eye of the calibrated stereo model corresponding to each position, the conversion relationship between each predicted space coordinate and each standard space coordinate is determined; according to the conversion relationship between each predicted space coordinate and each standard space coordinate, the camera extrinsic parameter is determined.
[0112] Optionally, the transformation relationship includes a rotation relationship and a translation relationship, wherein the rotation relationship includes a rotation angle such as a rotation matrix, wherein the rotation matrix is determined based on the rotation angle, and the rotation matrix can be used as a parameter of the camera extrinsic parameter. The rotation angle includes but is not limited to yaw (yaw angle), pitch (pitch angle), and roll (roll angle). The translation relationship includes a translation amount, and the translation amount can be used as a parameter of the camera extrinsic parameter, wherein the translation amount includes but is not limited to the translation amount in any direction of the X axis, Y axis, and Z axis.
[0113] In some embodiments, the conversion relationship between each predicted space coordinate and each standard space coordinate is determined by any of the following methods:
[0114] Method 1) Determine the rotation relationship between each prediction space coordinate and each standard space coordinate by singular value decomposition;
[0115] In practice, a general matrix solution can be used. As shown in Figure 6, this embodiment provides a schematic diagram of the transformation relationship between datasets. In this example, we want to find the optimal rotation and translation to align points in dataset A with those in dataset B. This transformation is called a Euclidean transformation or a rigid transformation because it preserves shape and size. The formula is: B = R × A + t (Equation (3)).
[0116] Where R and t are the transformations applied to dataset A to align it as closely as possible with dataset B. R represents the rotation matrix, and t represents the translation. In this embodiment, the predicted spatial coordinates are used as dataset A, and the standard spatial coordinates are used as dataset B to solve for the rotation and translation relationships.
[0117] Specifically, solving the optimal rigid transformation matrix can be decomposed into the following steps:
[0118] (a) Solve the centroid of two data sets;
[0119] The centroid is the average point, which can be calculated by the following formula:
[0120] Where P represents the spatial coordinate, represents the i-th prediction space coordinate in dataset A, Represents the i-th standard space coordinate in dataset B, centroid ARepresents the centroid of data set A, centroid B Denotes the centroid of dataset B; N denotes the number of spatial coordinates, where the number of predicted spatial coordinates in dataset A is the same as the number of standard spatial coordinates in dataset B. Dataset A includes N predicted spatial coordinates in the camera coordinate system, and dataset B includes N standard spatial coordinates in the world coordinate system, where one predicted spatial coordinate corresponds to one standard spatial coordinate.
[0121] (b) Bring the two data sets to the origin and then find the optimal rotation matrix R (i.e., the rotation relationship);
[0122] There are several ways to find the optimal rotation between points. The simplest method is to use singular value decomposition (SVD). To find the optimal rotation, we first re-center the two data sets so that both center points are at the origin. Figure 7 shows a schematic diagram of a method for determining the rotation relationship. Using the singular value decomposition method, the formula for solving the rotation matrix is as follows:
[0123] Where H represents the covariance matrix, represents the i-th prediction space coordinate in dataset A, Represents the i-th standard space coordinate in dataset B, centroid A Represents the centroid of data set A, centroid B represents the centroid of the dataset B, N represents the number of spatial coordinates, SVD represents the singular value, R represents the rotation matrix, and T represents the matrix transpose.
[0124] (c) Using the center of mass data, calculate the translation amount using the following formula: t = BR × A Formula (6);
[0125] Among them, t represents the translation amount (i.e., the translation relationship), R represents the rotation matrix, B can be taken as the center of mass of data set B, and A can also be taken as the center of mass of data set A; or, B can be taken as each standard space coordinate in data set B, and A can also be taken as each predicted space coordinate in data set A. Finally, the multiple t obtained by calculating multiple sets of standard space coordinates and predicted space coordinates are averaged to obtain the final translation amount.
[0126] Method 2) By constructing a loss function, the conversion relationship between each predicted space coordinate and each standard space coordinate is determined when each predicted space coordinate is closest to each standard space coordinate.
[0127] In some embodiments, the loss function is constructed and the conversion relationship is determined by:
[0128] 2a) constructing a loss function based on the predicted spatial coordinates, the standard spatial coordinates, and the rotation angle variable, and updating the rotation angle variable based on the loss function value;
[0129] Optionally, construct a loss function by following the steps below:
[0130] Determining a first centroid of each predicted space coordinate and a second centroid of each standard space coordinate;
[0131] constructing a standard space variable corresponding to each prediction space coordinate according to the first center of mass, the second center of mass, the rotation angle variable, and each prediction space coordinate;
[0132] Construct a loss function based on each standard space variable and each standard space coordinate.
[0133] 2b) Determine the rotation relationship between each predicted space coordinate and each standard space coordinate based on the rotation angle variable corresponding to the minimum loss function value.
[0134] Since calculating the optimal rotation matrix by singular value decomposition has limitations and requires calculating the inverse matrix, which makes the result unreliable, this embodiment also provides a method for solving the optimal rotation matrix by constructing a loss function.
[0135] For example, N sets of collected data are known, each set of collected data includes a predicted space coordinate and a standard space coordinate, where the predicted space coordinate is expressed as The standard space coordinates are expressed as
[0136] The unknown camera external parameters include the rotation angle (radian) variable θ = (θ x ,θ y ,θ z ) and the translation amount, i.e., the position offset (mm) d = (d x ,d y ,d z );
[0137] The loss function in this embodiment can be defined as follows:
[0138] in:
[0139] In formula (7), Respectively represent the preset weight ratios of the X-axis, Y-axis, and Z-axis, θ x ,θ y ,θ z represents the rotation angle variable, Represents the rotation matrix variable, mean(*) represents the average value, represents the prediction space coordinates, Represents standard space coordinates.
[0140] According to the above formula, we can get (x n ,y n ,z n )and The closer it is, the lower the loss function value. Calculate the gradient of L(θ) Adjust the gradient according to the loss function value for backpropagation.
[0141] Update θ according to the gradient value: Where t represents the number of iterations, l r Represents the set learning rate. By gradually updating θ, we find the optimal rotation angle, thereby obtaining the optimal rotation matrix, and then the optimal translation is obtained through the following steps.
[0142] In some embodiments, the transformation relationship between each predicted space coordinate and each standard space coordinate includes a rotation relationship and a translation relationship; the translation relationship is determined by:
[0143] The translation relationship is determined according to the predicted space coordinates, the standard space coordinates and the rotation relationship.
[0144] In practice, the translation amount can be solved by the following formula: t = BR × A Formula (8);
[0145] Among them, t represents the translation amount (i.e., the translation relationship), R represents the (optimal) rotation matrix, B can be taken as the center of mass of data set B, and A can also be taken as the center of mass of data set A; or, B can be taken as each standard space coordinate in data set B, and A can also be taken as each predicted space coordinate in data set A. Finally, the multiple t obtained by calculating multiple sets of standard space coordinates and predicted space coordinates are averaged to obtain the final translation amount.
[0146] As shown in Figure 8, this embodiment also provides a schematic diagram of the structure of an external parameter calibration system. Taking the calibration stereo model as a head model as an example, it includes a display screen, a calibration fixture, and a head model for simulating a human. Optionally, the display screen includes a 3D naked-eye display screen, and the display screen also includes an FPGA (Field Programmable Gate Array), a SoC (System on Chip), etc. The calibration fixture includes a built-in host computer control interface (using PAD as the control interface of the calibration fixture, including but not limited to a camera calibration algorithm operation interface and a camera real-time image viewing interface), an SDK, and a calibration PC (display). The calibration PC establishes a communication connection with the SoC of the display screen via HDMI and USB, and the display screen includes a camera.
[0147] During implementation, a calibration jig is used to control the movement of the head model, with the jig's motion position output as the standard space coordinates. The output of the SoC eye tracking solution serves as the predicted space coordinates to be calibrated. The extrinsic parameter calibration algorithm in this embodiment calculates the camera extrinsic parameters and converts them into a dataset A in a known camera coordinate system and a dataset B in a world coordinate system. The dataset A includes one or more predicted space coordinates, and the dataset B includes one or more standard space coordinates. The optimal rotation matrix R and translation t are calculated so that the value of R×A+t is close to B.
[0148] For fully automated camera extrinsic calibration, the head model is placed in front of the display screen and a calibration fixture equipped with sensors controls the head model's movement to simulate human motion. The calibration PC controls the head model's movement along the X, Y, and Z axes. Because a standard fixture is used (with the center of the screen as the origin and a right-hand rule governing the coordinate system (world coordinate system)), the SoC obtains data from the display's embedded camera to calculate the head model's eye position in the camera coordinate system and outputs it to the calibration PC.
[0149] The above-mentioned calibration system structure ensures fully automated labeling of camera extrinsic parameters. The calibration PC obtains the standard spatial coordinates by controlling the movement of the head model through the SoC. The predicted spatial coordinates obtained by the SoC can be transmitted to the calibration PC via Socket communication. The calibration PC outputs the camera extrinsic parameters to the SoC.
[0150] As shown in FIG9 , this embodiment also provides an implementation process of camera extrinsic calibration, which is as follows:
[0151] Step 900: Start calibration. Set the center of the display screen as the origin through the calibration fixture SDK and control the calibration 3D model to move to the zero position.
[0152] During implementation, the calibration model is first roughly moved to the center of the display screen using the calibration jig's built-in control interface. To ensure that the calibration jig's mechanical components, including the calibration jig, use the center of the display screen as their origin, a hole is opened between the two eyes of the calibration model. A large laser is positioned at the rear of the model to ensure that the laser passes through the model, forming a red spot on the display screen. The calibration model is then moved so that the red laser appears at the center of the display screen, which is the zero point. Ensure that the laser, the center of the eye, and the center of the screen are collinear.
[0153] Step 901: Determine whether the number of times the calibration 3D model has been moved has reached M times. If not, proceed to step 902; if so, proceed to step 904.
[0154] Step 902: Control the movement of the calibration stereo model through the calibration fixture SDK, record the standard spatial coordinates of the calibration stereo model in the world coordinate system, and receive the predicted spatial coordinates in the camera coordinate system from the SoC end;
[0155] During implementation, the 3D model is initially positioned at z0_gt = 650, x0_gt = 0, and y0_gt = 0. The standard spatial coordinates (x0_gt, y0_gt, z0_gt) of the current 3D model relative to the center of the display are recorded. The calibration PC then calls the SDK setup interface provided by the manufacturer to position the 3D model at the specified location. Once the 3D model reaches the specified location, the SoC is notified to run the algorithm program to obtain the predicted spatial coordinates (x0_pred, y0_pred, z0_pred) of the 3D model. The predicted spatial coordinates are then recorded and sent to the calibration PC. Because the predicted spatial coordinates obtained by the SoC require 2 seconds to stabilize, the SoC must obtain the predicted spatial coordinates before notifying the calibration PC that the 3D model can continue its next movement.
[0156] The calibration fixture SDK controls the movement of the calibration stereo model along the X, Y, and Z axes in the world coordinate system to obtain its standard spatial coordinates (xi_gt, i_gt, zi_gt), where i = 1, 2, 3, or M. Simultaneously, the SoC runs an algorithm to obtain the predicted spatial coordinates (xi_pred, yi_pred, zi_pred) of the calibration stereo model. This yields M pairs of standard spatial coordinates and predicted coordinates. The calibration stereo model remains within the display camera's field of view during movement.
[0157] Step 903: Call the calibration program based on the two sets of data, the standard space coordinates and the predicted space coordinates, calculate the rotation matrix and translation parameters, and convert them into rotation angles and translation amounts;
[0158] Step 904: Transmit the rotation angle and translation as camera external parameters to the SoC end via USB.
[0159] During implementation, the calibration PC starts the calibration program to obtain the rotation matrix and translation parameters, and converts them into three rotation angles (yaw, pitch, roll) and three translation amounts as camera extrinsics, which are then transferred to the SoC via USB.
[0160] Optionally, the calibration stereo model in the above calibration system can be fixed after moving to the zero position, and a binocular camera at the eye position of the calibration stereo model is used to capture the image of the display screen, and the position of the calibration stereo model is adjusted according to the image quality to calculate the camera extrinsic parameters.
[0161] Based on the same inventive concept, the embodiment of the present disclosure also provides a camera extrinsic parameter calibration system. Since the principle of solving the problem by the system is similar to that of the method, the implementation of the system can refer to the implementation of the method, and the repeated parts will not be repeated.
[0162] As shown in FIG10 , the camera extrinsic calibration system includes a display screen 1000 , a calibration fixture 1001 , and a calibration stereo model 1002 . The display screen 1000 and the calibration fixture 1001 establish a communication connection, wherein:
[0163] The calibration fixture 1001 is configured to control the calibration stereo model to move to a zero position, wherein the calibration stereo model 1002 at the zero position faces the display screen, and a line connecting the first center point of the calibration stereo model and the center point of the display screen is perpendicular to the plane of the display screen;
[0164] The display screen 1000 is configured to use a camera to capture a calibration stereo model image, determine the predicted spatial coordinates of the calibration stereo model eyes based on the calibration stereo model image; and determine the camera extrinsic parameters of the camera of the display screen 1000 based on the predicted spatial coordinates of the calibration stereo model eyes.
[0165] As an optional implementation, the calibration fixture is specifically configured to perform:
[0166] Adjusting the predicted spatial coordinates multiple times, rearranging pixels of an image to be displayed on a display screen using the adjusted spatial coordinates, and determining image quality after the pixel rearrangement;
[0167] Filtering out adjustment space coordinates corresponding to image qualities that meet preset requirements, where one adjustment space coordinate corresponds to one image quality;
[0168] The camera extrinsic parameters are determined according to the filtered adjustment space coordinates and the predicted space coordinates.
[0169] As an optional implementation, the calibration fixture is specifically configured to perform:
[0170] The coordinate components of the predicted spatial coordinates on the X-axis and the Z-axis are adjusted multiple times according to a preset step size.
[0171] As an optional implementation, the calibration fixture is specifically configured to perform:
[0172] The camera extrinsic parameters are determined according to the translation relationship between the filtered adjustment space coordinates and the predicted space coordinates.
[0173] As an optional implementation, the calibration fixture is specifically configured to perform:
[0174] Determining difference components between the adjusted spatial coordinates and the predicted spatial coordinates in the X-axis and Z-axis directions;
[0175] The translation relationship is determined based on the difference components in the X-axis and Z-axis directions and the offset component of the camera of the display screen relative to the center point of the display screen in the Y-axis direction.
[0176] As an optional embodiment, the system further includes a binocular camera disposed at the eyes of the calibration stereo model; the calibration fixture is specifically configured to perform:
[0177] Controlling the binocular camera to capture the image after pixel rearrangement displayed on the display screen, and calculating the light leakage rate of the image after pixel rearrangement;
[0178] The image quality of the image after pixel rearrangement displayed on the display screen is determined according to the light leakage rate.
[0179] As an optional implementation, the calibration fixture is further configured to perform:
[0180] Controlling the calibration stereo model to move to different positions other than the zero position, and determining the standard spatial coordinates of the eyes of the calibration stereo model corresponding to each position, wherein the standard spatial coordinates are determined based on the position of the eyes of the calibration stereo model in a world coordinate system, wherein the world coordinate system is created according to the right-hand rule with the center point of the display screen as the origin;
[0181] Receiving predicted spatial coordinates of the eyes of the calibration stereo model corresponding to each position sent by the display screen, wherein the predicted spatial coordinates are determined according to the positions of the eyes of the calibration stereo model in the camera coordinate system;
[0182] According to the predicted spatial coordinates and the standard spatial coordinates of the eye of the calibration stereo model corresponding to each position, the camera extrinsic parameters of the camera of the display screen are determined, and the camera extrinsic parameters are sent to the display screen.
[0183] As an optional implementation, the calibration fixture is further configured to perform:
[0184] Determine the conversion relationship between each predicted space coordinate and each standard space coordinate according to the predicted space coordinate and the standard space coordinate of the eye of the calibration stereo model corresponding to each position;
[0185] The camera extrinsic parameters are determined according to the conversion relationship between each predicted space coordinate and each standard space coordinate.
[0186] As an optional implementation, the calibration fixture is further configured to perform:
[0187] Determine the rotation relationship between each prediction space coordinate and each standard space coordinate by singular value decomposition; or
[0188] By constructing a loss function, the conversion relationship between each predicted space coordinate and each standard space coordinate is determined when each predicted space coordinate is closest to each standard space coordinate.
[0189] As an optional implementation, the calibration fixture is further configured to perform:
[0190] Construct a loss function based on the predicted space coordinates, the standard space coordinates and the rotation angle variable, and update the rotation angle variable based on the loss function value;
[0191] According to the rotation angle variable corresponding to the minimum loss function value, the rotation relationship between each predicted space coordinate and each standard space coordinate is determined.
[0192] As an optional implementation, the calibration fixture is further configured to perform:
[0193] Determining a first centroid of each predicted space coordinate and a second centroid of each standard space coordinate;
[0194] constructing a standard space variable corresponding to each prediction space coordinate according to the first center of mass, the second center of mass, the rotation angle variable, and each prediction space coordinate;
[0195] Construct a loss function based on each standard space variable and each standard space coordinate.
[0196] As an optional implementation, the transformation relationship between each predicted space coordinate and each standard space coordinate includes a rotation relationship and a translation relationship; the calibration fixture is further configured to determine the translation relationship in the following manner:
[0197] The translation relationship is determined according to the predicted space coordinates, the standard space coordinates and the rotation relationship.
[0198] Based on the same inventive concept, the embodiment of the present disclosure also provides an electronic device. Since the principle of solving the problem by the electronic device is similar to that of the method, the implementation of the electronic device can refer to the implementation of the method, and the repeated parts will not be repeated.
[0199] As shown in FIG11 , the electronic device includes a processor 1100 and a memory 1101 . The memory 1101 is used to store programs executable by the processor 1100 . The processor 1100 is used to read the programs in the memory 1101 and perform the following steps:
[0200] Controlling the calibration three-dimensional model to move to a zero position, wherein the calibration three-dimensional model at the zero position faces the display screen, and a line connecting the first center point of the calibration three-dimensional model and the center point of the display screen is perpendicular to the plane of the display screen;
[0201] Capturing a calibration stereo model image through a camera of the display screen, and determining predicted spatial coordinates of the eye of the calibration stereo model according to the calibration stereo model image;
[0202] The camera extrinsic parameters of the display screen's camera are determined based on the predicted spatial coordinates of the eye of the calibrated stereo model.
[0203] As an optional implementation manner, the processor 1100 is specifically configured to execute:
[0204] Adjusting the predicted spatial coordinates multiple times, rearranging pixels of an image to be displayed on a display screen using the adjusted spatial coordinates, and determining image quality after the pixel rearrangement;
[0205] Filtering out adjustment space coordinates corresponding to image qualities that meet preset requirements, where one adjustment space coordinate corresponds to one image quality;
[0206] The camera extrinsic parameters are determined according to the filtered adjustment space coordinates and the predicted space coordinates.
[0207] As an optional implementation manner, the processor 1100 is specifically configured to execute:
[0208] The coordinate components of the predicted spatial coordinates on the X-axis and the Z-axis are adjusted multiple times according to a preset step size.
[0209] As an optional implementation manner, the processor 1100 is specifically configured to execute:
[0210] The camera extrinsic parameters are determined according to the translation relationship between the filtered adjustment space coordinates and the predicted space coordinates.
[0211] As an optional implementation manner, the processor 1100 is specifically configured to execute:
[0212] Determining difference components between the adjusted spatial coordinates and the predicted spatial coordinates in the X-axis and Z-axis directions;
[0213] The translation relationship is determined based on the difference components in the X-axis and Z-axis directions and the offset component of the camera of the display screen relative to the center point of the display screen in the Y-axis direction.
[0214] As an optional implementation manner, the processor 1100 is specifically configured to determine the image quality obtained by pixel rearrangement of the image to be displayed in the following manner:
[0215] The binocular camera at the eye of the calibrated stereo model is used to capture the pixel-rearranged image displayed on the display screen, and the light leakage rate of the pixel-rearranged image is calculated;
[0216] The image quality of the image after pixel rearrangement displayed on the display screen is determined according to the light leakage rate.
[0217] As an optional implementation manner, the processor 1100 is specifically configured to execute:
[0218] Control the calibration stereo model to move to different positions except the zero point;
[0219] Determining predicted spatial coordinates and standard spatial coordinates of the eye of the calibrated stereo model corresponding to each position, wherein the predicted spatial coordinates are determined based on the position of the eye of the calibrated stereo model in a camera coordinate system, and the standard spatial coordinates are determined based on the position of the eye of the calibrated stereo model in a world coordinate system, where the world coordinate system is created according to the right-hand rule with the center point of the display screen as the origin;
[0220] The camera extrinsic parameters of the camera of the display screen are determined according to the predicted spatial coordinates and the standard spatial coordinates of the eye of the calibration stereo model corresponding to each position.
[0221] As an optional implementation manner, the processor 1100 is specifically configured to execute:
[0222] Determine the conversion relationship between each predicted space coordinate and each standard space coordinate according to the predicted space coordinate and the standard space coordinate of the eye of the calibration stereo model corresponding to each position;
[0223] The camera extrinsic parameters are determined according to the conversion relationship between each predicted space coordinate and each standard space coordinate.
[0224] As an optional implementation manner, the processor 1100 is specifically configured to execute:
[0225] Determine the rotation relationship between each prediction space coordinate and each standard space coordinate by singular value decomposition; or
[0226] By constructing a loss function, the conversion relationship between each predicted space coordinate and each standard space coordinate is determined when each predicted space coordinate is closest to each standard space coordinate.
[0227] As an optional implementation manner, the processor 1100 is specifically configured to execute:
[0228] Construct a loss function based on the predicted space coordinates, the standard space coordinates and the rotation angle variable, and update the rotation angle variable based on the loss function value;
[0229] According to the rotation angle variable corresponding to the minimum loss function value, the rotation relationship between each predicted space coordinate and each standard space coordinate is determined.
[0230] As an optional implementation manner, the processor 1100 is specifically configured to execute:
[0231] Determining a first centroid of each predicted space coordinate and a second centroid of each standard space coordinate;
[0232] constructing a standard space variable corresponding to each prediction space coordinate according to the first center of mass, the second center of mass, the rotation angle variable, and each prediction space coordinate;
[0233] Construct a loss function based on each standard space variable and each standard space coordinate.
[0234] As an optional implementation manner, the transformation relationship between each predicted space coordinate and each standard space coordinate includes a rotation relationship and a translation relationship; the processor 1100 is specifically configured to determine the translation relationship in the following manner:
[0235] The translation relationship is determined according to the predicted space coordinates, the standard space coordinates and the rotation relationship.
[0236] Based on the same inventive concept, the embodiment of the present disclosure also provides a camera extrinsic parameter calibration device. Since the principle of solving the problem by the device is similar to that of the method, the implementation of the device can refer to the implementation of the method, and the repeated parts will not be repeated.
[0237] As shown in FIG12 , the device includes:
[0238] The zero setting module 1200 is used to control the calibration 3D model to move to a zero position, wherein the calibration 3D model at the zero position faces the display screen, and the line connecting the first center point of the calibration 3D model and the center point of the display screen is perpendicular to the plane of the display screen;
[0239] Prediction module 1201, configured to capture a calibration stereo model image through a camera of a display screen, and determine predicted spatial coordinates of the eyes of the calibration stereo model based on the calibration stereo model image;
[0240] The calibration module 1202 is configured to determine the camera extrinsic parameters of the camera of the display screen according to the predicted spatial coordinates of the eyes of the calibrated stereo model.
[0241] As an optional implementation, the calibration module 1202 is specifically configured to:
[0242] Adjusting the predicted spatial coordinates multiple times, rearranging pixels of an image to be displayed on a display screen using the adjusted spatial coordinates, and determining image quality after the pixel rearrangement;
[0243] Filtering out adjustment space coordinates corresponding to image qualities that meet preset requirements, where one adjustment space coordinate corresponds to one image quality;
[0244] The camera extrinsic parameters are determined according to the filtered adjustment space coordinates and the predicted space coordinates.
[0245] As an optional implementation, the calibration module 1202 is specifically configured to:
[0246] The coordinate components of the predicted spatial coordinates on the X-axis and the Z-axis are adjusted multiple times according to a preset step size.
[0247] As an optional implementation, the calibration module 1202 is specifically configured to:
[0248] The camera extrinsic parameters are determined according to the translation relationship between the filtered adjustment space coordinates and the predicted space coordinates.
[0249] As an optional implementation, the calibration module 1202 is specifically configured to:
[0250] Determining difference components between the adjusted spatial coordinates and the predicted spatial coordinates in the X-axis and Z-axis directions;
[0251] The translation relationship is determined based on the difference components in the X-axis and Z-axis directions and the offset component of the camera of the display screen relative to the center point of the display screen in the Y-axis direction.
[0252] As an optional implementation, the calibration module 1202 is specifically configured to determine the image quality obtained by rearranging pixels of the image to be displayed in the following manner:
[0253] The binocular camera at the eye of the calibrated stereo model is used to capture the pixel-rearranged image displayed on the display screen, and the light leakage rate of the pixel-rearranged image is calculated;
[0254] The image quality of the image after pixel rearrangement displayed on the display screen is determined according to the light leakage rate.
[0255] As an optional implementation, the calibration module 1202 is specifically configured to:
[0256] Control the calibration stereo model to move to different positions except the zero point;
[0257] Determining predicted spatial coordinates and standard spatial coordinates of the eye of the calibrated stereo model corresponding to each position, wherein the predicted spatial coordinates are determined based on the position of the eye of the calibrated stereo model in a camera coordinate system, and the standard spatial coordinates are determined based on the position of the eye of the calibrated stereo model in a world coordinate system, where the world coordinate system is created according to the right-hand rule with the center point of the display screen as the origin;
[0258] The camera extrinsic parameters of the camera of the display screen are determined according to the predicted spatial coordinates and the standard spatial coordinates of the eye of the calibration stereo model corresponding to each position.
[0259] As an optional implementation, the calibration module 1202 is specifically configured to:
[0260] Determine the conversion relationship between each predicted space coordinate and each standard space coordinate according to the predicted space coordinate and the standard space coordinate of the eye of the calibration stereo model corresponding to each position;
[0261] The camera extrinsic parameters are determined according to the conversion relationship between each predicted space coordinate and each standard space coordinate.
[0262] As an optional implementation, the calibration module 1202 is specifically configured to:
[0263] Determine the rotation relationship between each prediction space coordinate and each standard space coordinate by singular value decomposition; or
[0264] By constructing a loss function, the conversion relationship between each predicted space coordinate and each standard space coordinate is determined when each predicted space coordinate is closest to each standard space coordinate.
[0265] As an optional implementation, the calibration module 1202 is specifically configured to:
[0266] Construct a loss function based on the predicted space coordinates, the standard space coordinates and the rotation angle variable, and update the rotation angle variable based on the loss function value;
[0267] According to the rotation angle variable corresponding to the minimum loss function value, the rotation relationship between each predicted space coordinate and each standard space coordinate is determined.
[0268] As an optional implementation, the calibration module 1202 is specifically configured to:
[0269] Determining a first centroid of each predicted space coordinate and a second centroid of each standard space coordinate;
[0270] constructing a standard space variable corresponding to each prediction space coordinate according to the first center of mass, the second center of mass, the rotation angle variable, and each prediction space coordinate;
[0271] Construct a loss function based on each standard space variable and each standard space coordinate.
[0272] As an optional implementation, the transformation relationship between each predicted space coordinate and each standard space coordinate includes a rotation relationship and a translation relationship; the calibration module 1202 is specifically configured to determine the translation relationship in the following manner:
[0273] The translation relationship is determined according to the predicted space coordinates, the standard space coordinates and the rotation relationship.
[0274] Based on the same inventive concept, embodiments of the present disclosure provide a computer storage medium comprising computer program code. When executed on a computer, the computer executes any of the camera extrinsic calibration methods discussed above. Because the principles underlying the aforementioned computer storage medium are similar to those of the camera extrinsic calibration method, the implementation of the aforementioned computer storage medium can be referenced to the implementation of the method, and any repetitions will be omitted.
[0275] In a specific implementation process, computer storage media may include: Universal Serial Bus Flash Drive (USB), mobile hard disk, Read-Only Memory (ROM), Random Access Memory (RAM), magnetic disk or optical disk, and other storage media that can store program code.
[0276] Based on the same inventive concept, embodiments of the present disclosure further provide a computer program product comprising computer program code that, when executed on a computer, causes the computer to execute any of the camera extrinsic calibration methods discussed above. Because the principles underlying the problems solved by the computer program products are similar to those of the camera extrinsic calibration methods, the implementation of the computer program products can be referenced to the implementation of the methods, and any repetitions will not be repeated.
[0277] The computer program product can employ any combination of one or more readable media. The readable medium can be a readable signal medium or a readable storage medium. The readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or component, or any combination thereof. More specific examples (non-exhaustive list) of readable storage media include: an electrical connection with one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof.
[0278] Those skilled in the art will appreciate that the embodiments of the present disclosure may be provided as methods, systems, or computer program products. Therefore, the present disclosure may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Furthermore, the present disclosure may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage and optical storage, etc.) containing computer-usable program code.
[0279] The present disclosure is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present disclosure. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device produce a device for implementing the functions specified in one or more processes in the flowchart and / or one or more boxes in the block diagram.
[0280] These computer program instructions may also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce a manufactured product including an instruction device that implements the functions specified in one or more processes in the flowchart and / or one or more boxes in the block diagram.
[0281] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, so that the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one or more processes in the flowchart and / or one or more boxes in the block diagram.
[0282] Obviously, those skilled in the art may make various changes and modifications to the present disclosure without departing from the spirit and scope of the present disclosure. Thus, if these modifications and variations of the present disclosure fall within the scope of the claims of the present disclosure and their equivalents, the present disclosure is intended to include these modifications and variations.
Claims
1. A method for calibrating the external parameters of a camera, wherein, The method includes: Controlling the calibrated stereo model to move to the zero position, where the calibrated stereo model at the zero position faces the display screen, and the line connecting the first center point of the calibrated stereo model and the center point of the display screen is perpendicular to the plane where the display screen is located; Collecting an image of the calibrated stereo model through the camera of the display screen, and determining the predicted spatial coordinates of the eyes of the calibrated stereo model according to the image of the calibrated stereo model; Determining the external camera parameters of the camera of the display screen according to the predicted spatial coordinates of the eyes of the calibrated stereo model.
2. The method according to claim 1, wherein, The determining the external camera parameters of the camera of the display screen according to the predicted spatial coordinates of the eyes of the calibrated stereo model includes: Adjusting the predicted spatial coordinates multiple times, using the adjusted adjusted spatial coordinates to perform pixel rearrangement on the image to be displayed on the display screen, and determining the image quality after pixel rearrangement; Selecting the adjusted spatial coordinates corresponding to the image quality that meets the preset requirements, where one adjusted spatial coordinate corresponds to one image quality; Determining the external camera parameters according to the selected adjusted spatial coordinates and the predicted spatial coordinates.
3. The method according to claim 2, wherein The adjusting the predicted spatial coordinates multiple times includes: Adjusting the coordinate components of the predicted spatial coordinates on the X-axis and Z-axis multiple times according to a preset step size.
4. The method according to claim 2, wherein, The determining the external camera parameters according to the selected adjusted spatial coordinates and the predicted spatial coordinates includes: Determining the external camera parameters according to the translation relationship between the selected adjusted spatial coordinates and the predicted spatial coordinates.
5. The method according to claim 4, wherein, The determining the translation relationship between the selected adjusted spatial coordinates and the predicted spatial coordinates includes: Determining the difference components of the selected adjusted spatial coordinates and the predicted spatial coordinates in the X-axis and Z-axis directions; Determining the translation relationship according to the difference components in the X-axis and Z-axis directions and the offset component of the camera of the display screen relative to the center point of the display screen in the Y-axis direction.
6. The method according to claim 2, wherein, The image quality obtained by pixel rearrangement of the image to be displayed is determined in the following manner: Taking a picture of the image after pixel rearrangement displayed on the display screen through the binocular cameras of the eyes of the calibrated stereo model, and calculating the light leakage rate of the image after pixel rearrangement; Determining the image quality of the image after pixel rearrangement displayed on the display screen according to the light leakage rate.
7. The method according to claim 1, wherein The determining the external camera parameters of the camera of the display screen according to the predicted spatial coordinates of the eyes of the calibrated stereo model includes: Controlling the calibrated stereo model to move to different positions other than the zero position; Determining the predicted spatial coordinates and standard spatial coordinates of the eyes of the calibrated stereo model corresponding to each position, where the predicted spatial coordinates are determined according to the position of the eyes of the calibrated stereo model in the camera coordinate system, and the standard spatial coordinates are determined according to the position of the eyes of the calibrated stereo model in the world coordinate system, and the world coordinate system is created with the center point of the display screen as the origin according to the right-hand rule; Determining the external camera parameters of the camera of the display screen according to the predicted spatial coordinates and standard spatial coordinates of the eyes of the calibrated stereo model corresponding to each position.
8. The method according to claim 7, wherein, The determining the external camera parameters of the camera of the display screen according to the predicted spatial coordinates and standard spatial coordinates of the eyes of the calibrated stereo model corresponding to each position includes: Determine the conversion relationship between each predicted spatial coordinate and each standard spatial coordinate according to the predicted spatial coordinates and standard spatial coordinates of the calibrated stereo model eye corresponding to each position; Determine the external camera parameters according to the conversion relationship between each predicted spatial coordinate and each standard spatial coordinate.
9. The method according to claim 8, wherein, The determination of the conversion relationship between each predicted spatial coordinate and each standard spatial coordinate includes: Determine the rotation relationship between each predicted spatial coordinate and each standard spatial coordinate by means of singular value decomposition; or, Determine the conversion relationship between each predicted spatial coordinate and each standard spatial coordinate when each predicted spatial coordinate and each standard spatial coordinate are closest by constructing a loss function.
10. The method according to claim 9, wherein, The determination of the conversion relationship between each predicted spatial coordinate and each standard spatial coordinate when each predicted spatial coordinate and each standard spatial coordinate are closest by constructing a loss function includes: Construct a loss function according to each predicted spatial coordinate, each standard spatial coordinate and the rotation angle variable, and update the rotation angle variable based on the loss function value; Determine the rotation relationship between each predicted spatial coordinate and each standard spatial coordinate according to the rotation angle variable corresponding to the minimum loss function value.
11. The method according to claim 10, wherein The construction of the loss function according to each predicted spatial coordinate, each standard spatial coordinate and the rotation angle variable includes: Determine the first centroid of each predicted spatial coordinate and the second centroid of each standard spatial coordinate; Construct a standard spatial variable corresponding to each predicted spatial coordinate according to the first centroid, the second centroid, the rotation angle variable and each predicted spatial coordinate; Construct a loss function according to each standard spatial variable and each standard spatial coordinate.
12. The method according to claim 9, wherein, The conversion relationship between each predicted spatial coordinate and each standard spatial coordinate includes a rotation relationship and a translation relationship; the translation relationship is determined by the following method: Determine the translation relationship according to each predicted spatial coordinate, each standard spatial coordinate and the rotation relationship.
13. An external camera calibration system, wherein, The system includes a display screen, a calibration fixture and a calibrated stereo model, and the display screen and the calibration fixture establish a communication connection, wherein: The calibration fixture is configured to control the calibrated stereo model to move to the zero position, where the calibrated stereo model at the zero position faces the display screen, and the line connecting the first center point of the calibrated stereo model and the center point of the display screen is perpendicular to the plane where the display screen is located; The display screen is configured to use a camera to collect an image of the calibrated stereo model, and determine the predicted spatial coordinates of the eye of the calibrated stereo model according to the image of the calibrated stereo model; according to the predicted spatial coordinates of the eye of the calibrated stereo model, determine the external camera parameters of the camera of the display screen.
14. The system according to claim 13, wherein, The calibration fixture is specifically configured to execute: Adjust the predicted spatial coordinates multiple times, perform pixel rearrangement on the image to be displayed on the display screen using the adjusted spatial coordinates, and determine the image quality after pixel rearrangement; Determine the external camera parameters according to the adjusted spatial coordinates corresponding to the image quality meeting the preset requirements and the predicted spatial coordinates.
15. The system according to claim 14, wherein The system further includes a binocular camera disposed at the eye of the calibrated stereo model; the calibration fixture is specifically configured to execute: Control the binocular camera to capture the image after pixel rearrangement displayed on the display screen, and calculate the light leakage rate of the image after pixel rearrangement; Determine the image quality of the image after pixel rearrangement displayed on the display screen according to the light leakage rate.
16. The system according to claim 13, wherein, The calibration fixture is specifically further configured to execute: Control the calibration stereo model to move to different positions other than the zero position, and determine the standard spatial coordinates of the eyes of the calibration stereo model corresponding to each position. The standard spatial coordinates are determined according to the position of the eyes of the calibration stereo model in the world coordinate system, and the world coordinate system is created with the center point of the display screen as the origin according to the right-hand rule; Receive the predicted spatial coordinates of the eyes of the calibration stereo model corresponding to each position sent by the display screen, where the predicted spatial coordinates are determined according to the position of the eyes of the calibration stereo model in the camera coordinate system; Determine the extrinsic parameters of the camera of the display screen according to the predicted spatial coordinates and the standard spatial coordinates of the eyes of the calibration stereo model corresponding to each position, and send the extrinsic parameters of the camera to the display screen.
17. An electronic device, wherein, The electronic device includes a processor and a memory. The memory is used to store programs executable by the processor, and the processor is used to read the programs in the memory and execute the steps of the method according to any one of claims 1 to 12.
18. A computer storage medium having a computer program stored thereon, wherein, When the program is executed by the processor, the steps of the method according to any one of claims 1 to 12 are implemented.
Citation Information
Patent Citations
Correction system and method for human eye tracking naked eye 3D display system
CN108063940A
Outer parameter correction jig of human eye tracking system and correction method
CN108108021A
Screen-camera calibration method and system, calibration equipment and storage medium
CN114820807A
Camera external parameter calibration method and device, vehicle and computer storage medium
CN115393450A
Camera external parameter calibration method and system and electronic equipment
CN117911531A