A three-dimensional positioning method based on a light field camera
By acquiring and processing object light field images using a light field camera, and combining parallax extraction and projection models, the resolution and universality issues of traditional 3D positioning methods in indoor and complex environments are solved, achieving high-precision panoramic 3D positioning.
Patent Information
- Application Number
- CN202210129190.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-02-11
- Publication Date
- 2025-11-04
- Estimated Expiration
- 2042-02-11
AI Technical Summary
Traditional 3D positioning methods have low positioning resolution and poor versatility in indoor and complex environments, and rely on complex optical components or precision mechanics, which cannot meet the needs of panoramic positioning.
A light field camera is used to acquire light field images of objects. By extracting parallax and calculating projection models, combined with a light field camera calibration program, a mapping relationship between spatial three-dimensional points and image sensor coordinates is established to achieve high-precision three-dimensional positioning.
It achieves high-precision, portable panoramic 3D positioning, solving the problems of low positioning resolution and poor universality of traditional methods, and achieving positioning accuracy superior to commercial Kinect 2.0.
Smart Images

Figure CN116630389B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to a three-dimensional positioning method based on a light field camera, and belongs to the field of visual positioning. BACKGROUND
[0002] Three-dimensional positioning technology plays an indispensable role in human-computer interaction, virtual reality (VR) and augmented reality (AR), national defense and military, and aerospace development. Traditional three-dimensional positioning methods such as GPS and wireless sensor networks require the target object as a signal source and the positioning resolution depends on the distribution density of the sensor, which cannot meet the panoramic positioning needs in indoor and complex environments.
[0003] Machine vision is a very key component in intelligent manufacturing. The positioning method based on digital image uses an optical imaging system to obtain the image and distance information of the target object, and obtains the appearance size, position coordinates, and attitude of the target object through related algorithms, so as to realize intelligent interaction and has important application value in industrial detection, automatic driving, and smart city. The main visual positioning methods include the structured light measurement method proposed by scholars Song Shaozhe, Zhu Lijun, and Foix S (Song Shaozhe, Niu Jinxing, Zhang Tao. Three-dimensional measurement technology based on structured light [J]. Henan Science and Technology, 2019(22):14-16; Zhu Lijun. Digital moire fringe three-dimensional surface measurement technology research [D]. Jinan, Shandong University, 2016; Foix S, GAlenya, C Torras. Lock-in time-of-flight (ToF) cameras: A survey [J]. IEEE Sensors Journal, 2011, 11(9):1917-1926.), the monocular camera visual positioning method proposed by Hartley R (Hartley R. Multiple view geometry in computer [M]. Cambridge university press. 2011.), and the binocular visual positioning method proposed by Scharstein D (Scharstein D, R Szeliski. A Taxonomy and Evaluation of Dense Two-Frame Stereo Correspondence Algorithms [J]. International Journal of Computer Vision. 2002. 47(1-3):7-42.).
[0004] The structured light measurement method uses a laser or a projection device to generate structured light with certain coding structure characteristics, and projects the structured light onto the object to be measured. The object surface causes the stripe to deform due to different depths. The three-dimensional structure and attitude information of the object are obtained by capturing and processing the deformed stripe by a camera. The monocular vision positioning method shoots the object by a camera, then extracts the feature points on the object, and calculates the position according to the distribution of the coordinates of the feature points in different coordinate systems and the camera model. The binocular vision needs to first calibrate the two cameras, calculate the corresponding internal parameters and the positional relationship between the two cameras, then perform stereo matching according to the corresponding feature point pairs on the left and right camera images, and then calculate the disparity to obtain the object pose information by using multiple point pairs.
[0005] However, the above method needs to rely on complex optical components or precise mechanical displacement to complete the accurate measurement of the depth information of the target object; at the same time, the binocular positioning method is not ideal in the stereo matching of the object with the occlusion relationship due to the lack of multi-view information. The dense and regular sampling characteristics of the light field camera make it possible to develop a high-accuracy visual positioning method using light field data.
[0006] Chinese invention patent CN107038719A discloses a depth estimation method and system of angle domain pixels based on light field images. The depth estimation algorithm of the technology is complex, the processor burden is too large, and the precision of the acquired depth map is poor. Chinese invention patent CN107578437 discloses a depth extraction method based on refocusing stack and object-image relationship, but the scheme and the geometry structure of the light field camera are not universal, and the mapping relationship between the image point and the spatial object point is not found in the two methods. SUMMARY
[0007] In order to solve the problems of poor universality and low positioning resolution of the current visual positioning, the present application provides a three-dimensional positioning method based on a light field camera, which comprises the following steps:
[0008] Step 1: collecting the light field image of the measured object: using a light field camera to shoot the image of the measured object, decoding the obtained image to obtain the light field image and converting it into a standard light field description form;
[0009] Step 2: disparity extraction: converting the light field image in step 1 into a sub-aperture image array, taking the center view angle of the sub-aperture image array as a reference, and obtaining a disparity map set M D ;
[0010] Step 3: projection model calculation: using a light field camera to shoot a checkerboard image, and obtaining the mapping relationship from a spatial three-dimensional point to an image sensor coordinate through a light field camera calibration program to form a projection model;
[0011] Step 4: Three-dimensional coordinate calculation: the parallax map set M obtained in step 2 D The unique parallax map D of adjacent views is obtained by mean fitting, and the parallax-depth conversion relationship of the sub-aperture image is obtained by combining the projection model in step 3, so as to obtain the three-dimensional coordinates of the measured object.
[0012] Optionally, the step 1 comprises:
[0013] Step 11: The light field camera collects the measured object to obtain an initial light field image L o ;
[0014] Step 12: The initial light field image L o is corrected to a standard light field image L s ;
[0015] Step 13: The standard light field image L s is rearranged to convert the hexagonal light field image into a standard four-dimensional light field in a square arrangement.
[0016] Optionally, the step 2 comprises:
[0017] Step 21: Extract the pixel points at the same position of all sub-aperture images in the light field image, and then combine them into a sub-aperture image; perform the same operation on all position pixel points of the sub-aperture image to obtain a light field sub-aperture image array;
[0018] Step 22: Take the sub-aperture image corresponding to the center view point of the sub-aperture image array as a reference view, and calculate the parallax of other views relative to the reference view by using image processing methods including but not limited to stereo matching method, optical flow method and machine learning, to obtain a parallax map set M D ;
[0019] Step 23: Combine the parallax map set M D obtained in step 22, and calculate a parallax map D representing adjacent two views by fitting mean.
[0020] Optionally, the step 3 comprises: step 31: establishing a mapping relationship between the camera coordinates [x, y, z] T and the space four-dimensional coordinates [u, v, s, t] T The camera coordinate system is a space three-dimensional coordinate system with the center of the main lens of the camera as the origin, and the space four-dimensional coordinate system is a space four-dimensional representation [s, t, u, v] T of the position coordinates [s, t] and direction coordinates [u, v] of the object source point m(x, y, z) in a certain plane in the three-dimensional coordinate system. TThe conversion relationship between the three-dimensional coordinates m(x, y, z) corresponding thereto is shown in formula (1):
[0021]
[0022] wherein λ is the distance from the arbitrary point to the origin of the camera coordinate system;
[0023] Step 32: Obtain the spatial four-dimensional coordinate system [s, t, u, v] by the chessboard calibration method T to the sensor light field coordinates [i, j, k, l] T , wherein [s, t] represents the spatial coordinates, [u, v] represents the angle coordinates, [i, j] represents the microlens coordinates, and [k, l] represents the local pixel coordinates corresponding to each microlens on the sensor, as shown in formula (2):
[0024]
[0025] wherein h (·) is the intrinsic matrix parameter;
[0026] Step 33: The mapping equation of the three-dimensional spatial coordinates [x, y, z] T to the sensor light field coordinates [i, j, k, l] T can be obtained by combining formula (1) and formula (2), as shown in formula (3):
[0027]
[0028] Optionally, the step 4 comprises:
[0029] Step 41: The linear slope of the two formulas in formula (3) represents the parallax of adjacent viewing angles, and the mapping relationship between the parallax and the depth is:
[0030]
[0031] wherein z is the depth coordinate in the three-dimensional coordinates;
[0032] Step 42: Obtain a group of spatial-angle coordinates [i, k] T or [j, l] T from the parallax map D;
[0033] Step 43: Calculate the (x, y) coordinates in the three-dimensional coordinates according to formula (3), and the calculation formula is shown in formula (5):
[0034]
[0035] Optionally, the method for calculating the parallax of each viewing angle relative to the central viewing angle is a stereo matching method.
[0036] Optionally, the method for calculating the parallax of each view relative to the center view is an optical flow method or a deep learning method.
[0037] A second object of the present application is to provide a light field camera-based three-dimensional positioning system, which uses the light field camera-based three-dimensional positioning method described above to realize three-dimensional positioning of an object.
[0038] The system comprises an image acquisition module, an image processing module, and a coordinate positioning module.
[0039] The image acquisition module uses a light field camera to acquire a light field image of the measured object.
[0040] The image processing module converts the light field image acquired by the image acquisition module into a sub-aperture image array, and obtains a parallax map set with the center view of the sub-aperture image array as a reference.
[0041] The coordinate positioning module uses a light field camera to capture a checkerboard image, obtains a mapping relationship from a spatial four-dimensional coordinate to an image sensor coordinate through a light field camera calibration program, forms a projection model, and then calculates the spatial three-dimensional coordinates of the measured object according to the projection model and the fitted parallax map obtained by the image processing module.
[0042] A third object of the present application is to provide a camera positioning device, which uses the light field camera-based three-dimensional positioning system described above to realize positioning of a measured object.
[0043] The present application also provides the application of the light field camera-based three-dimensional positioning method and / or the light field camera-based three-dimensional positioning system described above in the field of light field positioning.
[0044] The present application has the following advantages:
[0045] 1. The present application proposes a three-dimensional panoramic positioning technology based on light field imaging, which uses a light field camera to sample data and solves the technical requirements of traditional visual positioning systems, such as complexity and the need for additional scanning, by using the dense and regular sampling characteristics of the light field camera, thereby achieving the effect of miniaturized and portable visual positioning.
[0046] 2. The present application uses a high-precision three-dimensional panoramic light field positioning model based on joint optimization of light field camera calibration and light field parallax estimation, which combines the special optical structure of the light field camera and existing parallax estimation methods, solves the problems of poor universality and low precision of light field visual positioning, and achieves an effect superior to the positioning precision of commercial Kinect 2.0. BRIEF DESCRIPTION OF DRAWINGS
[0047] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings needed to be used in the embodiments will be briefly introduced as follows. Obviously, the drawings in the following description only constitute some embodiments of the present application, and for those skilled in the art, other drawings can also be obtained from these drawings without any creative effort.
[0048] Figure 1 Flow chart of three-dimensional positioning of the present application.
[0049] Figure 2 Schematic diagram of light field camera model of the present application.
[0050] Figure 3 Schematic diagram of light field image structure of the present application.
[0051] Figure 4 Schematic diagram of sub-aperture image conversion in the embodiments of the present application.
[0052] Figure 5 Schematic diagram of disparity-depth conversion of the present application.
[0053] Figure 6 Geometric projection model of the present application.
[0054] Figure 7 Schematic diagram of calculating xy coordinates of disparity map of the present application. DETAILED DESCRIPTION
[0055] In order to make the purpose, technical solutions and advantages of the present application more clear, the embodiments of the present application will be further described in detail below with reference to the drawings.
[0056] Embodiment one:
[0057] The present embodiment provides a three-dimensional positioning method based on a light field camera, which comprises:
[0058] Step 1: Collecting light field image of the measured object: using a light field camera to take an image of the measured object, decoding the obtained image to obtain a light field image and converting it into a standard light field description form;
[0059] Step 2: Disparity extraction: converting the light field image in step 1 into a sub-aperture image array, taking the center view angle of the sub-aperture image array as a reference, and obtaining a disparity map set M of each view angle relative to the center view angle D ;
[0060] Step 3: Projection model calculation: using a light field camera to take a checkerboard image, and through a light field camera calibration program, obtaining a mapping relationship of spatial three-dimensional points to image sensor coordinates to form a projection model;
[0061] Step 4: Three-dimensional coordinate calculation: the parallax map set solved in step 2 D The unique parallax map D of adjacent view angles is obtained by mean fitting, and the parallax-depth conversion relationship of the sub-aperture image is obtained by combining the projection model of step 3, and the three-dimensional coordinates of the measured object are obtained.
[0062] Embodiment two:
[0063] The embodiment provides a three-dimensional positioning method based on a light field camera, adopts an Illum light field camera as an acquisition device, uses an electrically controlled translation stage and a target doll as a measured object, and collects a three-dimensional scene of the measured object by the light field camera and calculates three-dimensional coordinates by the method.
[0064] The specific process is as shown in Figure 1 , and includes the following steps:
[0065] Step 1: Collecting a measured object, in this embodiment, an Illum light field camera is used to collect a doll located on an optical platform, and the model of the light field camera is as shown in Figure 2 . The distance between the doll and the camera is controlled by an electrically controlled translation stage, the image of the doll is shot, the image file is obtained by the camera, and the original light field image L o is obtained.
[0066] The original light field image L o is converted into a standard light field image L Figure 3 as shown by a light field image decoding program. s The hexagonal arrangement of L s is converted into a standard light field form in a square arrangement, and the standard light field image contains 625x434 sub-images, and each sub-image contains 7x7 pixels.
[0067] Step 2: Converting the light field image in the standard light field form into a sub-aperture image array, and the conversion process is as shown in Figure 4 . The pixel points in the first row and the first column in each sub-image are extracted and combined into a sub-aperture image, which represents a directional view image collected by the light field camera. The above operation is performed on all pixel points in the sub-image, and a sub-aperture image array is obtained, in this embodiment, each sub-aperture image obtained has a size of 623x434 pixels, and the sub-aperture image array contains scene information of 7x7 view angles.
[0068] The parallax of adjacent view angles is calculated from the extracted multi-view sub-aperture image array, in this embodiment, the parallax of adjacent view angles is calculated in a stereo matching manner, and the parallax calculation manner is as follows:
[0069] D = LabelUnit x Label x ViewIndex (1)
[0070] Where LabelUnit is the disparity unit, Label is the disparity calculation range, and ViewIndex is the view index of the sub-aperture. In multi-view sub-aperture images, the disparity is less than 1 pixel. When performing stereo matching on the light field image, the disparity cannot be directly calculated using Label. It is necessary to add a LabelUnit that is much smaller than 1 and the ViewIndex of the matching view.
[0071] Step 3: For the doll being tested, assume that the coordinates of a point m on it are [x, y, z]. T Representing the coordinates of point m using a four-dimensional spatial light field, the transformation conforms to the formula:
[0072]
[0073] Where [s,t] T Represents spatial coordinates, [u,v] T Representing angular coordinates, for a light field camera, the external light field [s,t,u,v] T The optical field distribution [i,j,k,l] on the internal sensor T The mapping relationship between them is as follows:
[0074]
[0075] Among them, h (·) Here, H is the intrinsic parameter matrix parameter, which can be obtained by the checkerboard calibration method. Combining formulas (12) and (13), the mapping relationship between spatial point m and the light field distribution of the camera sensor is obtained, as shown in formula (14):
[0076]
[0077] The formula defines the linear relationship between the sensor's light field coordinates k and i, and l and v, representing the epipolar equations of the sub-aperture image in the horizontal and vertical directions. The slope of the equation represents the parallax between adjacent viewpoints, such as... Figure 5 As shown. Combining the slope expression of the epipolar equation and its physical meaning, the mapping relationship between disparity and depth can be obtained, as shown in formula (15):
[0078]
[0079] Where D is the parallax of the center view of the sub-aperture image array.
[0080] Step 4: Calculate the three-dimensional coordinates using the disparity map, obtained from formulas (12) and (13).
[0081]
[0082] At the same time, the parallax map of the intermediate view point obtained from S2 contains a set of light field coordinates (0, 0, k, l), which are substituted into formula (16) to obtain the solving equation of (x, y) coordinates:
[0083]
[0084] The z coordinate obtained by combining the formula is used to restore the three-dimensional coordinates of the object represented by all the collected pixels in the light field image in this embodiment.
[0085] In order to better understand the technical solutions of the present application, Figure 1 The present application shows a schematic diagram of the three-dimensional positioning process based on the light field camera, Figure 4 The present application shows a schematic diagram of the sub-aperture image conversion in the embodiment. Figure 5 The present application shows a schematic diagram of the parallax-depth conversion. Figure 6 The present application shows a schematic diagram of the parallax-depth conversion. Figure 7 The present application shows a schematic diagram of the parallax-depth conversion.
[0086] As Figure 7 It can be seen that the edge texture of the parallax map based on the present application is clear, and the resolution of the parallax map is high, which reflects the advantages of the present application in high-resolution panoramic three-dimensional positioning.
[0087] Embodiment three:
[0088] The present application provides a three-dimensional positioning system based on a light field camera, which uses the three-dimensional positioning method based on a light field camera described in embodiment two to realize three-dimensional positioning of an object;
[0089] The system comprises an image acquisition module, an image processing module, and a coordinate positioning module.
[0090] The image acquisition module uses a light field camera to acquire a light field image of the measured object.
[0091] The image processing module converts the light field image acquired by the image acquisition module into a sub-aperture image array, and obtains a parallax map set with the center view angle of the sub-aperture image array as a reference.
[0092] The coordinate positioning module uses a light field camera to shoot a checkerboard image, obtains the mapping relationship from the spatial four-dimensional coordinates to the image sensor coordinates through a light field camera calibration program, forms a projection model, and then calculates the spatial three-dimensional coordinates of the measured object according to the projection model and the fitted parallax map obtained by the image processing module.
[0093] Some steps in the embodiment of the present application can be realized by software, and the corresponding software program can be stored in a readable storage medium, such as an optical disc or a hard disk.
[0094] The above description is only the preferred embodiment of the present application, and is not intended to limit the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.
Claims
1. A method for three-dimensional positioning based on a light field camera, characterized in that, The method comprises: Step 1: collecting the light field image of the measured object: using a light field camera to take an image of the measured object, decoding the obtained image to obtain a light field image and converting it into a standard light field description form; Step 2: disparity extraction: convert the light field image in step 1 into an array of sub-aperture images, and obtain a set of disparity maps M for each view relative to the center view of the array of sub-aperture images D ; Step 3: projection model calculation: using a light field camera to take a checkerboard image, obtaining the mapping relationship from the spatial three-dimensional point to the image sensor coordinates through a light field camera calibration program, and forming a projection model; Step 4: Three-dimensional coordinate calculation: the parallax map set M solved in step 2 D The unique parallax map D of the adjacent view angle is obtained by mean fitting, the parallax-depth conversion relationship of the sub-aperture image is obtained by combining the projection model of step 3, and the three-dimensional coordinates of the measured object are obtained. The step 3 comprises: Step 31: Establish camera coordinates [x, y, z] T To a four-dimensional coordinate system [u,v,s,t] T The mapping relationship is as follows: the camera coordinate system is a three-dimensional spatial coordinate system with the center of the camera's main lens as the origin; the four-dimensional spatial coordinate system is a four-dimensional spatial representation of the position coordinates [s,t] and orientation coordinates [u,v] of an object source point m(x,y,z) in a plane within this three-dimensional coordinate system. T Then the four-dimensional coordinates of point m are [s,t,u,v]. T The transformation relationship between the corresponding three-dimensional coordinates m(x,y,z) and the coordinates is shown in Equation (1). Wherein, λ is the distance from the arbitrary point to the origin of the camera coordinate system; Step 32: Obtain the spatial four-dimensional coordinate system [s, t, u, v] by chessboard calibration T to the conversion relationship of the sensor light field coordinates [i, j, k, l] T , wherein [s, t] represents the spatial coordinates, [u, v] represents the angle coordinates, [i, j] represents the microlens coordinates, and [k, l] represents the local pixel coordinates corresponding to each microlens on the sensor, as shown in formula (2): where h (·) is an intrinsic matrix parameter; Step 33: Combining equation (1) and equation (2) can get the mapping equation of three-dimensional space coordinates [x, y, z] to sensor light field coordinates [i, j, k, l], as shown in equation (3): T T Equation (3) The step 4 comprises: Step 41: the linear slope of the two formulas in formula (3) represents the parallax of adjacent viewing angles, and the mapping relationship between the parallax and the depth is: Wherein, z is the depth coordinate in the three-dimensional coordinates; Step 42: obtaining a set of space-angle coordinates [i, k] from the disparity map D T or [j, l] T ; Step 43: calculating the (x, y) coordinates in the three-dimensional coordinates according to formula (3), and the calculation formula is shown in formula (5):
2. The method of claim 1, wherein, The step 1 comprises: Step 11: The light field camera collects the measured object to obtain an initial light field image L o ; Step 12: using a light field image decoding procedure, decode the initial light field image L o correcting the initial light field image L s ; Step 13: Rearranging the standard light field image L s Converting the hexagonal light field image into a standard four- dimensional light field form arranged in a square.
3. The method of claim 1, wherein, The step 2 comprises: Step 21: extracting the pixel points at the same position of all microlens sub-images in the light field image, and then combining them into a sub-aperture image; performing the same operation on all position pixel points of the sub-image to obtain a light field sub-aperture image array; Step 22: taking the sub-aperture image corresponding to the center view point of the sub-aperture image array as the reference view, calculating the disparity of other views relative to the reference view by including stereo matching method, optical flow method and machine learning image processing method, obtaining the disparity map set M of all views relative to the center view D ; Step 23: combining the disparity map set M obtained in step 22 D A disparity map D representing two adjacent views is calculated by fitting the mean value.
4. The method of claim 3, wherein, The method for calculating the parallax of each viewing angle relative to the central viewing angle is a stereo matching method.
5. The method of claim 3, wherein, The method for calculating the parallax of each viewing angle relative to the central viewing angle is a light flow method or a deep learning method.
6. A light field camera based three-dimensional positioning system, characterized by The system adopts the three-dimensional positioning method based on the light field camera according to any one of claims 1-5 to realize three-dimensional positioning of the object. The system comprises an image acquisition module, an image processing module and a coordinate positioning module. The image acquisition module adopts a light field camera to collect the light field image of the measured object. The image processing module converts the light field image collected by the image acquisition module into a sub-aperture image array, and takes the central viewing angle of the sub-aperture image array as a reference to obtain a parallax map set; The coordinate positioning module uses a light field camera to take a checkerboard image, obtains the mapping relationship from the spatial four-dimensional coordinates to the image sensor coordinates through a light field camera calibration program, forms a projection model, and then calculates the spatial three-dimensional coordinates of the measured object according to the projection model and the fitted parallax map obtained by the image processing module.
7. A camera positioning apparatus, characterized by, The camera positioning device adopts the three-dimensional positioning system based on the light field camera according to claim 6 to realize positioning of the measured object.
Citation Information
Patent Citations
Depth estimation method and system based on light field image angle domain pixels
CN107038719A
Webpage end 3D (three-dimensional) model implementation method based on Cook-Torrance algorithm
CN112184889A
Image processing device and three-dimensional measuring system
US20220012905A1