Three-dimensional reconstruction method and device based on biomimetic stereo vision and storage medium

By using a biomimetic stereo vision approach, texture information is extracted using a spherical camera sensor and a multi-layered window frame mesh to generate a dense disparity map. This solves the problem of low feature point matching rate in stereo vision 3D reconstruction using traditional dual-camera devices, and achieves a more efficient 3D reconstruction effect.

CN115965677BActive Publication Date: 2025-12-23张国流
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202210295307.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-03-24
Publication Date
2025-12-23
Estimated Expiration
2042-03-24

AI Technical Summary

Technical Problem

Traditional dual-camera devices are limited by the parallel view pattern and the difficulty of feature point matching in stereo vision 3D reconstruction, resulting in limited focusing function and low feature point matching rate, which affects the 3D reconstruction effect.

Method used

A biomimetic stereo vision-based approach is adopted, using left and right camera sensors with spherical structures to form kappa angles. Texture information is extracted through a multi-layer window frame mesh to generate a dense disparity map, calculate the 3D coordinates of common feature points and estimate the 3D coordinates of non-common feature points.

Benefits of technology

It improves the feature point matching rate, enhances the flexibility of stereo vision and the density of 3D reconstruction, and achieves more efficient 3D reconstruction results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115965677B_ABST
    Figure CN115965677B_ABST
Patent Text Reader

Abstract

The application provides a three-dimensional reconstruction method and device based on biomimetic stereovision and a storage medium, the method comprising: acquiring left and right images captured by a double-camera device when focusing on an interest target; performing texture extraction on the left and right images respectively to obtain a total texture map of the left image and a total texture map of the right image; performing feature matching on the total texture map of the left image and the total texture map of the right image to generate a dense disparity map; calculating three-dimensional coordinates of common feature points matched in the dense disparity map, and estimating three-dimensional coordinates of non-common feature points according to the three-dimensional coordinates of the common feature points to obtain a three-dimensional coordinate dataset; and performing three-dimensional reconstruction according to the three-dimensional coordinate dataset to generate a stereoscopic view of the interest target; the application can effectively improve the matching rate of feature points and improve the stereoscopic effect.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of stereovision technology, and in particular to a three-dimensional reconstruction method and device based on biomimetic stereovision and a storage medium. BACKGROUND

[0002] Current stereovision and three-dimensional reconstruction are mostly used for processing non-time-parallel images, that is, using multiple frames of images taken by the same camera at different time points and different angles as raw materials to calculate the depth of the photographed object and restore the three-dimensional structure. In addition, there are also techniques for generating real-time three-dimensional scenes using dual cameras, such as machine navigation. This technique uses two parallel image dual cameras to take real-time images of the same scene at different angles, and then uses triangulation to convert the parallax information of the dual cameras into depth information of the object, achieving real-time three-dimensional reconstruction.

[0003] The core of binocular stereovision is to use the parallax information of left and right images to calculate the depth information of the object. Parallax information is the difference between the imaging position information of an object in the left camera and the imaging position information in the right camera. In parallel images, the feature points of the same object will fall on the same scan line in the left and right images. The horizontal coordinate of the feature point on the left image minus the horizontal coordinate on the right image is the parallax of the feature point. Then, using triangulation, the depth value of the feature point from the dual cameras can be directly calculated.

[0004] Triangulation is the simplest method for depth calculation, but the prerequisite for using triangulation to calculate depth information is to require that the dual cameras use a parallel view pattern and that the imaging positions of the feature points in the two images can be accurately obtained. The application of traditional dual camera devices is limited by these two conditions, resulting in obvious performance deficiencies. On the one hand, traditional dual camera devices need to deflect the camera angle for focusing, which breaks the parallel view pattern and makes triangulation no longer valid. To avoid this problem, traditional dual camera devices often do not have focusing functions. On the other hand, due to the similar structure and color of different points of the same object, it is easy to mistakenly match similar feature points in actual situations. The position and angle difference of the dual cameras also easily leads to differences in the imaging results of the same object in the left and right views, making it impossible to achieve matching. Therefore, the traditional dual camera has not been very effective in obtaining feature matching and obtaining dense disparity maps. SUMMARY

[0005] To solve the above problems, the present application provides a three-dimensional reconstruction method and device based on biomimetic stereovision and a storage medium, which can have a focusing function and effectively improve the matching rate of feature points, and improve the flexible stereovision effect of the device.

[0006] In a first aspect, the embodiments of the present application provide a three-dimensional reconstruction method based on biomimetic stereo vision, comprising:

[0007] obtaining left and right images taken by a dual-camera device when focusing on an object of interest;

[0008] performing texture extraction on the left and right images respectively to obtain total texture maps of the left and right images;

[0009] performing feature matching on the total texture maps of the left and right images to generate a dense disparity map;

[0010] calculating three-dimensional coordinates of common feature points matched in the dense disparity map, and estimating three-dimensional coordinates of non-common feature points according to the three-dimensional coordinates of the common feature points to obtain a three-dimensional coordinate dataset;

[0011] performing three-dimensional reconstruction according to the three-dimensional coordinate dataset to generate a stereo view of the object of interest.

[0012] As an improvement of the above scheme, the dual-camera device comprises a left camera and a right camera, the left camera is provided with a left sensor, and the right camera is provided with a right sensor; wherein the photosensitive wafers of the left sensor and the right sensor are both spherical structures, and the inner side of the sphere is the signal receiving surface; the center point of the left sensor is located on the left side of the optical lens main axis of the left camera, and the center point of the right sensor is located on the right side of the optical lens main axis of the right camera, so that the optical lens main axis of the left camera and the left visual axis, and the optical lens main axis of the right camera and the right visual axis form a kappa angle respectively, and the center point section of the left sensor and the center point section of the right sensor always remain parallel during the shooting process.

[0013] As an improvement of the above scheme, obtaining left and right images taken by a dual-camera device when focusing on an object of interest comprises:

[0014] setting a camera coordinate system: taking a straight line where the center points of the left sensor and the right sensor are located as the horizontal coordinate axis of the camera coordinate system; taking the midpoint of the line connecting the center points of the left sensor and the right sensor as the origin of the camera coordinate system; taking a straight line where the midpoint of the line connecting the lens optical center of the left camera and the lens optical center of the right camera is located as the depth coordinate axis of the camera coordinate system; and taking a straight line passing through the origin of the camera coordinate system and perpendicular to the horizontal coordinate axis and the depth coordinate axis as the vertical coordinate axis of the camera coordinate system;

[0015] rotating the dual-camera device to align the depth coordinate axis of the camera coordinate system with the object of interest;

[0016] Adjusting focal length of the left camera and the right camera respectively to make the target of interest project on the center point of the left sensor and the right sensor respectively, and locking the focal length to focus on the target of interest;

[0017] Obtaining a left image captured by the left camera and a right image captured by the right camera after focusing.

[0018] As an improvement of the above scheme, the extracting the total texture map of the left image and the total texture map of the right image comprises:

[0019] The total texture map of the left image and the total texture map of the right image are extracted by using a multi-layer window frame grid side-antagonistic extraction method.

[0020] As an improvement of the above scheme, the using a multi-layer window frame side-antagonistic extraction method to extract the total texture map of the left image and the total texture map of the right image respectively comprises:

[0021] The color value of each pixel of the left image and the right image is converted into a digital array by a conversion method, wherein the base of the color value of the converted pixel is lower than the base of the color value of the unconverted pixel.

[0022] According to the digital array of each pixel of the left image and the right image, the left image and the right image are decomposed into a plurality of layers.

[0023] The texture information of each layer is extracted by using a window mask to obtain a sub-texture map of each layer.

[0024] The total texture map of the left image is synthesized by superimposing the weight values of all the sub-texture maps of the left image.

[0025] The total texture map of the right image is synthesized by superimposing the weight values of all the sub-texture maps of the right image.

[0026] As an improvement of the above scheme, the feature matching of the total texture map of the left image and the total texture map of the right image to generate a dense disparity map comprises:

[0027] The texture blocks of the total texture map of the left image and the total texture map of the right image are used as feature points for feature matching to obtain a dense disparity map, wherein the texture blocks are right-angled triangles, and the directions of the four adjacent texture blocks in the same row are different.

[0028] As an improvement of the above scheme, the method further comprises:

[0029] For each total texture map, a region of pixel size N*N is divided into a herringbone shape to obtain a plurality of right-angled triangular texture patches.

[0030] As an improvement of the above scheme, the calculation of the three-dimensional coordinates of the matched common feature points in the dense disparity map comprises:

[0031] According to the structure parameters of the dual-camera device and the left-view pixel coordinates of the common feature points on the left sensor and the right-view pixel coordinates of the common feature points on the right sensor, the three-dimensional coordinates of the common feature points in the camera coordinate system are calculated; wherein the structure parameters include a first distance between the optical lens main optical axis of the left camera and the optical lens main optical axis of the right camera, a second distance between the center point of the left sensor and the optical lens main optical axis of the right camera, a lens distance, a focal length, the number of texture patches of each row of the left sensor, and the horizontal width of the texture patches.

[0032] According to the position parameters and the attitude parameters of the dual-camera device, the three-dimensional coordinates of the common feature points in the camera coordinate system are transformed into three-dimensional coordinates in the world coordinate system; wherein the attitude parameters include the horizontal coordinate axis deflection amount, the vertical coordinate axis deflection amount, and the depth coordinate axis deflection amount of the camera coordinate system of the dual-camera device relative to the world coordinate system; and the position parameters include the horizontal distance, the vertical distance, and the depth distance of the origin of the camera coordinate system of the dual-camera device relative to the origin of the world coordinate system.

[0033] As an improvement of the above scheme, the calculation of the three-dimensional coordinates of the matched common feature points in the dense disparity map comprises:

[0034] According to the left-view pixel coordinates of the common feature points on the left sensor and the right-view pixel coordinates of the common feature points on the right sensor, the disparity, the horizontal offset amount, and the vertical offset amount of the common feature points are calculated.

[0035] According to the first distance, the second distance, the number of texture patches, the horizontal width, and the disparity and the horizontal offset amount of the common feature points, the horizontal coordinates of the common feature points in the camera coordinate system are calculated.

[0036] According to the first distance, the second distance, the number of texture patches, the horizontal width, and the disparity and the vertical offset amount of the common feature points, the vertical coordinates of the common feature points in the camera coordinate system are calculated.

[0037] According to the first distance, the second distance, the focal length, the lens distance, the horizontal width, and the parallax of the common feature point, depth coordinates of the common feature point in a camera coordinate system are calculated.

[0038] Therefore, the transformation of the three-dimensional coordinates of the common feature point in the camera coordinate system into the three-dimensional coordinates in the world coordinate system according to the position parameters and the attitude parameters of the dual-camera device comprises:

[0039] According to the attitude parameters and the position parameters, the horizontal coordinates, the vertical coordinates, and the depth coordinates of the common feature point in the camera coordinate system are transformed into the three-dimensional coordinates of the common feature point in the world coordinate system.

[0040] As an improvement of the above-mentioned scheme, the calculation of the horizontal coordinates of the common feature point in the camera coordinate system according to the first distance, the second distance, the number of texture blocks, the horizontal width, and the parallax and the horizontal offset of the common feature point comprises:

[0041] The horizontal coordinates H of the common feature point are calculated according to formula (1) i ;

[0042]

[0043] wherein e represents the horizontal width of the texture block, n represents the number of texture blocks per row of the left / right sensor, b represents the first distance between the optical lens principal axis of the left camera and the optical lens principal axis of the right camera, d represents the second distance between the center point of the left sensor and the optical lens principal axis of the left camera, and between the center point of the right sensor and the optical lens principal axis of the right camera, (L i , W i ) represents the left view pixel coordinates of the common feature point i on the left sensor, (R i , W i ) represents the right view pixel coordinates of the common feature point i on the right sensor, (L i -R i e represents the parallax of the common feature point i, -L i -R i represents the horizontal offset of the common feature point i.

[0044] As an improvement of the above-mentioned scheme, the calculation of the vertical coordinates of the common feature point in the camera coordinate system according to the first distance, the second distance, the number of texture blocks, the horizontal width, and the parallax and the vertical offset of the common feature point comprises:

[0045] The vertical coordinates V of the common feature point are calculated according to formula (2)i ;

[0046]

[0047] wherein, W i represents the vertical offset of the common feature point i.

[0048] As an improvement of the above-mentioned scheme, the calculating the depth coordinate of the common feature point in the camera coordinate system according to the first distance, the second distance, the focal length, the mirror distance, the horizontal width and the parallax of the common feature point comprises:

[0049] The depth coordinate D of the common feature point is calculated according to formula (3) i ;

[0050]

[0051] wherein, v represents the mirror distance, and f represents the focal length.

[0052] As an improvement of the above-mentioned scheme, the estimating the three-dimensional coordinates of the non-common feature points according to the three-dimensional coordinates of the common feature points comprises:

[0053] The three-dimensional coordinates of the non-common feature points are estimated by using a single feature point estimation method according to the three-dimensional coordinates of the common feature points.

[0054] Or,

[0055] The three-dimensional coordinates of the non-common feature points are estimated by using a double feature point estimation method according to the three-dimensional coordinates of the common feature points.

[0056] As an improvement of the above-mentioned scheme, the estimating the three-dimensional coordinates of the non-common feature points according to the three-dimensional coordinates of the common feature points by using a single feature point estimation method comprises:

[0057] The three-dimensional coordinates of the non-common feature points which are not matched successfully are estimated by using a single feature point estimation method according to the feature point double gradient law in the total texture map of the left image and the total texture map of the right image.

[0058] For the non-common feature point, a common feature point adjacent to the non-common feature point and having a color difference ratio within a set color difference range is obtained as a reference point.

[0059] The depth coordinate of the reference point is taken as the depth coordinate of the non-common feature point in the camera coordinate system.

[0060] The horizontal coordinate of the non-common feature point in the camera coordinate system is calculated by formula (1) after the horizontal offset of the reference point is added or subtracted by 1; wherein, when the non-common feature point is on the left side of the common feature point, the horizontal offset of the reference point is subtracted by 1, and when the non-common feature point is on the right side of the common feature point, the horizontal offset of the reference point is added by 1.

[0061] The vertical coordinate of the non-common feature point in the camera coordinate system is calculated by formula (2) after the vertical offset of the reference point is added or subtracted by 1; wherein, when the non-common feature point is above the common feature point, the vertical offset of the reference point is subtracted by 1, and when the non-common feature point is below the common feature point, the vertical offset of the reference point is added by 1.

[0062] As an improvement of the above scheme, the three-dimensional coordinates of the non-common feature point are estimated by using a double feature point estimation method according to the three-dimensional coordinates of the common feature point, comprising:

[0063] The three-dimensional coordinates of the non-common feature point which fails to be matched are estimated by using a double feature point estimation method according to the double gradient rules of the feature points in the total texture map of the left image and the total texture map of the right image;

[0064] For the non-common feature point, two common feature points in the left image or the right image which have the same distance but opposite directions with the non-common feature point and have color difference ratios in a set color difference range are obtained as reference points;

[0065] The mean value of the depth coordinates of the two reference points is calculated as the depth coordinate of the non-common feature point in the camera coordinate system;

[0066] The mean value of the horizontal coordinates of the two reference points is calculated as the horizontal coordinate of the non-common feature point in the camera coordinate system;

[0067] The mean value of the vertical coordinates of the two reference points is calculated as the vertical coordinate of the non-common feature point in the camera coordinate system.

[0068] In a second aspect, an embodiment of the present application provides a three-dimensional reconstruction device based on bionic stereo vision, comprising:

[0069] An image acquisition module is configured to acquire a left image and a right image captured by a double camera device when focusing on an interest target;

[0070] A texture extraction module is configured to extract textures from the left image and the right image respectively to obtain a total texture map of the left image and a total texture map of the right image;

[0071] The feature matching module is configured to perform feature matching on the total texture map of the left image and the total texture map of the right image, and generate a dense disparity map;

[0072] The three-dimensional coordinate calculation module is configured to calculate three-dimensional coordinates of the common feature points matched in the dense disparity map, and estimate three-dimensional coordinates of non-common feature points according to the three-dimensional coordinates of the common feature points, to obtain a three-dimensional coordinate dataset;

[0073] The three-dimensional reconstruction module is configured to perform three-dimensional reconstruction according to the three-dimensional coordinate dataset, and generate a stereoscopic view of the target of interest.

[0074] In a third aspect, an embodiment of the present application provides a three-dimensional reconstruction device based on biomimetic stereovision, comprising a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor, and the processor implements the three-dimensional reconstruction method based on biomimetic stereovision according to any one of the first aspect.

[0075] In a fourth aspect, an embodiment of the present application provides a computer readable storage medium, comprising a stored computer program, wherein the computer readable storage medium controls a device where the computer readable storage medium is located to execute the three-dimensional reconstruction method based on biomimetic stereovision according to any one of the first aspect when the computer program runs.

[0076] Compared with the prior art, the embodiment of the present application has the beneficial effects that: the left image and the right image captured when a double-camera device focuses on a target of interest are obtained; total texture maps of the left image and the right image are obtained by performing texture extraction on the left image and the right image, respectively; a dense disparity map is generated by performing feature matching on the total texture map of the left image and the total texture map of the right image; a three-dimensional coordinate dataset is obtained by calculating three-dimensional coordinates of common feature points matched in the dense disparity map; and a stereoscopic view of the target of interest is generated by performing three-dimensional reconstruction according to the three-dimensional coordinate dataset. The three-dimensional coordinates of other feature points are estimated by the three-dimensional coordinates of the common feature points, which can improve the density of the depth map and improve the stereoscopic visual effect. BRIEF DESCRIPTION OF DRAWINGS

[0077] In order to more clearly illustrate the technical solutions of the present application, the following will briefly introduce the drawings needed in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can also be obtained by those skilled in the art without creative labor.

[0078] Figure 1 is a flowchart of a three-dimensional reconstruction method based on biomimetic stereovision provided by an embodiment of the present application;

[0079] Figure 2 is a top view of a dual-camera device provided by an embodiment of the present application;

[0080] Figure 3 is a schematic diagram of a camera coordinate system provided by an embodiment of the present application;

[0081] Figure 4 is a schematic diagram of a comparison between an image pixel and a texture patch provided by an embodiment of the present application;

[0082] Figure 5 is a schematic diagram of a same-feature interval and depth difference provided by an embodiment of the present application;

[0083] Figure 6 is a schematic diagram of left and right sensors in parallel in a pixel coordinate system provided by an embodiment of the present application;

[0084] Figure 7 is a schematic diagram of vertical coordinate calculation provided by an embodiment of the present application;

[0085] Figure 8 is a schematic diagram of horizontal coordinate calculation provided by an embodiment of the present application;

[0086] Figure 9 is a schematic diagram of coordinate system conversion provided by an embodiment of the present application;

[0087] Figure 10 is a schematic diagram of layer decomposition provided by an embodiment of the present application;

[0088] Figure 11 is a schematic diagram of a window frame mask provided by an embodiment of the present application;

[0089] Figure 12 is a schematic diagram of opposite detection points of a window frame provided by an embodiment of the present application;

[0090] Figure 13 is a schematic diagram of texture filling provided by an embodiment of the present application;

[0091] Figure 14 is a schematic diagram of layer merging provided by an embodiment of the present application;

[0092] Figure 15 is a schematic diagram of image texture information provided by an embodiment of the present application;

[0093] Figure 16 is a schematic diagram of layer weight superposition provided by an embodiment of the present application;

[0094] Figure 17 is a schematic diagram of a three-dimensional reconstruction device based on biomimetic stereovision provided by an embodiment of the present application;

[0095] Figure 18is a structural schematic diagram of a three-dimensional reconstruction device based on bionic stereovision provided by an embodiment of the present application. DETAILED DESCRIPTION

[0096] The technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by a person of ordinary skill in the art without creative work fall within the protection scope of the present application.

[0097] Please refer to Figure 1 The first embodiment of the present application provides a three-dimensional reconstruction method based on bionic stereovision, comprising:

[0098] S1: acquiring left and right images photographed when a double-camera device focuses on an interest target;

[0099] Further, the double-camera device comprises a left camera and a right camera, the left camera is provided with a left sensor, and the right camera is provided with a right sensor; wherein the photosensitive wafer of the left sensor and the photosensitive wafer of the right sensor are both spherical structures, and the spherical concave side is the signal receiving surface; the center point of the left sensor is located on the left side of the optical lens main optical axis of the left camera, and the center point of the right sensor is located on the right side of the optical lens main optical axis of the right camera, so that the optical lens main optical axis of the left camera and the left visual axis and the optical lens main optical axis of the right camera and the right visual axis form a kappa angle respectively, and the center point section of the left sensor and the center point section of the right sensor always remain parallel during the photographing process.

[0100] The left camera and the right camera are fixed by a rigid body structure to constitute a head of the device; during the photographing process, the device head rotates as a whole with the camera coordinate system origin as the axis, and the orientations of the left and right cameras always coincide; the center point section of the left sensor and the center point section of the right sensor always remain parallel to each other.

[0101] Further, acquiring left and right images photographed when a double-camera device focuses on an interest target comprises:

[0102] The camera coordinate system is defined as follows: the horizontal axis of the camera coordinate system is the straight line connecting the center points of the left and right sensors; the origin of the camera coordinate system is the midpoint of the line connecting the center points of the left and right sensors; the depth axis of the camera coordinate system is the straight line connecting the origin of the camera coordinate system, the midpoint of the line connecting the optical center of the lens of the left camera and the optical center of the lens of the right camera; and the vertical axis of the camera coordinate system is the straight line passing through the origin of the camera coordinate system and perpendicular to both the horizontal and depth axes.

[0103] Rotate the dual-camera device to align the depth coordinate axis of the camera coordinate system with the target of interest;

[0104] The focal lengths of the left and right cameras are adjusted so that the target of interest is projected onto the center points of the left and right sensors, respectively, and the focal lengths are locked to focus on the target of interest.

[0105] The left image captured by the left camera and the right image captured by the right camera are obtained after focusing.

[0106] For example, such as Figure 2 and Figure 3 As shown, the center point of the left sensor is offset to the left of the main optical axis of the left camera's lens, and the center point of the right sensor is offset to the right of the main optical axis of the right camera's lens, so that the main optical axis of the left camera's lens and the viewing axis form a kappa angle of θ, and the main optical axis of the right camera's lens and the viewing axis also form a kappa angle of θ. At this time, the midpoint of the line connecting the center points of the left and right sensors is the origin C of the camera coordinate system. The straight line passing through the center points K1 and K2 of the left and right sensors is the horizontal axis of the camera coordinate system, defined as the H-axis; the straight line passing through the origin C and point O is the depth coordinate axis, defined as the D-axis; point O is the straight line connecting the optical centers O1 and O2 of the left and right cameras' lenses; and the straight line passing through the origin C and perpendicular to both the H-axis and D-axis is the vertical coordinate axis, defined as the V-axis. Therefore, the depth coordinate axis D coincides with the alignment line. The dual-camera device has the following structural features:

[0107] (1) The dual-camera device consists of a left camera, a right camera, an image information processing module, and several structural supports with a specific structure.

[0108] (2) The left camera is fixed with the right camera rigid structure, constituting the head of the device; during the shooting process, the device head rotates as a whole with the camera coordinate system origin as the axis, and the orientations of the left and right cameras are always consistent; the left sensor center point section and the right sensor center point section are always parallel to each other.

[0109] (3) The sensors of the left and right cameras both adopt spherical photosensitive wafers, and the inner side of the sphere is used as the signal receiving surface; the shapes and sizes of the left and right sensors are the same;

[0110] (4) The sensors of the left and right cameras are respectively translated to the temporal side of the optical lens main optical axis, and the center points of the sensors are located on the temporal side of the optical lens main optical axis of the corresponding camera; the optical axis of the camera and the main optical axis of the optical lens form a Kappa angle;

[0111] (5) The double-camera device takes the midpoint C of the line connecting the center points K1 and K2 of the left and right sensors as the camera coordinate system origin, and takes the extension line of the line connecting the points C and O as the alignment line, wherein the point O is the midpoint of the line connecting the lens centers O1 and O2 of the left and right cameras.

[0112] (6) When performing a visual task, the head of the double-camera device rotates as a whole with the origin point C as the axis to achieve the locking of different targets.

[0113] Through the above structure, it can be ensured that the device focuses on the target of interest while maintaining the parallel view pattern. After determining the target of interest P to be focused on, the double-camera device first adjusts its camera coordinate system by moving the position or rotating the head, so that the D-axis of the camera coordinate system is aligned with the target of interest. After the target of interest is located on the D-axis, the left and right cameras adjust the focal length through the optical lens until the target of interest P is projected on the center points K1 and K2 of the left and right sensors respectively. At this time, the intersection of the optical axis and the main optical axis of the optical lens is the focal point position F1 and F2. After locking the focal length, the focusing on the target of interest is achieved. After completing the focusing on the target of interest, the left and right cameras respectively shoot the current scene to obtain left and right images.

[0114] S2: Extracting textures from the left image and the right image respectively to obtain a total texture map of the left image and a total texture map of the right image;

[0115] S3: Performing feature matching on the total texture map of the left image and the total texture map of the right image to generate a dense disparity map;

[0116] Further, the feature matching on the total texture map of the left image and the total texture map of the right image to generate a dense disparity map comprises:

[0117] S3: performing feature matching on the texture patches of the total texture map of the left image and the total texture map of the right image as feature points to obtain a dense disparity map; wherein the texture patches are right-angled triangles, and the orientations of four adjacent texture patches in the same row are different.

[0118] wherein, for each total texture map, a region with a pixel size of N*N is H-shaped segmented to obtain a plurality of right-angled triangular texture patches.

[0119] For example, there are not only differences in color values between the texture patches, but also differences in shapes. Figure 4 As shown, H-shaped segmentation is performed on a region with a pixel size of 3*3 to obtain 8 right-angled triangular texture patches, wherein the orientations of the four texture patches in the upper half region are different, and at this time, the size of each texture patch is 9 / 8 of a pixel, close to a single-pixel size. If the image captured by the camera is 1 million pixels, then the texture patch matching can theoretically generate a dense disparity map with 88.9 million disparity values, which is much better than the SIFT algorithm that can only generate a few hundred disparity values and the single-pixel matching with a very high mismatch rate. The texture patch matching scheme adopted in this embodiment can not only ensure a low mismatch rate, but also ensure dense data output, and has obvious performance advantages.

[0120] S4: calculating three-dimensional coordinates of the common feature points matched in the dense disparity map, and estimating three-dimensional coordinates of non-common feature points according to the three-dimensional coordinates of the common feature points to obtain a three-dimensional coordinate dataset;

[0121] S5: performing three-dimensional reconstruction according to the three-dimensional coordinate dataset to generate a stereoscopic view of the target of interest.

[0122] In the embodiment of the application, the views captured by the left and right cameras are first subjected to texture extraction processing to generate total texture maps, and then the texture patches of the total texture maps are used as feature point units for matching. The texture patches have the performance advantages of small granularity and low similarity, so that the correctness of feature point matching can be ensured while the density of the depth map is effectively improved, a dense depth map close to the pixel level is generated, and better three-dimensional reconstruction effects are achieved, thereby improving the stereoscopic view effect.

[0123] In an optional embodiment, the calculation of the three-dimensional coordinates of the common feature points matched in the dense disparity map comprises:

[0124] According to the structural parameters of the dual-camera device and the left-view pixel coordinates of the common feature point on the left sensor and the right-view pixel coordinates of the common feature point on the right sensor, the three-dimensional coordinates of the common feature point in the camera coordinate system are calculated; wherein the structural parameters include a first distance between the optical lens main optical axis of the left camera and the optical lens main optical axis of the right camera, a second distance from the center point of the left sensor to the optical lens main optical axis of the left camera, a lens distance, a focal length, a number of texture color blocks per row of the left sensor, a horizontal width of the texture color blocks;

[0125] According to the position parameters and the attitude parameters of the dual-camera device, the three-dimensional coordinates of the common feature point in the camera coordinate system are transformed into three-dimensional coordinates in the world coordinate system to obtain a three-dimensional coordinate dataset; wherein the attitude parameters include a horizontal coordinate axis deflection amount, a vertical coordinate axis deflection amount, and a depth coordinate axis deflection amount of the camera coordinate system of the dual-camera device relative to the world coordinate system; and the position parameters include a horizontal distance, a vertical distance, and a depth distance of the origin of the camera coordinate system of the dual-camera device relative to the origin of the world coordinate system.

[0126] Further, the calculation of the three-dimensional coordinates of the common feature point in the camera coordinate system according to the structural parameters of the dual-camera device and the left-view pixel coordinates of the common feature point on the left sensor and the right-view pixel coordinates of the common feature point on the right sensor includes:

[0127] According to the left-view pixel coordinates of the common feature point on the left sensor and the right-view pixel coordinates of the common feature point on the right sensor, the disparity, the horizontal offset amount, and the vertical offset amount of the common feature point are calculated;

[0128] According to the first distance, the second distance, the number of texture color blocks, the horizontal width, and the disparity and the horizontal offset amount of the common feature point, the horizontal coordinates of the common feature point in the camera coordinate system are calculated;

[0129] The horizontal coordinates H of the common feature point are calculated according to formula (1) i ;

[0130]

[0131] wherein e represents the horizontal width of the texture color block, n represents the number of texture color blocks per row of the left / right sensor, b represents the first distance between the optical lens main optical axis of the left camera and the optical lens main optical axis of the right camera, d represents the second distance from the center point of the left sensor to the optical lens main optical axis of the left camera, (L i , W i ) represents the left-view pixel coordinates of the common feature point i on the left sensor, (Ri , W i ) represents the right view pixel coordinate of the common feature point i on the right sensor, (L i -R i )e represents the disparity of the common feature point i, -L i -R i represents the horizontal offset of the common feature point i.

[0132] wherein the distance from the center point of the right sensor to the optical lens principal axis of the right camera is equal to the second distance.

[0133] According to the first distance, the second distance, the number of texture blocks, the horizontal width, and the disparity and vertical offset of the common feature point, the vertical coordinate of the common feature point in the camera coordinate system is calculated.

[0134] According to formula (2), the vertical coordinate V i of the common feature point is calculated.

[0135]

[0136] wherein W i represents the vertical offset of the common feature point.

[0137] According to the first distance, the second distance, the focal length, the lens distance, the horizontal width, and the disparity of the common feature point, the depth coordinate of the common feature point in the camera coordinate system is calculated.

[0138] According to formula (3), the depth coordinate D i of the common feature point is calculated.

[0139]

[0140] wherein v represents the lens distance, and f represents the focal length.

[0141] Then, the method for transforming the three-dimensional coordinates of the common feature point in the camera coordinate system into the three-dimensional coordinates in the world coordinate system according to the position parameters and the attitude parameters of the dual-camera device comprises:

[0142] According to the attitude parameters and the position parameters, the horizontal coordinate, the vertical coordinate, and the depth coordinate of the common feature point in the camera coordinate system are transformed in the coordinate system, so as to obtain the three-dimensional coordinates of the common feature point in the world coordinate system.

[0143] Specifically, according to the horizontal coordinate axis deflection amount, the vertical coordinate axis deflection amount, the depth coordinate axis deflection amount in the attitude parameter and the horizontal distance, the vertical distance, the depth distance in the position parameter, a rotation matrix R and a translation matrix T can be constructed. Through the conversion of the rotation matrix R and the translation matrix T, the three-dimensional coordinates of the common feature point i in the world coordinate system can be obtained: That is, the three-dimensional coordinates P of all the common feature points in the world coordinate system under the common field of view can be obtained in real time i (X i , Y i , Z i ), that is, a three-dimensional image of the overall field of view is generated in real time.

[0144] After obtaining the three-dimensional coordinate data set of all the common feature points through the above steps, the three-dimensional reconstruction of the field of view scene can be realized based on the three-dimensional coordinate data, so as to generate a stereoscopic view. Since the conversion of the image into the total texture map, the texture block matching, the calculation of the disparity to obtain the dense disparity map, and the calculation of the three-dimensional coordinates of the common feature points by using the fixed formula do not involve complex operations, only programmed operation is needed to quickly complete, therefore, the embodiment of the present application can quickly complete the three-dimensional reconstruction, and realize the effect of generating a stereoscopic view in real time.

[0145] In combination with Figure 2 , the derivation process of the depth coordinates of the common feature points is as follows:

[0146] Suppose that P is the target of interest, and K1 and K2 are the center points of the left and right sensors respectively. The line K1K2 is drawn, and the intersection point of the optical lens main axis of the left and right cameras with the straight line K1K2 is called the camera center point, which is represented by C1 and C2 respectively; the center point of the left sensor is located on the left side of the corresponding left camera center point, and the center point of the right sensor is located on the right side of the corresponding right camera center point; the distance from the camera bottom center point is a fixed value, which is represented by d, that is, K1C1=K2C2=d. The angle between the left and right optical lens main axes and the left and right optical axes is denoted as θ. The line connecting the center points of the left and right cameras is called the baseline, which is represented by b. C is the midpoint of the baseline K1K2; O1 and O2 are the lens centers of the left and right cameras respectively, and the distance between the lens center and the corresponding camera center point is called the lens distance, which is represented by v, that is, C1O1=C2O2=v. The line connecting O1O2 is called the pupil distance line, and O is the midpoint of the pupil distance line. The extension line of the line connecting C and O is called the alignment line.

[0147] In the camera coordinate system, the coordinates of the C point are (0, 0, 0), and the coordinates of the P point are (0, 0, D0). The three-dimensional coordinates of the P point in the camera coordinate system are obtained by solving P in the depth coordinate D0.

[0148] Since CP / / C2F2, there is ∠CPK2=∠C2F2K2=θ;

[0149] In RtΔCPK2, ∠PCK2=90°, ∠CPK2=θ, CK2=(b+2d) / 2;

[0150] Therefore, the distance D0 of P point is:

[0151] D0=CP=(b+2d) / (2tanθ) (I)

[0152] From RtΔC2F2K2, we have:

[0153] tanθ=d / (v-f) (II)

[0154] Substituting (II) into (I), we have D0=((b+2d)(v-f)) / 2d (III)

[0155] The formula (III) is the depth coordinate calculation formula of the target of interest. Wherein, b, v, d are all the structural parameters of the double camera device, which are fixed values. Therefore, the distance D0 of P point is only related to the focal length f of the optical lens. After obtaining the focal length parameter, the depth coordinate of the target of interest can be directly calculated by formula (III).

[0156] It can be known from formula (III) that the depth coordinate value of the target of interest can be obtained at the moment of focusing the target of interest by the double camera device, and only the focal length needs to be adjusted, so that the switching focusing of different depth distances can be realized, thereby realizing the focusing of objects at different distances.

[0157] The depth coordinates of other common feature points of non-target of interest are derived as follows:

[0158] Suppose there are 7 objects P1-P7 with different depth values, wherein P4 is the target of interest, and the distance of other objects from P4 is called depth difference, denoted by ΔD i The imaging point distance of the same feature point on the left and right sensors is called the same feature distance, denoted by Δl i The same feature distance and depth value of P1-P7 objects are shown in Table 1. Figure 5

[0159] ​When the device focuses on object P, the focal points Fl and F2 are locked, and objects with the same distance as the target of interest will project on the same points of the left and right sensors, which is called the "same point imaging law". The same point here refers to the point where the horizontal coordinate value and the vertical coordinate value are equal. That is, in this embodiment, as long as the object has the same point imaging feature, its depth coordinate is equal to that of the focused object. Most objects in the field of view do not have the same distance as the target of interest, and for these objects, the device also has a fixed imaging rule that can be used to calculate their depth coordinates. For example, if the left and right cameras focus on object P4, the depth coordinates of P1, P2, and P3 are smaller than that of the focused object, and the depth coordinates of P5, P6, and P7 are greater than that of the focused object. From this, the following rule can be seen: placing the left and right sensors side by side, it can be found that the distance between the imaging coordinates of the same feature points on the left and right sensors (referred to as the same feature distance) becomes smaller as the object is farther away. It is largest at P1 and smallest at P7. P1 is the position where the left and right fields of view first intersect, and P7 is the position where the left and right fields of view last intersect.

[0160] Therefore, a linear relationship between the same feature distance Δl i and the depth difference ΔD i can be established. Let the relationship between the same feature distance Δl i and the depth difference ΔD i be consistent with the expression:

[0161] ΔD i = αΔl i + β; (IV)

[0162] Substitute the same feature distances and depth differences of objects P1, P4, and P7 into equation (IV) to obtain the following system of equations:

[0163]

[0164] The solution is:

[0165]

[0166] Therefore, the relationship between the same feature distance Δl i and the depth difference ΔD i is:

[0167]

[0168] Combining Figure 6The schematic diagram of the left and right sensors in parallel under the pixel coordinate system. The sensor is composed of n x n imaging units (n is an odd number), and the size of the imaging unit is e x e. The upper left corner of the sensor is the pixel coordinate origin (0, 0), and the coordinate of the center point K of the sensor is ((n+1) / 2, (n+1) / 2). Since the dual-camera device is parallel vision, the imaging result is subject to the epipolar constraint, and the imaging points of any common feature point on the left and right sensors are always on the same scanning line. Let the first imaging point coordinate of any common feature point i on the left sensor be (L i , W i ), and the second imaging point coordinate on the right sensor be (R i , W i ). Where L i and R i are the column numbers of the imaging points in the left and right sensors, respectively, and W i is the row number of the imaging point in the sensor. Then the homologous distance Δl i of the common feature point can be expressed by the pixel coordinate value as:

[0169] Δl i = ne+(R i -L i )e, (Ⅵ)

[0170] Substituting (Ⅵ) into the above (V) can obtain:

[0171]

[0172] Substituting (VII) into D i =D0-ΔD i can obtain:

[0173]

[0174] Finally, we can obtain:

[0175]

[0176] Formula (VIII) represents the depth coordinate calculation formula of any common feature point in the common field of view. Where b, e, v, d are all structure parameters, which are fixed values. (L i -R i )e is the parallax of the common feature point on the left and right sensors. According to the formula, the depth value D i of the common feature point i is determined by the focal length f and the parallax (L i -R i )e. When the spherical camera focuses on the target of interest, the current focal length value f can be obtained through the muscle clues of the ciliary muscle, and the left and right imaging point horizontal coordinates L i of all common feature points can be obtained through the imaging results of the left and right sensors.and R i Then, the depth values of all common feature points in the common field of view are calculated in batch using formula (VIII).

[0177] The derivation of the horizontal coordinate and the vertical coordinate of the common feature point in the camera coordinate system is described. Figure 7 and Figure 8 The derivation of the horizontal coordinate and the vertical coordinate of the common feature point in the camera coordinate system is described.

[0178] The vertical coordinate of the common feature point in the camera coordinate system is derived as follows:

[0179] Because ΔP'P i F~ΔKQ i F, so:

[0180]

[0181] From the above formula, we can get:

[0182]

[0183] Where,

[0184]

[0185] Substitute (X) into (IX), we can get:

[0186]

[0187] When the spatial position of object i is higher than the coordinate origin C(0, 0, 0) of the device, its vertical coordinate W i will be greater than the vertical coordinate value of the center point n+1 / 2. According to formula (XI), when W i >n+1 / 2, V i is less than zero. That is, under formula (XI), the object whose spatial position is higher than the origin has a negative vertical coordinate value. Our cognitive habit is more inclined to be positive up and negative down, so here the sign of formula (XI) is reversed as a whole, that is:

[0188]

[0189] From formula (XI) to formula (XII), only the representation of the spatial vertical coordinate is adjusted from negative up and positive down to positive up and negative down. The relationship between the dependent variable and the independent variable in formula (XI) has not changed, so formula (XII) still holds.

[0190] The horizontal coordinate of the common feature point in the camera coordinate system is derived as follows:

[0191] P i (H i , Vi D i ( ) represents any common feature point within the common field of view. When the device focuses on the target of interest P(0, 0, D0), the locked focus pair is F1 and F2. i The horizontal coordinates projected onto the left and right sensors by the focal point are L and L, respectively. i and R i Find the horizontal coordinate H of Pi at this moment. i .

[0192] Draw the alignment line and P i The intersection point P′(0, 0, D) of the equidistant spheres i Point P', with its focal point projected onto the left and right sensors, has horizontal coordinates L' and R' respectively. i The extension of P′ is P i The equidistant sphere. Q1 and Q2 are the left and right visual axes and P, respectively. i The intersection of equidistant spheres. The straight line connecting Q1 and F2 intersects the right sensor at point Q′. The principal optical axis of the right optical lens intersects P. i The equidistant spheres intersect at point G.

[0193] P i Horizontal distance H from P′ i This is the target we are looking for, and its length is X+Y.

[0194] Because ΔP′P i F2~ΔR′R i F2, ΔP′GF2~ΔR′C2F2, therefore:

[0195]

[0196] From the above formula, we can obtain:

[0197] Calculation (L) i -K1) pixel distance and (R i The sum of pixel distances from -K2) yields:

[0198] [(L i -K1)+(R i -K2)]e=y+(x+x+y)=2(x+y);

[0199] Where the horizontal coordinate of the sensor center point K in the pixel coordinate system is (n+1) / 2, therefore:

[0200] K1+K2=(n+1) / 2+(n+1) / 2=n+1;

[0201] Then the expression for x+y is:

[0202]

[0203] Substituting (i) and (x) into formula (ii), we get:

[0204]

[0205] Since the structural parameters b, e, v, d, and n are known, when the dual-camera device focuses on the target of interest P, the left and right focal points F1 and F2 are locked. The device obtains the focal length f through muscle cues from the ciliary muscle and obtains the coordinates of the first imaging point (L) of all common feature values ​​within the common field of view through the left and right sensors. i W i ) and second imaging point coordinates (R i W i Therefore, the three-dimensional coordinates P of all common feature points in the camera coordinate system can be calculated in batches using the following fixed formula. i (H i V i D i ):

[0206]

[0207] The above formula considers the case where the lateral space of each imaging unit does not overlap. However, in the overall texture map, the horizontal width *e* of one unit contains two texture patches. Therefore, when using texture patches as imaging units, *e* in the above formula needs to be replaced with *e / 2*. Finally, the formula for calculating the 3D coordinates in the camera coordinate system can be obtained as follows:

[0208]

[0209] The process of transforming the three-dimensional coordinates of the common feature points in the camera coordinate system to the three-dimensional coordinates in the world coordinate system is as follows:

[0210] From camera coordinate system P i (H i V i D i Transform to world coordinate system P i (X i Y i Z i This involves rotating the axes of the camera coordinate system to align with the world coordinate system and translating the origin C of the camera coordinate system to the origin of the world coordinate system.

[0211] like Figure 9 As shown, first, rotate the camera coordinate system along the D-axis to correspond to the world coordinate system. For point P in the camera coordinate system... i (H i V i, D i ), which are X-axis coordinate X i and Y-axis coordinate Y i in the world coordinate system respectively:

[0212] X i = H i cos α - V i sin α

[0213] Y i = H i sin α + V i cos α

[0214] Since the three-dimensional coordinates rotate around the D / Z axis, the depth coordinates of P i in the two coordinate systems are equal, so Z i = D i .

[0215] In matrix form:

[0216]

[0217] The 3x3 transformation matrix in it is an axis-based rotation matrix, denoted as r1; in the same way, the rotation conversion of the camera coordinate system along the H axis and the V axis is performed, and the corresponding rotation matrices are denoted as r2 and r3 respectively; and the total rotation matrix R = r1 r2 r3. Therefore, the rotation conversion from the camera coordinate system to the world coordinate system is expressed by the formula:

[0218]

[0219] In addition to the rotation conversion to make the coordinate axes overlap with each other, the camera coordinate system also needs to be translated to make its origin coincide with the origin of the world coordinate system. The translation amounts of the camera coordinate system in the horizontal direction, the vertical direction and the depth direction are denoted as t1, t2 and t3 respectively; then the coordinate system conversion formula is supplemented as:

[0220]

[0221] Denote the total translation matrix Then the above formula can be expressed as an augmented formula:

[0222]

[0223] Wherein, R and T are both attitude parameters and position parameters that can be obtained in real time by the dual-camera device. When the three-dimensional data sets (H i , V i , D i) and then converting it into a three-dimensional data set (X i , Y i , Z i ) in the world coordinate system in real time through the conversion formula.

[0224] In an alternative embodiment, the estimating the three-dimensional coordinates of the non-common feature points according to the three-dimensional coordinates of the common feature points comprises:

[0225] adopting a single-feature-point estimation method to estimate the three-dimensional coordinates of the non-common feature points according to the three-dimensional coordinates of the common feature points;

[0226] alternatively,

[0227] adopting a double-feature-point estimation method to estimate the three-dimensional coordinates of the non-common feature points according to the three-dimensional coordinates of the common feature points.

[0228] In an alternative embodiment, the adopting a single-feature-point estimation method to estimate the three-dimensional coordinates of the non-common feature points according to the three-dimensional coordinates of the common feature points comprises:

[0229] adopting a single-feature-point estimation method to estimate the three-dimensional coordinates of the non-common feature points that fail to be matched successfully by using the double-gradient law of the feature points in the total texture map of the left image and the total texture map of the right image;

[0230] for a non-common feature point, obtaining a common feature point adjacent to the non-common feature point and having a color difference ratio within a set color difference range as a reference point;

[0231] taking the depth coordinate of the reference point as the depth coordinate of the non-common feature point in the camera coordinate system;

[0232] adding 1 or subtracting 1 from the horizontal offset of the reference point and then calculating the horizontal coordinate of the non-common feature point in the camera coordinate system through formula (1); wherein when the non-common feature point is on the left side of the common feature point, the horizontal offset of the reference point is subtracted by 1, and when the non-common feature point is on the right side of the common feature point, the horizontal offset of the reference point is added by 1;

[0233] adding 1 or subtracting 1 from the vertical offset of the reference point and then calculating the vertical coordinate of the non-common feature point in the camera coordinate system through formula (2); wherein when the non-common feature point is above the common feature point, the vertical offset of the reference point is subtracted by 1, and when the non-common feature point is below the common feature point, the vertical offset of the reference point is added by 1.

[0234] In an alternative embodiment, the adopting a double-feature-point estimation method to estimate the three-dimensional coordinates of the non-common feature points according to the three-dimensional coordinates of the common feature points comprises:

[0235] The three-dimensional coordinates of the non-common feature points which are not matched successfully are estimated by using the double feature point estimation method according to the double gradient rules of the feature points in the total texture map of the left image and the total texture map of the right image.

[0236] For the non-common feature points, two common feature points in the left image or the right image which are the same distance but opposite direction from the non-common feature points and have color difference ratio in the set color difference range are obtained as reference points.

[0237] The mean of the depth coordinates of the two reference points is calculated as the depth coordinate of the non-common feature point in the camera coordinate system.

[0238] The mean of the horizontal coordinates of the two reference points is calculated as the horizontal coordinate of the non-common feature point in the camera coordinate system.

[0239] The mean of the vertical coordinates of the two reference points is calculated as the vertical coordinate of the non-common feature point in the camera coordinate system.

[0240] Further, the horizontal coordinate, the vertical coordinate and the depth coordinate of the non-common feature point in the camera coordinate system are transformed according to the pose parameters and the position parameters to obtain the three-dimensional coordinates of the non-common feature point in the world coordinate system. For details, refer to the conversion process of the common feature points, which will not be repeated here.

[0241] The double gradient rules refer to the similarity and gradient of the three-dimensional coordinates and the color difference ratio of the feature points from the same object. The color difference ratio refers to the red, green and blue ratio R:G:B of the texture color block.

[0242] Therefore, the feature points in the field of view space generally have the following two gradient rules:

[0243] (1) Color difference ratio gradient rule: adjacent feature points with similar color difference ratio or color difference ratio gradient are probably from the same object.

[0244] (2) Coordinate gradient rule: the three-dimensional coordinates of the feature points of the same object are close, and the three-dimensional coordinates of adjacent feature points of the same object have gradient.

[0245] Since the total texture map contains all the texture information of the image, and the characteristics of each feature point are intuitively reflected by the color difference value. By using this performance of the total texture map, the three-dimensional coordinates of most feature points in the field of view space which are not matched successfully are estimated.

[0246] In an alternative embodiment, the total texture map of the left image and the total texture map of the right image are extracted, including:

[0247] extracting a total texture map of the left image and a total texture map of the right image by using the multi-layer window grid contralateral opponent extraction method.

[0248] Further, the texture extraction of the left image and the right image by using the multi-layer window grid contralateral opponent extraction method respectively to obtain the total texture map of the left image and the total texture map of the right image comprises:

[0249] respectively converting the color value of each pixel of the left image and the right image into a digital array, wherein the base of the color value of the converted pixel is lower than the base of the color value of the unconverted pixel;

[0250] According to the digital array of each pixel of the left image and the right image, the left image and the right image are respectively decomposed into a plurality of layers;

[0251] respectively extracting texture information of each layer by using a window mask to obtain a sub-texture map of each layer;

[0252] superimposing the weight values of all the sub-texture maps of the left image to synthesize a total texture map of the left image;

[0253] superimposing the weight values of all the sub-texture maps of the right image to synthesize a total texture map of the right image.

[0254] Texture is actually the difference in optical properties between two light surfaces. Texture exists only when there is a color difference between the two light surfaces. Based on this, texture can be detected by detecting the color difference between the two light surfaces. Specifically, an image can be represented in color by using a variety of different color spaces, such as HSV, HSB, CMY, L*a*b*, etc. In the implementation of the present application, the RGB color space is used as an example to present the steps of the process of the present application, but in principle, any other color space can also be used to apply the method implemented by the present application to extract the texture information of the image.

[0255] For left and right images, the specific process of the multi-layer window grid contralateral opponent extraction method (MGCO) for extracting texture information is as follows:

[0256] respectively converting the color value of each pixel of the left image and the right image into a digital array, wherein the base of the color value of the converted pixel is lower than the base of the color value of the unconverted pixel;

[0257] Wherein, the color scale value of each pixel of R, G, B color channel is expressed as N-base form respectively, the RGB value of each pixel is expressed as 3×M-bit digital array; wherein, M represents the digit corresponding to the maximum color value of the pixel in N-base form.

[0258] In the embodiment of the present application, the color value of each pixel of the left image and the right image can be converted from decimal system to N-base system, N<10, to reduce the base of the color value of each pixel of the left image and the right image, such as conversion to binary system, ternary system, quaternary system, etc. The lower the base is, the higher the number of layers of the left image and the right image is decomposed, and the higher the accuracy of texture extraction is.

[0259] Further, the conversion of the color value of each pixel of the left image and the right image to N-base system respectively to obtain the digital array of each pixel of the left image and the right image comprises:

[0260] Converting the color value of each pixel of the left image and the right image to binary system to obtain the 3×8-bit digital array of each pixel;

[0261] In the embodiment of the present application, it is preferred to convert the color value of the pixel in decimal system to binary system. Since the RGB color space is composed of three color channels of red, green and blue, the value of each color channel is an integer value of [0, 255], and there are 256 amplitude values. That is, each color channel can be represented by 8-bit binary. If three color channels are represented by binary, the RGB value of each pixel can be represented as a digital array composed of 3×8=24 0 / 1 signals. Thus, the color value comparison between two light surfaces can be understood as the numerical value comparison between two such digital arrays by computer language. The difference value between two digital arrays is the texture value between two light surfaces.

[0262] The embodiment of the present application utilizes the simple and clear contrast characteristics of binary signal to convert the task of identifying complete texture information into the task suitable for computer language, so that the computer can more simply and accurately extract all the texture information in the left image and the right image.

[0263] For example, for a pixel with color value of RGB(50, 120, 200), the binary digital array corresponding to the color value is shown in the following table.

[0264]

[0265] Wherein:

[0266] R50=128×0+64×0+32×1+16×1+4×0+2×1+1×0;

[0267] G120 = 128 x 0 + 64 x 1 + 32 x 1 + 16 x 1 + 4 x 1 + 2 x 0 + 1 x 0;

[0268] B200 = 128 x 1 + 64 x 1 + 32 x 0 + 16 x 0 + 4 x 1 + 2 x 0 + 1 x 0;

[0269] Further, based on the binary conversion, according to the digital array of each pixel, each digit of each color channel in all pixels is respectively clustered and merged into a layer, and a total of 24 layers are obtained.

[0270] Each digit of each color channel can generate a layer. The color value of each pixel is composed of three color channels R, G and B, and each color channel has 8 digits of 0 / 1. Each digit is named according to its "channel + binary digit". For example, the digit in the R channel at the 6th position is named R6, the digit in the G channel at the 4th position is named G4, and so on. Each pixel contains R1-R8, G1-G8 and B1-B8, a total of 24 digits of 0 / 1. Each digit of each color channel in all pixels is respectively clustered and merged, for example, the R1 digits of all pixels are clustered and merged into a layer, the R2 digits of all pixels are clustered and merged to form a second layer, and so on. Finally, 24 layers are obtained. Each layer is consistent with the number of pixels of the original image, but the pixel value has only two values of 0 and 1.

[0271] Each layer is named by the corresponding same name number constituting the layer, for example, the layer clustered and merged by the G5 digits of each pixel is called the G5 layer, and so on. The 24 layers are R1-R8, G1-G8 and B1-B8, as shown in Figure 10 .

[0272] If the texture information extraction work is performed on the left and right images, each pixel has 256x256x256 = 1677 million color values, and there are 256^6 = 28 trillion different textures between two light surfaces. Therefore, the operation amount of directly extracting all texture information on the left and right images is extremely large. Through layer decomposition of the left and right images, the pixel of each layer has only two values of 0 and 1, and there are only four difference combinations between two pixels, so that the complexity of extracting texture information on the layer is greatly reduced.

[0273] Assigning a value to the sub-texture map: according to the binary digit number reflected in the name of the sub-texture map, a value is assigned to the sub-texture map. The relationship between the weight value y and the binary digit number n is y = 2^(n-1). For example: the weight value y(G5) of the G5 sub-texture map is 2^4 = 16. That is, the weight value of the G5 sub-texture map is the green value 16.

[0274] Weighted superposition of all the sub-texture maps of the left image / right image forms the total texture map of the left image / right image.

[0275] For example, the total texture map of the left image is obtained by weighted superposition of the 24 sub-texture maps corresponding to the 24 layers of the left image, as shown in the following figure. Figure 11 The total texture map of the left image obtained at this time will depict the texture of the left image, and intuitively represent the color difference value of each feature point with an RGB value.

[0276] For example, the principle of texture weighted superposition in the window frame is described. Figure 12 In the example, the RGB value of the left side of the window frame in the figure is 200, 90 and 20, and the RGB value of the right side of the window frame is 210, 220 and 240. Figure 8 After weighted superposition of the 24 sub-texture maps corresponding to the 24 layers according to their weights, the total texture color difference between the left and right sides of the window frame can be obtained as follows.

[0277] R = layer R5 + layer R4 + layer R2 = -16 + 4 - 2 = -10;

[0278] G = layer G8 + layer G3 + layer G2 = -128 - 4 + 2 = -130;

[0279] B = layer B8 + layer B7 + layer B6 + layer B3 = -128 - 64 - 32 + 4 = -220.

[0280] In the implementation of the present application, based on the 8 digits of the R, G and B values in the binary digital array, the left image / right image is divided into 24 layers, and then texture extraction is performed for each layer.

[0281] Further, window frame masks are used for texture information extraction for each layer to obtain the sub-texture map of each layer, including:

[0282] The window frame mask is used to cover each layer, and the window frame mask is composed of a plurality of window frames.

[0283] For each layer, at least two comparison groups are used to extract the texture of the layer region where the window frame is located to obtain the texture information of the layer region where the window frame is located.

[0284] In the example, the points on the boundary of the window frame are taken as the detection points for detecting the pixel value of the layer at the position of the point, and the two detection points opposite to each other in the window frame are taken as a pair of comparison groups.

[0285] Splicing the texture information of all window frames in the window frame mask to obtain a sub-texture map of the corresponding layer.

[0286] Further, a first contrast group and a second contrast group are set at the boundary of each window frame, wherein the included angle between the line connecting the two detection points of the first contrast group and the line connecting the two detection points of the second contrast group is 360° / 2n, n representing the number of contrast groups.

[0287] Then, at least two contrast groups are used for each window frame to extract the texture information of the layer region where the window frame is located, including:

[0288] The pixel values of the sites where the detection points of the first contrast group and the second contrast group are located are extracted respectively and compared numerically.

[0289] When the two detection points of the first contrast group are of the same value and the two detection points of the second contrast group are of the same value, it is determined that there is no texture between the first contrast group and the second contrast group, and the window frame does not need to be filled with color.

[0290] When the two detection points of the first contrast group are of different values and the two detection points of the second contrast group are of the same value, it is determined that there is texture between the first contrast group and no texture in the second contrast group. The window frame is divided into two halves along the median line of the first contrast group, and the half where the detection point with a value of 1 is located is filled with color, and the other half is not filled with color.

[0291] When the two detection points of the first contrast group are of different values and the two detection points of the second contrast group are of different values, it is determined that there is common texture between the first contrast group and the second contrast group. The window frame is divided into two halves along the median line of any two detection points with different values among the four detection points of the first contrast group and the second contrast group, and the half where the two detection points with a value of 1 are located is filled with color, and the half where the two detection points with a value of 0 are located is not filled with color.

[0292] The filling result of the window frame is the texture information of the corresponding window frame.

[0293] In the embodiment of the application, the window frame mask (opposite antagonistic window frame mask) is composed of a plurality of window frames with a self-defined size. For example, in the embodiment of the application, the window frame is set to have a size of 3x3 pixels. For an original target image with a size of 11x11 pixels, each layer is covered by a mask composed of 5x5 window frames. Then, each window frame covers 3x3, i.e. 9 pixels of the image, as shown in the following figure. Figure 3

[0294] ​Structure of window frame: each window frame is provided with at least two contrast groups, and the feature points of the contrast groups are: 1) one contrast group is composed of two detection points located on the boundary of the window frame, and the two detection points are symmetrical along the midpoint and located on opposite sides of the window frame; 2) the included angle between the directions of the two contrast groups is 360° / 2n, and n represents the number of contrast groups. Taking the setting of the first contrast group A and the second contrast group B as an example, the extraction of the texture information of the window frame is described:

[0295] Texture identification: the values of the pixel points where the detection points of the contrast groups A and B are located are extracted respectively, and the values are compared; if the two detection points of the contrast groups A and B are the same (both 0 or both 1), it is determined that there is no texture between the contrast groups A and B; if the two detection points of the contrast groups A or B are different (one is 0 and the other is 1), it is determined that there is texture between the contrast groups A or B;

[0296] Texture extraction: the pixel region covered by the four detection points of the contrast groups A and B is extracted. According to the different detection results of the contrast groups A and B, the following texture result output is performed respectively:

[0297] 1) if the two detection points of the contrast group A are the same and the detection points of the contrast group B are the same: the window frame has no texture and is not filled with color

[0298] 2) if the detection points of the contrast group A are different and the detection points of the contrast group B are the same: the window frame is divided into two halves along the median line of the contrast group A, and the half of the window frame where the detection point with a value of 1 of the contrast group A is located is filled with color, and the other half is not filled with color;

[0299] 3) if the detection points of the contrast group A are the same and the detection points of the contrast group B are different: the window frame is divided into two halves along the median line of the contrast group B, and the half of the window frame where the detection point with a value of 1 of the contrast group B is located is filled with color, and the other half is not filled with color;

[0300] 4) if the detection points of the contrast group A are different and the detection points of the contrast group B are different: the window frame is divided into two halves along the median line of any two detection points with different values among the four detection points, and the half of the window frame where the two detection points with a value of 1 are located is filled with color, and the other half is not filled with color;

[0301] For example, it is assumed that the contrast group A is horizontally oriented and the contrast group B is vertically oriented.

[0302] When the detection points of the contrast group A are the same and the detection points of the contrast group B are also the same, it is determined that there is no texture in the window frame;

[0303] When the detection points of the A comparison group are different values and the detection points of the B comparison group are the same value, the window frame is divided into two halves by the median line of the A comparison group. At this time, if the left detection point of the A comparison group is 1 and the right detection point is 0, the left half of the window frame is filled with color and the right half is maintained without color; if the 1 value is in the right detection point, the right half of the window frame is filled with color and the left half is maintained without color.

[0304] When the detection points of the A comparison group are the same value and the detection points of the B comparison group are different values, the window frame is divided into two halves by the median line of the B comparison group. At this time, if the upper detection point of the B group is 1 and the lower detection point is 0, the upper half of the window frame is filled with color and the lower half is maintained without color; otherwise, the lower half is filled with color and the upper half is maintained without color.

[0305] When the detection points of the A comparison group are different values and the detection points of the B comparison group are also different values, the window frame is divided into two halves by the median line of the two detection points that are different values, and the half in which the two 1 value detection points are located is filled with color. At this time, if the left detection point and the upper detection point are 1 values, the upper left half of the window frame is filled with color and the lower right half is maintained without color; if the right detection point and the lower detection point are 1 values, the lower right half of the window frame is filled with color and the lower left half is maintained without color.

[0306] The four detection points are the midpoints of the four edges of the window frame, wherein a pair of detection points opposite each other is taken as a comparison group, and the four detection points of the window frame can be divided into a pair of detection points in the vertical direction and a pair of detection points in the horizontal direction. The values of the pixels where the detection points are located are compared, and if there is a different value in any pair of detection points in the comparison, it indicates that there is texture between the two detection points, i.e., there is texture in the window frame; if there is no different value in both pairs of detection points in the comparison, there is no texture in the window frame. For example, the values of the pixels where a pair of detection points in the vertical direction of a window frame on layer B1 are 1 and 0, and the values of the pixels where a pair of detection points in the horizontal direction are 0 and 0, indicating that there is texture in the window frame; for another example, the values of the pixels where a pair of detection points in the vertical direction of another window frame on layer B1 are 0 and 0, and the values of the pixels where a pair of detection points in the horizontal direction are 0 and 0, indicating that there is no texture in the window frame. Specifically, the window frame can be divided into two halves by the median line of a pair of detection points with different pixel values, and the half where the pixel value is 1 is filled with color and the half where the pixel value is 0 is not filled with color; wherein the window frame is filled with color according to the color weight of the corresponding layer, and the color weight depends on the color channel corresponding to the layer and the weight corresponding to its digit, and the relationship between the weight y and the digit n where it is located is y = 2^(n-1). For example, the window frame of layer R1 is filled with color according to R = 1, the window frame of layer G5 is filled with color according to G = 16, and the window frame of layer B8 is filled with color according to B = 128.

[0307] As shown in FIG. 1, the texture direction of each window frame in the target image is determined by comparing the pixel values of the detection points in the window frame. For example, taking the layer R1 as an example, one horizontal direction and one vertical direction contrast group are set for each window frame in the layer R1. When only one pair of detection points is different, i.e., one is 1 and the other is 0, the middle line of the pair of different detection points is taken as the texture direction; when two pairs of detection points are different, the texture direction is the diagonal line of the window frame. There are 16 combinations of the two pairs of detection points. When the two pairs of contrast groups are the same, i.e., both are 0 or both are 1, it represents that there is no texture in the window frame. The texture direction in the window frame is as shown in FIG. 1. By setting multiple pairs of detection points to compare the pixel values, the accuracy of the texture direction in the window frame can be improved. Figure 13

[0308] The texture information of all the window frames in the window frame mask is spliced to obtain a sub-texture map of the corresponding layer.

[0309] After the color filling processing of the region in each window frame is completed, the texture information of all the window frames is merged according to the position of the window frame in the window frame mask to form a sub-texture map of the layer.

[0310] The sub-texture map is named according to the layer in which it is located, for example, the sub-texture map of the R1 layer is called R1 sub-texture map, and so on. As shown in FIG. 1. Figure 15

[0311] In the embodiment of the present application, the total texture map of the target image is the result of superimposing the texture directions of the window frames in 24 layers. Since there are 8 possibilities for the texture direction of the window frame in each layer, and the texture direction of the same window frame in different layers may be different, after superimposing multiple layers, the texture of the window frame in the total texture map is often in the form of a rice character. Assuming that the area of the window frame is S, and the minimum color block unit of the total texture map is a right-angled triangle with an area of S / 8, therefore, by taking the window frame as a unit to extract the texture of the target image, the resolution of the total texture map will be 8 times higher than that of the window frame mask. When the window frame is set to be 3x3 pixels in size, the area of each pixel is S / 9, and the minimum resolvable color block area of the total texture map is S / 8, that is, it can be inferred that the minimum color block area of the total texture map is 1.125 times that of a single pixel. As can be seen, when the window frame is set to be 3x3 pixels in size, the total texture map obtained finally is close to the level of single pixel in detail resolution, and has a similar detail resolution to the target image, so that the extracted texture information has good integrity and fineness.

[0312] The traditional edge detection algorithm is good for edge detection of gray images with obvious texture, but the edge detection effect is poor for color images or natural environment images with weak or complex texture. Compared with the prior art, the present application has excellent texture extraction performance for various image types, including abstract graphs, scatter plots, black and white images, color images, human images, environment images, etc.

[0313] ​​Since the total texture map extracted by the present application not only retains all the texture information, but also intuitively reflects the color difference characteristics of each texture point with RGB values. Therefore, the textures of different objects in the target image can be separated by using the rule that the color difference characteristics of the textures of the same object are similar (need to be illustrated). There is also great application prospect in subsequent image processing technologies such as object separation and object recognition.

[0314] In an alternative embodiment, the method further comprises:

[0315] By setting the RGB value range, the total texture map is filtered to obtain the final texture map output result.

[0316] The present method can flexibly filter the texture information of the total texture map by setting the RGB value range to obtain texture results of different degrees of simplification.

[0317] The present method can flexibly filter the texture information of the total texture map by setting the RGB value range to obtain texture results of different degrees of simplification.

[0318] In the embodiment of the present application, the total texture map contains all the texture information of the target image, and the color difference characteristics of each texture point are displayed with intuitive RGB values, so that the texture data contained in the total texture map has a great mining space and great ease of operation. For example, by setting different numerical range of RGB value filtering conditions, the total texture map is filtered by the RGB value filtering conditions to obtain different texture results. For example, set R, G, B ∈ (8, 255] RGB value filtering condition, that is, filter out all the texture blocks with RGB values less than 8; or further improve the RGB value filtering condition, set R, G, B ∈ (16, 255] RGB value filtering condition, filter out the texture blocks with RGB values less than 16, and further simplify the texture. By setting different RGB value filtering conditions, the color blocks in the total texture map that do not meet the RGB value filtering conditions are directly deleted, and the remaining color blocks constitute a result image with clear texture, so that the texture result has the characteristics of self-definition.

[0319] At the same time, since the total texture map of the embodiment of the present application includes all the texture information of color difference values 1-255, it can detect textures that are easily ignored / difficult to identify by the naked eye.

[0320] In other embodiments, the total texture map can also be filtered and processed by using a regional ranking filtering method, which specifically includes:

[0321] The overall texture map is divided into i columns, j rows, and ixj sub-regions. Color patches are filtered independently within each sub-region. The filtering criteria are set to delete all color patches within the sub-region whose color difference value ranks below n%, and to filter out extremely weak textures with RGB values ​​less than m. Here, n and m are preset constants.

[0322] By filtering the overall texture map using a regional ranking method, it can be ensured that each sub-region retains a certain amount of texture, so that the relatively strong textures in the weak texture area can be preserved. These relatively strong textures are often important textures that are difficult to extract due to interference from lighting and shadow environment, thus avoiding interference from ambient lighting on texture extraction.

[0323] The advantages of using the total texture map for feature point matching are:

[0324] (1) Filter out secondary feature points.

[0325] Feature points of an object are key locations for locating its 3D coordinates. Using the overall texture map instead of the original image for feature matching is equivalent to filtering out low-value smooth surface signal points before matching, retaining only the key information related to 3D localization. This eliminates the need to calculate the 3D coordinates of low-value signal points, thus improving the efficiency of 3D coordinate calculation.

[0326] (2) Higher matching precision.

[0327] The smallest feature unit of the overall texture map is a right-angled triangular color patch. Although its area is 1.123 times that of a single pixel, from Figure 14 The comparison shows that within a 3x3 grid area of ​​the same size, each row, which originally had 3 pixels, becomes 4 texture blocks. Because the left and right sensors are constrained by epipolar lines, the same feature point will only be located in the same row. Therefore, after converting the original image to a total texture map, the number of rows of feature points is reduced to 2 / 3 of the original, but the number of columns increases to 4 / 3. Although the total number of feature points that need to be matched is less, the matching precision is actually 1 / 3 higher than the original image. Therefore, using the total texture map for feature point matching does not result in a coarser result than matching using the original image.

[0328] (3) The error rate of feature matching is lower.

[0329] First, the original image uses color values ​​as its feature elements, while the overall texture map uses color difference values ​​as its feature elements. It's highly probable that different locations of the same object have the same color, but the probability of them having the same color difference is much lower. Therefore, the possibility of mismatched texture color patches is far lower than that of mismatched pixels in the original image.

[0330] Secondly, the texture blocks not only have color difference value difference, but also have shape difference. The four blocks in the same row in the same grid have different shapes. The pixels of the original image can be mismatched as long as the colors are consistent, while the blocks of the total texture map need to be consistent in color difference value and shape to be mismatched, which makes the mismatch of the texture blocks become an extremely low probability event. In summary, the error rate will be greatly reduced by using the total texture map for feature matching.

[0331] Referring to Figure 17 The second embodiment of the present application provides a three-dimensional reconstruction device based on bionic stereovision, comprising:

[0332] An image acquisition module 1 is configured to acquire a left image and a right image captured by a double-camera device when focusing on an interest target;

[0333] A texture extraction module 2 is configured to extract textures from the left image and the right image respectively to obtain a total texture map of the left image and a total texture map of the right image;

[0334] A feature matching module 3 is configured to perform feature matching on the total texture map of the left image and the total texture map of the right image to generate a dense disparity map;

[0335] A three-dimensional coordinate calculation module 4 is configured to calculate three-dimensional coordinates of common feature points matched in the dense disparity map, and estimate three-dimensional coordinates of non-common feature points according to the three-dimensional coordinates of the common feature points to obtain a three-dimensional coordinate dataset;

[0336] A three-dimensional reconstruction module 5 is configured to perform three-dimensional reconstruction according to the three-dimensional coordinate dataset to generate a stereoscopic view of the interest target.

[0337] In an optional embodiment, the double-camera device comprises a left camera and a right camera, the left camera is provided with a left sensor, and the right camera is provided with a right sensor; wherein the photosensitive wafers of the left sensor and the right sensor are both spherical structures, and the inner side of the sphere is the signal receiving surface; the center point of the left sensor is located on the left side of the optical lens main optical axis of the left camera, and the center point of the right sensor is located on the right side of the optical lens main optical axis of the right camera, so that the optical lens main optical axis of the left camera and the left visual axis, and the optical lens main optical axis of the right camera and the right visual axis form a kappa angle respectively, and the center point section of the left sensor and the center point section of the right sensor always remain parallel during the shooting process.

[0338] In an optional embodiment, the image acquisition module 1 comprises:

[0339] The camera coordinate system setting unit is configured to set a camera coordinate system, wherein a straight line passing through the center points of the left sensor and the right sensor is set as a horizontal coordinate axis of the camera coordinate system; a midpoint of a line connecting the center points of the left sensor and the right sensor is set as an origin of the camera coordinate system; a straight line passing through the origin of the camera coordinate system, the lens optical center of the left camera and the lens optical center of the right camera is set as a depth coordinate axis of the camera coordinate system; and a straight line passing through the origin of the camera coordinate system and perpendicular to the horizontal coordinate axis and the depth coordinate axis is set as a vertical coordinate axis of the camera coordinate system

[0340] The focus control unit is configured to rotate the dual-camera device to lock the selected target of interest on the depth coordinate axis of the camera coordinate system.

[0341] The focal length adjustment unit is configured to respectively adjust the focal lengths of the left camera and the right camera, so that the target of interest is respectively projected on the center points of the left sensor and the right sensor, and the focal lengths are locked to focus on the target of interest.

[0342] The image capturing unit is configured to obtain a left image captured by the left camera and a right image captured by the right camera after focusing.

[0343] In an optional embodiment, the texture extraction module 2 is specifically configured to extract the total texture map of the left image and the total texture map of the right image by using a multi-layer window frame grid side antagonistic extraction method.

[0344] In an optional embodiment, the texture extraction module 2 comprises:

[0345] The base conversion unit is configured to respectively convert the color values of each pixel of the left image and the right image into a digital array, wherein the base of the color value of the converted pixel is lower than the base of the color value of the unconverted pixel.

[0346] The layer decomposition unit is configured to decompose the left image and the right image into a plurality of layers according to the digital array of each pixel of the left image and the right image.

[0347] The sub-texture extraction unit is configured to extract texture information of each layer by using a window frame mask, to obtain a sub-texture map of each layer.

[0348] The first texture merging unit is configured to superimpose the weight values of all the sub-texture maps of the left image to synthesize a total texture map of the left image.

[0349] A second texture merging unit is configured to superimpose all the sub-texture maps of the right image with weights to synthesize a total texture map of the right image.

[0350] In an alternative embodiment, the feature matching module 3 comprises:

[0351] A texture matching unit is configured to perform feature matching on texture patches of the total texture map of the left image and the total texture map of the right image as feature points to obtain a dense disparity map; wherein the texture patches are in the shape of right-angled triangles, and the orientations of four adjacent texture patches in the same row are different.

[0352] Further, for each total texture map, a region with a pixel size of N*N is divided into multiple texture patches in the shape of right-angled triangles.

[0353] In an alternative embodiment, the three-dimensional coordinate calculation module 4 comprises:

[0354] A first three-dimensional coordinate calculation unit is configured to calculate a three-dimensional coordinate of the common feature point in a camera coordinate system according to the structural parameters of the dual-camera device and the left-view pixel coordinate of the common feature point on the left sensor and the right-view pixel coordinate of the common feature point on the right sensor; wherein the structural parameters comprise a first distance between an optical lens main optical axis of the left camera and an optical lens main optical axis of the right camera, a second distance between a center point of the left sensor and the optical lens main optical axis of the left camera, a lens distance, a focal length, a number of texture patches per row of the left sensor, and a horizontal width of the texture patches.

[0355] A second three-dimensional coordinate calculation unit is configured to transform the three-dimensional coordinate of the common feature point in the camera coordinate system into a three-dimensional coordinate in a world coordinate system according to the position parameters and the attitude parameters of the dual-camera device; wherein the attitude parameters comprise a horizontal coordinate axis deflection amount, a vertical coordinate axis deflection amount, and a depth coordinate axis deflection amount of the camera coordinate system of the dual-camera device relative to the world coordinate system; and the position parameters comprise a horizontal distance, a vertical distance, and a depth distance of the origin of the camera coordinate system of the dual-camera device relative to the origin of the world coordinate system.

[0356] In an alternative embodiment, the first three-dimensional coordinate calculation unit comprises:

[0357] A parameter calculation unit is configured to calculate a disparity, a horizontal offset amount, and a vertical offset amount of the common feature point according to the left-view pixel coordinate of the common feature point on the left sensor and the right-view pixel coordinate of the common feature point on the right sensor.

[0358] a horizontal coordinate calculation unit configured to calculate a horizontal coordinate of the common feature point in a camera coordinate system according to the first distance, the second distance, the number of texture patches, the horizontal width, and a disparity and a horizontal offset of the common feature point;

[0359] Specifically, the horizontal coordinate H of the common feature point is calculated according to formula (1) i ;

[0360]

[0361] wherein e represents the horizontal width of the texture patch, n represents the number of texture patches per row of the left / right sensor, b represents the first distance between the optical lens principal axis of the left camera and the optical lens principal axis of the right camera, d represents the second distance from the center point of the left sensor to the optical lens principal axis of the left camera, (L i , W i ) represents the left view pixel coordinate of the common feature point i on the left sensor, (R i , W i ) represents the right view pixel coordinate of the common feature point i on the right sensor, (L i -R i e represents the disparity of the common feature point i, -L i -R i represents the horizontal offset of the common feature point i.

[0362] a vertical coordinate calculation unit configured to calculate a vertical coordinate of the common feature point in the camera coordinate system according to the first distance, the second distance, the number of texture patches, the horizontal width, and a disparity and a vertical offset of the common feature point;

[0363] Specifically, the vertical coordinate V of the common feature point is calculated according to formula (2) i ;

[0364]

[0365] wherein W i represents the vertical offset of the common feature point i.

[0366] a depth coordinate calculation unit configured to calculate a depth coordinate of the common feature point in the camera coordinate system according to the first distance, the second distance, the focal length, the lens distance, the horizontal width, and the disparity of the common feature point;

[0367] Specifically, the depth coordinate D of the common feature point is calculated according to formula (3) i ;

[0368]

[0369] wherein v represents the mirror distance, and f represents the focal length.

[0370] The second three-dimensional coordinate calculation unit comprises:

[0371] The coordinate conversion unit is configured to perform coordinate system conversion on the horizontal coordinate, the vertical coordinate and the depth coordinate of the common feature point in the camera coordinate system according to the attitude parameter and the position parameter, to obtain the three-dimensional coordinate of the common feature point in the world coordinate system.

[0372] In an alternative embodiment, the root three-dimensional coordinate calculation module 4 comprises:

[0373] The single-feature-point estimation unit is configured to estimate the three-dimensional coordinate of the non-common feature point by using a single-feature-point estimation method according to the three-dimensional coordinate of the common feature point.

[0374] Alternatively,

[0375] The double-feature-point estimation unit is configured to estimate the three-dimensional coordinate of the non-common feature point by using a double-feature-point estimation method according to the three-dimensional coordinate of the common feature point.

[0376] In an alternative embodiment, the single-feature-point estimation unit is configured to:

[0377] estimate the three-dimensional coordinate of the non-common feature point by using a single-feature-point estimation method according to the feature point double-gradation rule in the total texture map of the left image and the total texture map of the right image;

[0378] For the non-common feature point, a common feature point adjacent to the non-common feature point and having a color difference ratio within a set color difference range is obtained as a reference point;

[0379] The depth coordinate of the reference point is taken as the depth coordinate of the non-common feature point in the camera coordinate system;

[0380] After adding 1 or subtracting 1 from the horizontal offset of the reference point, the horizontal coordinate of the non-common feature point in the camera coordinate system is calculated by using formula (1); wherein, when the non-common feature point is on the left side of the common feature point, the horizontal offset of the reference point is subtracted by 1, and when the non-common feature point is on the right side of the common feature point, the horizontal offset of the reference point is added by 1;

[0381] After adding 1 or subtracting 1 from the vertical offset of the reference point, the vertical coordinate of the non-common feature point in the camera coordinate system is calculated by using formula (2); wherein, when the non-common feature point is above the common feature point, the vertical offset of the reference point is subtracted by 1, and when the non-common feature point is below the common feature point, the vertical offset of the reference point is added by 1.

[0382] In an alternative embodiment, the double feature point estimation unit is configured to:

[0383] using the double feature point estimation method, estimate the three-dimensional coordinates of the non-common feature points that are not successfully matched, according to the double gradient rules of the feature points in the total texture map of the left image and the total texture map of the right image;

[0384] For the non-common feature points, obtain two common feature points in the left image or the right image that are the same distance but opposite directions from the non-common feature points and have color difference ratios within a set color difference range, as reference points;

[0385] calculate the mean of the depth coordinates of the two reference points as the depth coordinate of the non-common feature point in the camera coordinate system;

[0386] calculate the mean of the horizontal coordinates of the two reference points as the horizontal coordinate of the non-common feature point in the camera coordinate system;

[0387] calculate the mean of the vertical coordinates of the two reference points as the vertical coordinate of the non-common feature point in the camera coordinate system.

[0388] It should be noted that the principle and technical effects of the three-dimensional reconstruction device based on biomimetic stereo vision described in the embodiments of the present application are the same as those of the three-dimensional reconstruction method based on biomimetic stereo vision described in the first embodiment, and will not be repeated here.

[0389] Referring to Figure 18 The fourth embodiment of the present application provides a three-dimensional reconstruction device based on biomimetic stereo vision, which comprises at least one processor 11, such as a CPU, at least one network interface 14 or other user interface 13, a memory 15, and at least one communication bus 12 for connecting and communicating between these components. The user interface 13 can optionally include a USB interface and other standard interfaces, wired interfaces. The network interface 14 can optionally include a Wi-Fi interface and other wireless interfaces. The memory 15 can include a high-speed RAM memory and can also include a non-volatile memory, such as at least one disk memory. The memory 15 can optionally include at least one storage device located remotely from the aforementioned processor 11.

[0390] In some embodiments, the memory 15 stores the following elements, executable modules or data structures, or a subset thereof, or an extended set thereof:

[0391] The operating system 151 contains various system programs for implementing various basic services and processing hardware-based tasks;

[0392] program 152 stored in the memory 15, and executes the method of three-dimensional reconstruction based on biomimetic stereo vision described in the above embodiments, for example

[0393] Specifically, the processor 11 is configured to invoke the program 152 stored in the memory 15, and execute the method of three-dimensional reconstruction based on biomimetic stereo vision described in the above embodiments, for example Figure 1 The processor 11 is configured to invoke the program 152 stored in the memory 15, and execute the method of three-dimensional reconstruction based on biomimetic stereo vision described in the above embodiments, for example

[0394] For example, the computer program can be divided into one or more modules / units, which are stored in the memory and executed by the processor to complete the present application. The one or more modules / units can be a series of computer program instruction segments capable of completing a specific function, which are used to describe the execution process of the computer program in the three-dimensional reconstruction device based on biomimetic stereo vision.

[0395] The three-dimensional reconstruction device based on biomimetic stereo vision can be a VCU, an ECU, a BMS, or other computing devices. The three-dimensional reconstruction device based on biomimetic stereo vision can include, but is not limited to, a processor and a memory. Those skilled in the art can understand that the schematic diagram is only an example of the three-dimensional reconstruction device based on biomimetic stereo vision, and does not constitute a limitation on the three-dimensional reconstruction device based on biomimetic stereo vision, which can include more or fewer components than the diagram, or combine certain components, or different components.

[0396] The processor 11 can be a microcontroller unit (MCU), a central processing unit (CPU), and can also be other general-purpose processors, digital signal processors (DSP), application specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or any conventional processor, etc. The processor 11 is the control center of the three-dimensional reconstruction device based on biomimetic stereo vision, which connects all parts of the three-dimensional reconstruction device based on biomimetic stereo vision through various interfaces and lines.

[0397] The memory 15 can be used to store the computer programs and / or modules, and the processor 11 realizes various functions of the three-dimensional reconstruction device based on biomimetic stereo vision by running or executing the computer programs and / or modules stored in the memory and calling the data stored in the memory. The memory 15 can mainly include a program storage area and a data storage area, wherein the program storage area can store an operating system, at least one application program required by a function (such as a sound playing function, an image playing function, etc.), and the like; and the data storage area can store data created according to the use of the mobile phone (such as audio data, a phone book, etc.), and the like. In addition, the memory 15 can include a high-speed random access memory, and can also include a non-volatile memory, for example, a hard disk, a memory, a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, at least one disk storage device, a flash memory device, or other volatile solid-state storage devices.

[0398] The modules / units of the three-dimensional reconstruction device based on biomimetic stereo vision are integrated in the form of software function units and sold or used as independent products, which can be stored in a computer readable storage medium. Based on such understanding, all or part of the processes in the above-mentioned embodiment methods can also be completed by a computer program instructing related hardware, and the computer program can be stored in a computer readable storage medium. When the processor executes the computer program, the steps of the above-mentioned various method embodiments can be realized. The computer program includes computer program code, which can be in the form of source code, object code, executable files or some intermediate forms, etc. The computer readable medium can include any entity or device capable of carrying the computer program code, recording medium, U disk, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal and software distribution medium, etc.

[0399] In a fourth aspect, an embodiment of the present application provides a computer readable storage medium, which includes a stored computer program, wherein the computer readable storage medium controls a device where the computer readable storage medium is located to perform the three-dimensional reconstruction method based on biomimetic stereo vision according to any one of the first aspect when the computer program runs.

[0400] It should be noted that the apparatus embodiments described above are merely illustrative, and the units described as separate units can or can not be physically separate, and the units shown as units can or can not be physical units, i.e., can be located in one place, or can be distributed to multiple network units. Part or all of the modules can be selected to achieve the purpose of the embodiment according to actual needs. In addition, the connection relationship between the modules in the apparatus embodiment provided by the present application indicates that there is a communication connection between them, which can be implemented as one or more communication buses or signal lines. Those skilled in the art can understand and implement it without creative labor.

[0401] The above describes the preferred embodiments of the present application. It should be noted that those skilled in the art can make several improvements and refinements without departing from the principles of the present application, and these improvements and refinements are also considered within the scope of protection of the present application.

Claims

1. A three-dimensional reconstruction method based on biomimetic stereo vision, characterized in that, include: The method acquires left and right images captured by a dual-camera device when it focuses on a target of interest. The dual-camera device includes a left camera and a right camera. The left camera is equipped with a left sensor, and the right camera is equipped with a right sensor. The photosensitive wafers of the left and right sensors are both spherical structures, with the inner side of the sphere serving as the signal receiving surface. The left and right sensors are respectively translated as a whole towards the temporal side of the camera's optical axis. The textures of the left and right images are extracted using a multi-layer window frame side-antagonistic extraction method to obtain the total texture map of the left image and the total texture map of the right image. Feature matching is performed on the total texture map of the left image and the total texture map of the right image to generate a dense disparity map; A camera coordinate system is established with the midpoint of the line connecting the center points of the left sensor and the right sensor as the origin. Using the fixed three-dimensional coordinate calculation formula in the camera coordinate system corresponding to the dual-camera device, the three-dimensional coordinates of the common feature points matched in the dense parallax map are calculated, and the three-dimensional coordinates of the non-common feature points are estimated based on the three-dimensional coordinates of the common feature points to obtain a three-dimensional coordinate dataset. Based on the three-dimensional coordinate dataset, a three-dimensional reconstruction is performed to generate a stereoscopic view of the target of interest.

2. The 3D reconstruction method based on biomimetic stereo vision as described in claim 1, characterized in that, The center point of the left sensor is located to the left of the main optical axis of the left camera's optical lens, and the center point of the right sensor is located to the right of the main optical axis of the right camera's optical lens, so that the main optical axis of the left camera's optical lens and the left viewing axis, and the main optical axis of the right camera's optical lens and the right viewing axis, respectively form a kappa angle. During the shooting process, the cross-section of the center point of the left sensor and the cross-section of the center point of the right sensor always remain parallel.

3. The 3D reconstruction method based on biomimetic stereo vision as described in claim 2, characterized in that, Acquire the left and right images captured by the dual-camera device when it is focused on the target of interest, including: The camera coordinate system is defined as follows: the horizontal axis of the camera coordinate system is the straight line connecting the center points of the left and right sensors; the origin of the camera coordinate system is the midpoint of the line connecting the center points of the left and right sensors; the depth axis of the camera coordinate system is the straight line connecting the origin of the camera coordinate system, the midpoint of the line connecting the optical center of the lens of the left camera and the optical center of the lens of the right camera; and the vertical axis of the camera coordinate system is the straight line passing through the origin of the camera coordinate system and perpendicular to both the horizontal and depth axes. Rotate the dual-camera device so that the depth coordinate axis of the camera coordinate system is aligned with the target of interest; The focal lengths of the left and right cameras are adjusted so that the target of interest is projected onto the center points of the left and right sensors, respectively, and the focal lengths are locked to focus on the target of interest. The left image captured by the left camera and the right image captured by the right camera are obtained after focusing.

4. The 3D reconstruction method based on biomimetic stereo vision as described in claim 1, characterized in that, The method employs a multi-layer window frame side-antagonistic extraction technique to extract textures from the left and right images respectively, obtaining the total texture map of the left image and the total texture map of the right image, including: The color values ​​of each pixel in the left and right images are converted into a base to obtain a number array of each pixel in the left and right images; wherein the base of the color values ​​of the converted pixels is lower than the base of the color values ​​of the pixels before conversion. Based on the numerical array of each pixel in the left and right images, each of the left and right images is decomposed into several layers; Texture information is extracted for each layer using a window frame mask to obtain a sub-texture map for each layer; The weights of all the sub-texture maps of the left image are superimposed to synthesize the total texture map of the left image; The weights of all the sub-texture maps of the right image are superimposed to synthesize the total texture map of the right image.

5. The 3D reconstruction method based on biomimetic stereo vision as described in claim 3, characterized in that, The step of performing feature matching on the total texture map of the left image and the total texture map of the right image to generate a dense disparity map includes: Feature matching is performed using the texture color patches of the total texture map of the left image and the total texture map of the right image as feature points to obtain a dense disparity map; wherein the texture color patches are right-angled triangles, and the four adjacent texture color patches in the same row have different orientations.

6. The 3D reconstruction method based on biomimetic stereo vision as described in claim 5, characterized in that, The method further includes: For each overall texture map, the N*N pixel area is divided into a cross shape to obtain multiple texture color blocks in the form of right-angled triangles.

7. The 3D reconstruction method based on biomimetic stereo vision as described in claim 5, characterized in that, The calculation of the three-dimensional coordinates of the common feature points matched in the dense disparity map includes: Based on the structural parameters of the dual-camera device and the left-view pixel coordinates of the common feature point on the left sensor and the right-view pixel coordinates of the common feature point on the right sensor, the three-dimensional coordinates of the common feature point in the camera coordinate system are calculated; wherein, the structural parameters include a first distance between the principal optical axis of the optical lens of the left camera and the principal optical axis of the optical lens of the right camera, a second distance between the center point of the left sensor and the principal optical axis of the optical lens of the right camera, lens distance, focal length, number of texture color blocks per row of the left sensor, and horizontal width of the texture color blocks; The three-dimensional coordinates of the common feature points in the camera coordinate system are transformed into three-dimensional coordinates in the world coordinate system based on the position and attitude parameters of the dual-camera device. The attitude parameters include the horizontal axis deflection, vertical axis deflection, and depth axis deflection of the camera coordinate system of the dual-camera device relative to the world coordinate system. The position parameters include the horizontal distance, vertical distance, and depth distance of the origin of the camera coordinate system of the dual-camera device relative to the origin of the world coordinate system.

8. The 3D reconstruction method based on biomimetic stereo vision as described in claim 7, characterized in that, The step of calculating the three-dimensional coordinates of the common feature point in the camera coordinate system based on the structural parameters of the dual-camera device and the left-view pixel coordinates of the common feature point on the left sensor and the right-view pixel coordinates of the common feature point on the right sensor includes: Based on the left view pixel coordinates of the common feature point on the left sensor and the right view pixel coordinates of the common feature point on the right sensor, calculate the disparity, horizontal offset, and vertical offset of the common feature point; Based on the first distance, the second distance, the number of texture color blocks, the horizontal width, and the parallax and horizontal offset of the common feature points, calculate the horizontal coordinates of the common feature points in the camera coordinate system; Based on the first distance, the second distance, the number of texture color blocks, the horizontal width, and the parallax and vertical offset of the common feature points, calculate the vertical coordinates of the common feature points in the camera coordinate system; Calculate the depth coordinates of the common feature point in the camera coordinate system based on the first distance, the second distance, the focal length, the lens distance, the horizontal width, and the parallax of the common feature point; Then, the step of transforming the three-dimensional coordinates of the common feature points in the camera coordinate system to the three-dimensional coordinates in the world coordinate system based on the position and attitude parameters of the dual-camera device includes: Based on the posture parameters and position parameters, the horizontal, vertical, and depth coordinates of the common feature points in the camera coordinate system are transformed to obtain the three-dimensional coordinates of the common feature points in the world coordinate system.

9. The 3D reconstruction method based on biomimetic stereo vision as described in claim 8, characterized in that, The step of calculating the horizontal coordinates of the common feature points in the camera coordinate system based on the first distance, the second distance, the number of texture color blocks, the horizontal width, and the disparity and horizontal offset of the common feature points includes: Calculate the horizontal coordinate H of the common feature point according to formula (1). i ; Where e represents the horizontal width of the texture patch, n represents the number of texture patches per row of the left / right sensors, b represents the first distance between the principal optical axis of the left camera's lens and the principal optical axis of the right camera's lens, and d represents the second distance between the center point of the left sensor and the principal optical axis of the left camera's lens, and between the center point of the right sensor and the principal optical axis of the right camera's lens. (L) i W i ) represents the left view pixel coordinates of the common feature point i on the left sensor, (R i W i ) represents the right view pixel coordinates of the common feature point i on the right sensor, (L i -R i )e represents the disparity of the common feature point i, -L i -R i This represents the horizontal offset of the common feature point i.

10. The 3D reconstruction method based on biomimetic stereo vision as described in claim 9, characterized in that, The step of calculating the vertical coordinates of the common feature points in the camera coordinate system based on the first distance, the second distance, the number of texture color blocks, the horizontal width, and the disparity and vertical offset of the common feature points includes: Calculate the vertical coordinate V of the common feature point according to formula (2). i ; Among them, W i This represents the vertical offset of the common feature point i.

11. The 3D reconstruction method based on biomimetic stereo vision as described in claim 10, characterized in that, The step of calculating the depth coordinates of the common feature point in the camera coordinate system based on the first distance, the second distance, the focal length, the lens distance, the horizontal width, and the disparity of the common feature point includes: Calculate the depth coordinates D of the common feature points according to formula (3). i ; Where v represents the lens distance and f represents the focal length.

12. The 3D reconstruction method based on biomimetic stereo vision as described in claim 11, characterized in that, The step of estimating the three-dimensional coordinates of non-common feature points based on the three-dimensional coordinates of the common feature points includes: Based on the three-dimensional coordinates of the common feature points, the three-dimensional coordinates of the non-common feature points are estimated using the single feature point estimation method. or, Based on the three-dimensional coordinates of the common feature points, the three-dimensional coordinates of the non-common feature points are estimated using the dual feature point estimation method.

13. The 3D reconstruction method based on biomimetic stereo vision as described in claim 12, characterized in that, The step of estimating the three-dimensional coordinates of non-common feature points using a single feature point estimation method based on the three-dimensional coordinates of the common feature points includes: Using the dual gradient pattern of feature points in the overall texture map of the left image and the overall texture map of the right image, the three-dimensional coordinates of unmatched non-common feature points are estimated using a single feature point estimation method; For non-common feature points, obtain a common feature point whose pixel coordinates are adjacent to it and whose color difference ratio is within the set color difference range, and use it as a reference point; The depth coordinates of the reference point are used as the depth coordinates of the non-common feature point in the camera coordinate system; After adding or subtracting 1 from the horizontal offset of the reference point, the horizontal coordinates of the non-common feature point in the camera coordinate system are calculated using formula (1); wherein, when the non-common feature point is to the left of the common feature point, the horizontal offset of the reference point is subtracted by 1, and when the non-common feature point is to the right of the common feature point, the horizontal offset of the reference point is added by 1. After adding or subtracting 1 from the vertical offset of the reference point, the vertical coordinates of the non-common feature point in the camera coordinate system are calculated using formula (2); wherein, when the non-common feature point is above the common feature point, the vertical offset of the reference point is subtracted by 1, and when the non-common feature point is below the common feature point, the vertical offset of the reference point is added by 1.

14. The 3D reconstruction method based on biomimetic stereo vision as described in claim 13, characterized in that, The step of estimating the three-dimensional coordinates of non-common feature points using the dual feature point estimation method based on the three-dimensional coordinates of the common feature points includes: By utilizing the dual gradient pattern of feature points in the overall texture map of the left image and the overall texture map of the right image, the three-dimensional coordinates of unmatched non-common feature points are estimated using the dual feature point estimation method. For non-common feature points, two common feature points in the left or right image that are equidistant from the non-common feature points but in opposite directions, whose color difference ratios are within a set color difference range, are obtained as reference points; The mean of the depth coordinates of the two reference points is calculated as the depth coordinate of the non-common feature point in the camera coordinate system; The mean of the horizontal coordinates of the two reference points is calculated as the horizontal coordinate of the non-common feature point in the camera coordinate system; The mean of the vertical coordinates of the two reference points is calculated as the vertical coordinate of the non-common feature point in the camera coordinate system.

15. A three-dimensional reconstruction device based on biomimetic stereo vision, characterized in that, include: An image acquisition module is used to acquire left and right images captured by a dual-camera device when it focuses on a target of interest. The dual-camera device includes a left camera and a right camera. The left camera is equipped with a left sensor, and the right camera is equipped with a right sensor. The photosensitive wafers of the left and right sensors are both spherical structures, with the inner side of the sphere serving as the signal receiving surface. The left and right sensors are respectively translated as a whole towards the temporal side of the camera's optical axis. The texture extraction module is used to extract textures from the left image and the right image respectively using a multi-layer window frame side-antagonistic extraction method to obtain the total texture map of the left image and the total texture map of the right image; The feature matching module is used to perform feature matching on the total texture map of the left image and the total texture map of the right image to generate a dense disparity map. The three-dimensional coordinate acquisition module is used to establish a camera coordinate system with the midpoint of the line connecting the center points of the left sensor and the right sensor as the origin. Using the three-dimensional coordinate calculation formula fixed in the camera coordinate system corresponding to the dual-camera device, it calculates the three-dimensional coordinates of the common feature points matched in the dense parallax map, and estimates the three-dimensional coordinates of non-common feature points based on the three-dimensional coordinates of the common feature points to obtain a three-dimensional coordinate dataset. The 3D reconstruction module is used to perform 3D reconstruction based on the 3D coordinate dataset to generate a stereoscopic view of the target of interest.

16. A three-dimensional reconstruction method and device based on biomimetic stereo vision, characterized in that, The method includes a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor, wherein the processor, when executing the computer program, implements the three-dimensional reconstruction method based on biomimetic stereo vision as described in any one of claims 1-14.

17. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes a stored computer program, wherein, when the computer program is executed, it controls the device on which the computer-readable storage medium is located to perform the three-dimensional reconstruction method based on biomimetic stereo vision as described in any one of claims 1-14.

Citation Information

Patent Citations

  • Binocular stereoscopic vision-based three dimensional human face reconstruction method

    CN106910222A

  • 3D reconstruction method and system for binocular vision

    CN110738731A

  • Improved feature stereo matching method based on binocular vision

    CN113887624A