Trinuclear ray image dense matching method and system
Through the three-core line image dense matching method, pyramid images are generated, parallax proportional factors are calculated, matching cost aggregation and parallax optimization are performed, and noise and occlusion problems in three-dimensional reconstruction are solved, achieving higher precision three-dimensional reconstruction.
Patent Information
- Application Number
- CN202510366806.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-26
- Publication Date
- 2025-07-11
AI Technical Summary
The prior art is difficult to fully utilize the advantages of triple stereo images, and there are limitations on the number of noise, occlusion and matching images, which affects the accuracy of three-dimensional reconstruction.
A three-core line image intensive matching method is provided. By generating pyramid images, calculating the parallax scale factor, performing matching cost aggregation and parallax optimization, using closed-loop projection mapping for consistency detection, and generating three-dimensional point clouds and DSM products.
It improves the accuracy of three-dimensional reconstruction, alleviates noise and occlusion problems, makes full use of the three-view image characteristics, and improves the accuracy and completeness of matching.
Smart Images

Figure CN120298728A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of satellite remote sensing image processing, relates to the field of high-precision three-dimensional reconstruction of three-line array images, and particularly relates to a method and system for dense matching of trinucleus line images. Background Art
[0002] As an important means of obtaining three-dimensional spatial information, the technology of image dense matching is particularly important in the construction of smart cities. However, the existing technologies still face some challenges. First, the noise generated during the sensor imaging process affects the accuracy of dense matching. Second, it is difficult for dense matching technology to handle foreground and background occlusion and the matching of weak texture or repetitive regions. In addition, there are limitations in the number of images for matching in the existing technologies, making it difficult to completely restore the shape of the target ground object. Moreover, the geometric and photometric consistency problems of multi-view images, as well as the significantly increased computational complexity with the increase in the number of viewing angles in multi-view stereo matching algorithms, are also problems that need to be solved by the current technologies.
[0003] Satellites such as Ziyuan-3, SPOT5, and Pleiades provide triple stereo images. Compared with the forward and backward view data of traditional satellite images, this configuration adds the result of the nadir image, allowing better retrieval of heights on the terrain and getting rid of the limitation of only having forward and backward stereo image pairs in classical photogrammetry. However, the mainstream existing technologies for matching using triple stereo images still remain on the basis of using two of the images for multiple groups of matching and fusion, and have not yet exploited the characteristics of triple stereo images.
[0004] Therefore, how to make full use of the advantages of triple stereo images, overcome the technical obstacles of noise, occlusion, and the number of matching images existing in the above classical image dense matching, and improve the three-dimensional reconstruction accuracy is something to be overcome and studied in this field. Summary of the Invention
[0005] The purpose of the present invention is to make full use of the advantages of triple stereo images, provide a triple stereo image matching technology with higher accuracy, and alleviate the problems of noise, occlusion, and the number of matching images existing in the above classical image dense matching. A method for dense matching of trinucleus line images provided by the technical solution of the present invention includes generating pyramid images for the trinucleus line images; calculating the matching cost of the trinucleus line images; performing cost aggregation on the matching cost matrix of the trinucleus line images; calculating the disparity according to the aggregated trinucleus line images; optimizing the disparity of the trinucleus line images; then swapping images for matching according to the characteristics of the trinucleus line images to improve the accuracy; generating a stereo three-dimensional point cloud according to the final matching image disparity map; and generating a DSM image according to the stereo three-dimensional point cloud.
[0006] For this reason, the present invention provides the following technical solutions:
[0007] On the one hand, a trinocular image dense matching method provided by the present invention includes the following steps:
[0008] S1: Obtain trinocular images and generate pyramid images, and select corresponding points to calculate the disparity scale factor between each image in the trinocular images. The trinocular images include a rear view image, a downward view image, and a front view image, and the front view image is selected as the left image, the downward view image is selected as the middle image, and the rear view image is selected as the right image;
[0009] S2: Match the corresponding points on the trinocular images in the epipolar direction, use the disparity scale factor, layer by layer calculate the matching cost of the trinocular images on the pyramid images and introduce cost aggregation to optimize the disparity result to obtain the underlying disparity maps between any two of the left, middle, and right images;
[0010] S3: Perform trinocular image consistency detection and optimization through the formation of a closed-loop projection mapping between the three underlying disparity maps, and output the optimized underlying disparity maps to complete the matching, so as to generate subsequent point clouds and DSM products.
[0011] Perform trinocular image consistency detection and optimization through the formation of a closed-loop projection mapping between the three underlying disparity maps, and output the optimized underlying disparity maps, so as to generate subsequent point clouds and Digital Surface Model (DSM) products.
[0012] Preferably, the calculation formula of the disparity scale factor in step S1 is:
[0013]
[0014] where α represents the disparity scale factor, Bwd, Nad, Fwd represent the rear view, downward view, and front view images, BR, MM, TL represent the lower right corner, middle, and upper left corner points on the trinocular images or in the set research area, col represents the abscissa, such as BwdBRcol represents the abscissa col of the corresponding point corresponding to the lower right corner BR in the rear view image Bwd, NadBRcol represents the abscissa col of the corresponding point corresponding to the lower right corner BR in the downward view image Nad, and by the same token, the meanings of other parameters can be analogized.
[0015] Preferably, the rule for generating the disparity map layer by layer on the pyramid images in step S2 is:
[0016] For the high-level pyramid images, select two perspective images as the main matching images to obtain the disparity map and pass it down;
[0017] For the bottom - layer pyramid image, the disparity ranges of the two viewpoints at the bottom layer are determined by cross - layer transmission using the disparity map of the upper layer. Then, the disparity ranges of other pairwise viewpoints are determined using the cross - view disparity conversion equation based on the disparity scale factor. Furthermore, based on their respective disparity ranges, the matching costs between the three epipolar images are calculated respectively and cost aggregation is introduced to optimize and obtain three bottom - layer disparity maps;
[0018] Among them, the cross - view disparity conversion equation based on the disparity scale factor is:
[0019] d mr =αd lm
[0020] d rl =-(1 + α)d lm
[0021] In the formula, d lm , d mr , d rl respectively represent the disparity values of the same point on the left - middle, middle - right, and right - left disparity maps, and α represents the disparity scale factor.
[0022] For example, for the high - layer pyramid image, two left - middle images or middle - right images are selected as the main matching images to obtain the left - middle disparity map or middle - right disparity map, and it is transmitted to the lower layer,
[0023] For the bottom - layer pyramid image, the left - middle or middle - right disparity ranges at the bottom layer are determined by cross - layer transmission using the left - middle disparity map or middle - right disparity map of the upper layer; then, the left - middle, middle - right, and left - right disparity ranges are determined using the cross - view disparity conversion equation based on the disparity scale factor; furthermore, based on their respective disparity ranges, the matching costs between the three epipolar images are calculated respectively and cost aggregation is introduced to optimize and obtain the three corresponding bottom - layer disparity maps of left - middle, middle - right, and left - right. Among them, the left - middle similar description not only represents the two viewpoints of left - middle, but does not limit the order.
[0024] Preferably, when the left - middle image pair is the main matching view in the high - layer pyramid image, there is a linear mapping relationship between the pixel displacement Δx of the middle image and the compensated displacement Δx' = Δx*(1 + α) of the right image;
[0025] When the left - right image pair is the main matching view in the bottom - layer pyramid image, a reverse displacement compensation mechanism is constructed, and a non - linear mapping relationship between the displacement Δy of the right image and the compensated displacement Δy' = Δy / (1 + α) of the middle image is defined.
[0026] Preferably, in step S2, the matching cost uses the Hamming distance as the cost value;
[0027] Among them, when the left - middle epipolar image is the main matching object, the Hamming distance C is expressed as:
[0028] C(u, v, d) = Hamming(C sl (u, v), C sm (u - d, v), C sr (u - d1, v))
[0029] or
[0030]
[0031] or
[0032]
[0033] wherein, C(u, v, d) is the corresponding Hamming distance C at pixel coordinates u, v and disparity d, u, v are pixel coordinates, d represents the disparity between the left-middle images, C s is the Census calculation, l, m, r represent the left, middle and right images, such as C sl (u, v) represents the result of the Census calculation corresponding to the pixel coordinates u, v in the left image, C sm (u - d, v) represents the result of the Census calculation corresponding to the corresponding point of the pixel coordinates u, v in the middle image, C sr (u - d1, v) represents the result of the Census calculation corresponding to the corresponding point of the pixel coordinates u, v in the right image, d1 represents the disparity between the left and right images, and there is: d1 = (1 + α)d, where α represents the disparity scale factor.
[0034] Preferably, a penalty parameter P3 is added in the calculation of the path energy value of cost aggregation in step S2, and the penalty parameter P3 is expressed as:
[0035]
[0036] wherein, P3(p, d) represents the penalty parameter P3 of pixel p at disparity d, d' represents the disparity of the upper layer pyramid, d'' represents the disparity transmitted to the lower layer pyramid, and the value of k ranges from 0 to 1; Mini is set to be less than the value of kP1, P1 is the normalized average value of the matching costs of all pixels, and the disparity d represents the disparity of the two epipolar images of the main matching object.
[0037] Preferably, in step S3, the three-epipolar image consistency detection by forming a closed-loop projection mapping between the three bottom-layer disparity maps is to obtain the back-projected coordinates by projecting the pixel coordinates and disparity values of a pixel point in one bottom-layer disparity map back to the original bottom-layer disparity map through two forward projections and one back-projection with the other two bottom-layer disparity maps, and then comparing the back-projected coordinates with the original pixel coordinates to complete the three-epipolar image consistency verification;
[0038] Specifically as follows:
[0039] If a, b, and c represent three underlying disparity maps, taking the coordinates of pixel p in the underlying disparity map a and the disparity value as the initial observation values, through the cross-view geometric projection coordinate transformation established based on the trinocular epipolar geometry constraint, the pixel p is orthogonally projected onto the underlying disparity map b to obtain the pixel coordinates and disparity value of the corresponding point in the underlying disparity map b;
[0040] The pixel coordinates and disparity value of the corresponding point in the underlying disparity map b are orthogonally projected onto the underlying disparity map c to obtain the pixel coordinates and disparity value of the corresponding point in the underlying disparity map c;
[0041] The pixel coordinates and disparity value of the corresponding point in the underlying disparity map c are back-projected onto the underlying disparity map a to obtain the back-projection coordinates;
[0042] The distance deviation between the pixel coordinates of pixel p in the underlying disparity map a and the back-projection coordinates is subjected to a threshold check. If it is less than the threshold, it is considered that the trinocular epipolar image consistency check is satisfied at the corresponding pixel p in the underlying disparity map a; otherwise, it is considered that the trinocular epipolar image consistency check is not satisfied.
[0043] Preferably, the process of disparity optimization by forming a closed-loop projection mapping between the three underlying disparity maps in step S3 is an adaptive disparity optimization based on the trinocular epipolar constraint;
[0044] Among them, for the pixel p in the underlying disparity map a that does not satisfy the trinocular epipolar image consistency check, the disparity values of the corresponding points of pixel p in the underlying disparity maps b and c are geometrically corrected using the disparity scale factor to obtain the disparity value projected onto the underlying disparity map a, and then three disparity values including the original disparity value of pixel p in the underlying disparity map a are obtained;
[0045] If the three disparity values are the same, the original disparity value is retained;
[0046] If two of the disparity values are the same, based on the principle of disparity continuity in the occlusion area, the majority-consistent disparity value is selected as the true value, and the other disparity values are updated using the disparity scale factor;
[0047] If all three are different, the disparity value closest to the background depth is selected as the true value according to the background priority criterion, and the other disparity values are updated using the disparity scale factor.
[0048] On the other hand, the present invention also provides a matching system based on the above matching method, including:
[0049] A disparity scale factor determination module for obtaining the trinocular epipolar image and generating a pyramid image, and selecting corresponding points to calculate the disparity scale factor between the images in the trinocular epipolar image. The trinocular epipolar image includes the rear view, the downward view, and the front view images, and the front view image is selected as the left image, the downward view image is selected as the middle image, and the rear view image is selected as the right image;
[0050] A cost calculation and cost aggregation module, configured to match corresponding points on a trinocular image in the epipolar line direction, and use the disparity scale factor to calculate the matching cost of the trinocular image layer by layer on the pyramid image and introduce cost aggregation to optimize the disparity result, so as to obtain the underlying disparity map between any two images in the left middle and lower images;
[0051] A consistency detection and optimization module, configured to perform trinocular image consistency detection and optimization through the formation of a closed-loop projection mapping between three underlying disparity maps, and output an optimized underlying disparity map for generating subsequent point cloud and DSM products.
[0052] Thirdly, the present invention further provides a computer device, including:
[0053] One or more processors;
[0054] A memory storing one or more computer programs;
[0055] Wherein, the processor calls the computer program to implement:
[0056] The steps of a trinocular image dense matching method.
[0057] Fourthly, the present invention further provides a computer-readable storage medium storing a computer program, and the computer program is called by a processor to implement:
[0058] The steps of a trinocular image dense matching method.
[0059] Beneficial effects
[0060] Compared with the existing methods, the advantages of the present invention are as follows:
[0061] 1. The technical solution of the present invention provides a brand-new process for dense matching using trinocular images. By matching corresponding points on the trinocular image in the epipolar line direction, using the disparity scale factor, calculating the matching cost of the trinocular image layer by layer on the pyramid image and introducing cost aggregation to optimize the disparity result to obtain three underlying disparity maps; finally, for the three underlying disparity maps, performing trinocular image consistency detection and optimization through the formation of a closed-loop projection mapping between the three underlying disparity maps, and outputting an optimized underlying disparity map. Through the above technical means, the present invention can make full use of the characteristics of the trinocular images and effectively improve the accuracy of matching, which is different from the classical image dense matching and effectively alleviates the problems of noise, occlusion, and the number of matching images existing in the classical image dense matching.
[0062] 2. The trinucleating line image cost calculation provided by the present invention is a cost calculation method applicable to stereo matching of three-view satellite images. It can simultaneously and synchronously solve the matching costs of three images. Through the calculation of a multi-layer image pyramid, it makes full use of the properties of the trinucleating line image itself. The introduced penalty coefficient P3 is determined by the disparity d in the upper-layer pyramid aggregation and the result of image matching. When the disparity d is within the range of the upper-layer pyramid aggregation disparity, P3 is 0; when d exceeds the pyramid aggregation disparity range and p is successfully matched in the upper-layer pyramid image, P3 is a constant smaller than P1; when d exceeds the pyramid aggregation disparity range and d fails to be matched in the upper-layer pyramid image, P3 is a smaller value.
[0063] 3. The trinucleating line image consistency detection method provided by the present invention is a disparity optimization method applicable to stereo matching of three-view satellite images. It can eliminate the regions with incorrect matches based on three trinucleating line images, and at the same time use the redundant view observation results of the trinucleating line images to fill the disparity hole regions, achieving an accurate matching effect for the image occlusion regions. BRIEF DESCRIPTION OF THE DRAWINGS
[0064] Figure 1 is a schematic diagram of the trinucleating line image matching cost calculation method of the present invention;
[0065] Figure 2 is a comparison effect diagram of DSM of the image data of Ziyuan-3 02 satellite, the fusion result using ERDAS software, and the result of the trinucleating line image semi-global stereo matching method provided by the present invention;
[0066] Figure 3 is a technical idea roadmap provided by an embodiment of the present invention;
[0067] Figure 4 is a schematic diagram of the trinucleating line image consistency detection provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0068] A completely new technology for dense matching using trinucleating line images provided by an embodiment of the present invention is as Figure 3 shown, and its technical idea is as follows:
[0069] S1: Obtain trinucleating line images and generate pyramid images, and select corresponding points to calculate the disparity scale factors between the images in the trinucleating line images. The trinucleating line images include the rear view, the nadir view, and the forward view images. The forward view image is selected as the left image, the nadir view image is selected as the middle image, and the rear view image is selected as the right image;
[0070] S2: Match corresponding points on the triple-epipolar images in the epipolar direction. Using the disparity scale factor, calculate the matching costs of the triple-epipolar images layer by layer on the pyramid images and introduce cost aggregation to optimize the disparity results to obtain the underlying disparity maps between any two of the left, middle, and lower images; among them, preferably, one disparity map is generated in the high-level pyramid image, and three disparity maps are generated in the low-level pyramid images. A cross-view disparity transformation equation is established through the disparity scale factor to achieve the composite expansion from the high-level binocular disparity range to the low-level trinocular disparity range, that is, to achieve multi-image disparity expansion; and the cost function proposed in the present invention is the cost calculation of the triple-epipolar images.
[0071] S3: Perform triple-epipolar image consistency detection and optimization through the formation of a closed-loop projection mapping between the three underlying disparity maps, so as to generate subsequent point cloud and DSM products, and make full use of the closed-loop projection mapping relationship between the three underlying disparity maps for disparity detection and optimization.
[0072] The present invention will be further described below in conjunction with embodiments.
[0073] Embodiment 1
[0074] A triple-epipolar image dense matching method provided by an embodiment of the present invention includes the following steps:
[0075] Step 1: Obtain triple-epipolar images and generate corresponding pyramid images at each level, and select corresponding points within the experimental area. At this time, use the rational function model to calculate the disparity scale factors of the corresponding points between the images.
[0076] The specific implementation method for calculating the disparity scale factor of the corresponding points in this embodiment is as follows:
[0077] Select three points at the upper left corner, lower right corner, and center point of the experimental area. Use UTM (Universal Transverse Mercator) geographic coordinates and global DEM (Digital Elevation Model) files to iteratively calculate the approximate elevation. Then, use the RPC (Rational Polynomial Coefficients) files of the three epipolar images, combined with the approximate three-dimensional coordinates of the three points at the upper left corner, lower right corner, and center point, and respectively back-project them onto the three epipolar images. Then, for the three points at the upper left corner, lower right corner, and center point, calculate the disparity by taking the difference between the abscissas of the corresponding points on the three epipolar images. Then, take the ratios to obtain the disparity scale factors of the corresponding points at the three points at the upper left corner, lower right corner, and center point respectively. Finally, take the average of the three groups of results of the three points at the upper left corner, lower right corner, and center point used as the finally calculated disparity scale factor. Its expression is:
[0078]
[0079] Among them, α represents the scale factor, Bwd, Nad, and Fwd represent the back view, downward view, and front view images, BR, MM, and TL represent the bottom right, middle, and top left points, and col represents the abscissa. For example, BwdBRcol represents the abscissa col of the corresponding homologous point of BR at the bottom right in the back view image Bwd. Similarly, the meanings of other parameters can be analogized. Since the ratio of the parallax of homologous points in the epipolar image to the photographic baseline is equal to the ratio of the focal length to the imaging height, the calculated scale factor of the parallax of homologous points remains fixed.
[0080] Step 2: Use the three epipolar images and the calculated scale factors of the parallax of homologous points between each image to simultaneously match the homologous points on the three images in the epipolar direction. Use the cost calculation method based on the three epipolar images to calculate the matching cost among the three. Calculate the matching cost of the three epipolar images layer by layer in the high-level pyramid and introduce the cost aggregation method in semi-global matching to optimize the parallax result to obtain three low-level parallax maps.
[0081] In the embodiment of the present invention, preferably, a parallax map is generated in the high-level pyramid image. When transmitted to the low level, three low-level parallax maps are generated. Among them, the present invention does not have specific requirements for the order of selecting viewpoints.
[0082] For example, in this embodiment, two left-middle epipolar images are selected as the main matching images in the high-level pyramid image to obtain the high-level parallax map between the left-middle images; while three parallax maps are generated in the low-level pyramid image. The cross-viewpoint parallax conversion equation is established through the scale factor of the parallax to achieve the composite expansion from the high-level dual-viewpoint parallax range to the low-level triple-viewpoint parallax range. That is, in this embodiment, the left-middle parallax map of the upper layer is used for cross-layer transmission to determine the left-middle parallax range of the low level; then the cross-viewpoint parallax conversion equation based on the scale factor of the parallax is used to determine the middle-right and left-right parallax ranges; furthermore, based on their respective parallax ranges, the matching costs between the three epipolar images are calculated respectively and the cost aggregation is introduced for optimization to obtain the three low-level parallax maps corresponding to left-middle, middle-right, and left-right. It should be understood that in other feasible embodiments, the middle-right image can also be selected as the main matching image in the high-level pyramid image.
[0083] Among them, the cross-viewpoint parallax conversion equation based on the scale factor of the parallax is as follows:
[0084] d mr = αd lm
[0085] d rl = -(1 + α)d lm
[0086] In the formula, d lm , d mr , d rlrespectively represent the parallax values of the same point on the left-middle, middle-right, and right-left parallax maps, and α represents the parallax scale factor.
[0087] It should be understood that in the process of calculating the trinocular epipolar image matching cost layer by layer in the high-level pyramid and introducing the cost aggregation method in semi-global matching to optimize the parallax result, a parallax space propagation model is established based on the pyramid hierarchical structure. After performing the cost aggregation step in the high-level pyramid, the initial parallax value is selected using the minimum matching cost criterion to complete the cost calculation step, obtaining the parallax map. Through left-right consistency detection, removing discontinuous regions, median filtering, etc., the parallax optimization step is completed to generate the high-level pyramid parallax map result, which is used as the high-level parallax constraint benchmark. Through the tSGM (publicly recorded in the paper: SURE: Photogrammetric Surface Reconstruction from Imagery) model, cross-level parallax range recursive constraints are performed for layer-by-layer loop to complete the aforementioned cost calculation - cost aggregation - parallax calculation - parallax optimization until the bottom-layer pyramid image is calculated. Among them, the pyramid transmission process can be realized by the prior art. The high-level parallax range constraint is projected to the next layer of the pyramid until the bottom-layer pyramid image. Therefore, the present invention does not elaborate on this.
[0088] The embodiment of the present invention preferably optimizes the cost function. The following will take the left-middle image as the main matching image as an example for illustration. The trinocular epipolar image matching cost calculation method of this embodiment is as follows:
[0089] Select the front view image as the left image, the bottom view image as the middle image, and the rear view image as the right image. For each pixel on the left image, search for the pixels to be matched one by one on the epipolar line of the middle image. After using the Census calculation method to calculate the bit strings of the pixels to be matched on the left image and the middle image, according to the parallax between the abscissa of the left image and the abscissa of the middle image and the previously calculated parallax scale factor α, obtain the abscissa of the right image that is also distributed on the epipolar line, and resample the window according to the offset value after the decimal point to calculate the bit string of the pixel to be matched on the right image. Perform an exclusive OR calculation of the same position for the three bit strings corresponding to the three pixels (assign 0 if the values of the three are the same, and assign 1 if any of the three values are different, as Figure 1 shown), accumulate the assignments of the exclusive OR calculation of the same position, calculate the Hamming distance between the three, and store the Hamming distance as the cost value C in a cost matrix with the length and width of the left image as the matrix length and width and the left-right parallax range as the matrix height;
[0090] The formula is as follows:
[0091] C(u,v,d)=Hamming(C sl (u,v),C sm (u-d,v),Csr (u - d1, v))
[0092] Among them, u and v are the pixel coordinates selected in the left image, and d represents the disparity between the left and middle images (in the high - level pyramid images of this embodiment, the left - middle is the main matching object. If it is the bottom - level pyramid image, such as left - right, then the disparity d1 between the left and right images is set here, and the Hamming distance between the pixel - corresponding bit - strings of the left, middle, and right images is still calculated. That is, according to the relationship of the trinocular - epipolar images, the cost function of any two main matching objects can be constructed by analogy in sequence). C s is for Census calculation. l, m, and r represent the left, middle, and right images respectively, d1 represents the disparity between the left and right images, and the calculation formula is as follows:
[0093] d1 = (1 + α)d
[0094] In addition, in some embodiments, the present invention also proposes two other cost calculation methods for trinocular - epipolar images, that is, calculating the average value and median of pairwise matches within three windows respectively as the result of the Hamming distance. The formulas are as follows:
[0095]
[0096] The above - mentioned cost calculation steps are used to perform dual - matching perspective collaborative cost calculation in the bottom - level pyramid in the high - level pyramid image, and a composite disparity field is generated by using a constrained fusion algorithm.
[0097] In this embodiment, preferably, the left - middle image pair with a short baseline is selected as the main matching perspective in the high - level pyramid image. Based on the trinocular - epipolar geometric relationship, a synchronous search mechanism is constructed, and a linear mapping relationship between the pixel displacement Δx of the middle image and the compensated displacement Δx' = Δx*(1 + α) of the right image is defined to determine the coordinates for cost calculation; the left - right image pair with a long baseline is selected as the main matching perspective in the bottom - level pyramid image, and a reverse displacement compensation mechanism is constructed. A non - linear mapping relationship between the displacement Δy of the right image and the compensated displacement Δy' = Δy / (1 + α) of the middle image is defined to determine the coordinates for cost calculation; and preferably, in the bottom - level pyramid image, the remaining matching perspectives all perform cost calculation according to the image pair with a long baseline.
[0098] After calculating the cost matrix of the left image, cost aggregation is required to optimize the obtained cost matrix. The cost aggregation method is as follows:
[0099] The eight - path aggregation method of SGM (Semi - Global Matching) is adopted to automatically determine the regularization parameter. In the embodiment of the present invention, preferably, a third penalty parameter is added to constrain the deviation of the pyramid disparity. This penalty parameter is determined by the relationship between the disparity d and the disparity range during aggregation of the upper - level pyramid and the quality of the matching result of the upper - level pyramid image. The principle formula is as follows:
[0100] L r (p, d) = C(p, d) + P3(p, d) + min(L r (p - r, d), L r (p - r, d - 1) + P1,
[0101]
[0102] Among them, L r (p, d) represents the energy value of pixel p in the direction of path r at disparity d. i represents the disparity at which the cost is minimized. P1 is the normalized average value of the matching costs of all pixels, and P2 is the maximum value among the matching costs of all pixels. The calculation formula is as follows:
[0103]
[0104] Among them, H and W represent the height and width of the image. ΔS(x, y, l) represents the normalization operation of the matching costs of all pixels within the disparity range. P3 is determined by the result of aggregating the disparity d in the upper - layer pyramid and image matching. When the disparity d is within the aggregating disparity range of the upper - layer pyramid, P3 is 0; when d exceeds the aggregating disparity range of the pyramid and p is successfully matched in the upper - layer pyramid image, P3 is a constant smaller than P1; when d exceeds the aggregating disparity range of the pyramid and d fails to be matched in the upper - layer pyramid image, P3 is a smaller value:
[0105]
[0106] Among them, d′ represents the upper - layer pyramid disparity, d″ represents the disparity passed to the lower - layer pyramid. Usually, d″ = 2d′; the value of k ranges from 0 to 1; Mini is set to a smaller value less than kP1. It should be understood that when there is no matching information in the highest - layer pyramid image, it can be determined by setting an empirical value.
[0107] After the aggregation is completed, an optimized cost matrix is obtained for disparity calculation of the image of this layer of pyramid. This process is a prior art and will not be described in detail herein.
[0108] Step 3: After calculating the bottom - layer disparity map, perform trinocular epipolar image consistency detection and optimization through the formation of a closed - loop projection mapping among the three bottom - layer disparity maps, and output the optimized bottom - layer disparity map for generating subsequent point cloud and DSM products.
[0109] Among them, the trinocular image consistency detection through the formation of a closed-loop projection mapping between three underlying disparity maps is to obtain the back-projected coordinates by projecting the pixel coordinates and disparity values of a pixel point in one underlying disparity map back to the original underlying disparity map through two forward projections and one backward projection with the other two underlying disparity maps, and then compare the back-projected coordinates with the original pixel coordinates to complete the trinocular image consistency verification.
[0110] Specifically as follows:
[0111] If a, b, and c represent three underlying disparity maps, taking the coordinates and disparity value of pixel p in underlying disparity map a as the initial observation values, the pixel p is forward-projected through the cross-view geometric projection coordinate transformation established based on the trinocular geometric constraint to obtain the pixel coordinates and disparity value of the corresponding point in underlying disparity map b;
[0112] The pixel coordinates and disparity value of the corresponding point in underlying disparity map b are forward-projected to underlying disparity map c to obtain the pixel coordinates and disparity value of the corresponding point in underlying disparity map c;
[0113] The pixel coordinates and disparity value of the corresponding point in underlying disparity map c are back-projected to underlying disparity map a to obtain the back-projected coordinates;
[0114] The distance deviation between the pixel coordinates of pixel p in underlying disparity map a and the back-projected coordinates is verified by a threshold. If it is less than the threshold, it is considered that the trinocular image consistency verification is satisfied at the corresponding pixel p in underlying disparity map a; otherwise, it is considered that the trinocular image consistency verification is not satisfied.
[0115] In this embodiment, taking left-middle as an example, as Figure 4 shown, the specific implementation process is as follows:
[0116] (1) Forward projection: First, extract the pixel p coordinates (u lm , v) and its disparity value d lm of the left-middle image disparity map as the initial observation values, establish the cross-view geometric projection coordinate transformation relationship based on the trinocular geometric constraint, and project the left-middle image coordinates (u lm , v) to the corresponding coordinates (u mr , v) of the middle-right image disparity map through coordinate mapping, and synchronously calculate its theoretical disparity value d mr ;
[0117] (2) Second forward projection: Further transmit the middle-right image projection result in the same direction to the corresponding coordinates (u rl , v) on the right-left image disparity map, and calculate its theoretical disparity value d rl ;
[0118] (3) Back-projection verification: Back-project at the corresponding points of the right-left images, and re-project back to the theoretical coordinates (u l ′ m , v) in the left-middle image coordinate system, and compare whether u l ′ m returns to u lm . Specifically, threshold verification is performed based on the distance deviation between the pixel coordinates and the back-projected coordinates. If it is less than the threshold, it is considered to meet the trinocular image consistency verification (the geometric consistency of the trinocular lines); otherwise, it is considered not to meet the trinocular image consistency verification.
[0119] It should be understood that the trinocular image consistency verification set by the present invention is based on the following theoretical formula:
[0120] d lm +d mr +d rl =0
[0121] Through the formula, trinocular line consistency detection can be completed in the mapping, the disparity values at the pixel points that meet this consistency detection are retained, the target points in the occluded, mismatched, and weak texture regions are initially detected, and at the same time, the multi-view expandability of the consistency detection is demonstrated. The idea of this section can be extended to multi-view matching under other corresponding conditions, and the robustness is improved through cyclic consistency.
[0122] For the pixel points that do not pass the multi-view geometric consistency, the technical solution of the present invention uses the disparity scale factor to geometrically correct the disparity parameters of the pixels.
[0123] For example, for the pixel points excluded from the left-middle image disparity map due to non-compliance with the epipolar line consistency, the middle-right d mr right-left d rl are geometrically corrected using the disparity scale factor, that is, using the following formula, the middle-right d mr right-left d rl are converted into the corresponding d lm observation values, which form a multi-source disparity observation set with the original left-middle d lm ;
[0124] d mr / α = d lm
[0125] -d rl / (1 + α) = d lm
[0126] For different spatial context relationships, a three-level decision-making mechanism is designed:
[0127] 1) When the trinocular disparity values are consistent, directly retain the original d lm ;
[0128] 2) When there is a match between the two, based on the principle of parallax continuity in the occluded area, select the most consistent parallax value as the true value estimate;
[0129] 3) When all three are inconsistent, then combine the scene depth distribution characteristics and select the parallax value closest to the background depth through the background priority criterion.
[0130] Finally, morphological filtering is used to remove discontinuous isolated noise regions, and adaptive window median filtering is used to smooth the disparity map, thereby generating an optimized disparity map to improve the integrity of point cloud reconstruction and the product quality of the digital surface model (DSM).
[0131] Embodiment 2
[0132] The embodiment of the present invention provides a matching system based on the above matching method, including:
[0133] A parallax ratio factor determination module, configured to obtain trinocular images and generate pyramid images, and select corresponding points to calculate the parallax ratio factor between each image in the trinocular images. The trinocular images include a rear view, a downward view, and a front view image, and the front view image is selected as the left image, the downward view image is selected as the middle image, and the rear view image is selected as the right image;
[0134] A cost calculation and cost aggregation module, configured to match corresponding points on the trinocular images in the epipolar line direction, use the parallax ratio factor to calculate the matching cost of the trinocular images layer by layer on the pyramid images, and introduce cost aggregation to optimize the disparity result to obtain the underlying disparity maps between any two of the left, middle, and right images;
[0135] A consistency detection and optimization module, configured to perform trinocular image consistency detection and optimization through the formation of a closed-loop projection mapping between the three underlying disparity maps, and output an optimized underlying disparity map for generating subsequent point cloud and DSM products.
[0136] It should be understood that for the specific implementation process of each module, please refer to the above method content, which will not be elaborated herein. The above division of functional modules is only for illustrative purposes. In some embodiments, some functional modules can be combined, some functional modules can be split, and each functional module can be implemented in software or hardware or a combination of software and hardware. Among them, the software and hardware devices include, but are not limited to, general computer devices, programmable gate arrays, digital signal processors, microprocessors, and their corresponding programming or burning software.
[0137] Embodiment 3
[0138] The embodiment of the present invention provides a computer device, including: one or more processors and a memory storing one or more computer programs;
[0139] Wherein, the processor calls the computer program to implement: the steps of a three-core line image dense matching method.
[0140] Among them, the specific implementation is:
[0141] Step 1: Obtain the three-core line images and generate the corresponding pyramid images of each level, and select the same-name points in the experimental area, and use the rational function model to calculate the disparity scale factor of the same-name points between the images;
[0142] Step 2: Use the three-core line images and the calculated disparity scaling factors of the same-name points between the images to match the same-name points on the three images in the epipolar direction at the same time, use the cost calculation method based on the three-core line images to calculate the matching cost between the three, calculate the matching cost of the three-core line images layer by layer in the high-level pyramid, and introduce the cost aggregation method in semi-global matching to optimize the disparity results to obtain three bottom-level disparity maps;
[0143] Step 3: After calculating the underlying disparity map, the three underlying disparity maps are used to form a closed-loop projection mapping to perform consistency detection and optimization of the three-core line images, and the optimized underlying disparity map is output to generate subsequent point cloud and DSM products.
[0144] For the specific implementation process of each step, please refer to the description of the above method.
[0145] It should be understood that in the embodiments of the present invention, the processor may be a central processing unit (CPU), and the processor may also be other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The memory may include a read-only memory and a random access memory, and provide instructions and data to the processor. A portion of the memory may also include a non-volatile random access memory. For example, the memory may also store information about the type of device.
[0146] Example 4
[0147] An embodiment of the present invention provides a computer-readable storage medium storing a computer program, wherein the computer program is called by a processor to implement: steps of a three-core line image dense matching method.
[0148] Among them, the specific implementation is:
[0149] Step 1: Obtain the trinocular image and generate the pyramid images corresponding to each level, and select corresponding points within the experimental area. At this time, use the rational function model to calculate the parallax scale factor of the corresponding points between each image.
[0150] Step 2: Use the trinocular image and the calculated parallax scale factor of the corresponding points between each image to simultaneously match the corresponding points on the three images in the epipolar direction. Use the cost calculation method based on the trinocular image to calculate the matching cost among the three. Calculate the trinocular image matching cost layer by layer in the high-level pyramid and introduce the cost aggregation method in semi-global matching to optimize the parallax result to obtain three low-level parallax maps.
[0151] Step 3: After calculating the low-level parallax maps, perform trinocular image consistency detection and optimization through the closed-loop projection mapping formed among the three low-level parallax maps, and output the optimized low-level parallax maps for generating subsequent point cloud and DSM products.
[0152] For the specific implementation process of each step, please refer to the description of the foregoing method.
[0153] The readable storage medium is a computer-readable storage medium, which can be the internal storage unit of the software and hardware device described in any of the foregoing embodiments, such as the hard disk or memory of the controller. The readable storage medium can also be an external storage device of the controller, such as a plug-in hard disk equipped on the controller, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc. Further, the readable storage medium can also include both the internal storage unit of the controller and the external storage device. The readable storage medium is used to store the computer program and other programs and data required by the controller. The readable storage medium can also be used to temporarily store the data that has been output or will be output.
[0154] Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present invention. And the foregoing readable storage medium includes: USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks or optical disks and other various media that can store program codes.
[0155] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-readable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code. The present application is a device that, when executed by a processor, generates instructions for implementing the functions specified in one or more processes of the flowchart and / or one or more boxes of the block diagram. These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device that implements the functions specified in one or more processes of the flowchart and / or one or more boxes of the block diagram. These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one or more processes of the flowchart and / or one or more boxes of the block diagram.
[0156] It should be emphasized that the examples described in the present invention are illustrative rather than restrictive. Therefore, the present invention is not limited to the examples described in the specific embodiments. Any other embodiments obtained by those skilled in the art based on the technical solutions of the present invention, whether modified or replaced, as long as they do not depart from the spirit and scope of the present invention, also belong to the protection scope of the present invention.
Claims
1. A trinuclear line image dense matching method, characterized in that: It includes the following steps: S1: Obtain the trinocular line images and generate the pyramid images, and select corresponding points to calculate the disparity scale factors between the images in the trinocular line images. The trinocular line images include the rear view, downward view, and front view images, and the front view image is selected as the left image, the downward view image as the middle image, and the rear view image as the right image; S2: Match the corresponding points on the trinocular line images in the epipolar direction. Using the disparity scale factors, calculate the matching costs of the trinocular line images layer by layer on the pyramid images and introduce cost aggregation to optimize the disparity results to obtain the underlying disparity maps between any two of the left, middle, and lower images; S3: Perform trinocular line image consistency detection and optimization through the formation of a closed-loop projection mapping between the three underlying disparity maps, and output the optimized underlying disparity maps to complete the matching, so as to generate subsequent point cloud and DSM products.
2. The method according to claim 1, wherein: The calculation formula for the disparity scale factor in step S1 is: where α represents the disparity scale factor, Bwd, Nad, Fwd represent the rear view, downward view, and front view images, BR, MM, TL represent the lower right corner point, middle point, and upper left corner point on the trinocular line images or in the set research area, and col represents the abscissa.
3. The method according to claim 1, characterized in that: The rule for generating the disparity maps layer by layer on the pyramid images in step S2 is: For the high-level pyramid images, select two perspective images as the main matching images to obtain the disparity maps and pass them down; For the low-level pyramid images, use the disparity maps of the previous layer for cross-layer transfer to determine the disparity ranges of the two perspectives at the bottom layer, and then use the cross-perspective disparity conversion equation based on the disparity scale factor to determine the disparity ranges of the other two perspectives. Then, based on their respective disparity ranges, calculate the matching costs between the trinocular line images and introduce cost aggregation to optimize and obtain three underlying disparity maps; Among them, the cross-perspective disparity conversion equation based on the disparity scale factor is: d mr = αd lm d rl = -(1 + α)d lm where d lm , d mr , d rl respectively represent the parallax values of the same point on the left-middle, middle-right, and right-left parallax maps, and α represents the parallax scale factor.
4. The method according to claim 3, wherein: When the left-middle image pair is the main matching perspective in the high-level pyramid images, the linear mapping relationship between the pixel displacement Δx of the middle image and the compensated displacement Δx' = Δx*(1 + α) of the right image; When the left-right image pair is the main matching perspective in the low-level pyramid images, construct a reverse displacement compensation mechanism, and define the non-linear mapping relationship between the displacement Δy of the right image and the compensated displacement Δy' = Δy / (1 + α) of the middle image.
5. The method according to claim 1, characterized in that: In step S2, the matching cost is the Hamming distance as the cost value; Among them, when the left-middle epipolar line image is the main matching object, the Hamming distance C is expressed as: C(u, v, d) = Hamming(C sl (u, v), C sm (u - d, v), C sr (u - d1, v)) or or Wherein, C(u, v, d) is the corresponding Hamming distance C at pixel coordinates u, v and disparity d, u and v are pixel coordinates, d represents the disparity between the left and middle images, and C s is the Census calculation, l, m, and r represent the left, middle, and right images, d1 represents the disparity between the left and right images, and there is: d1 = (1 + α)d, where α represents the disparity scale factor.
6. The method according to claim 1, characterized in that: In the calculation of the path energy value of cost aggregation in step S2, add a penalty parameter P3, and the penalty parameter P3 is expressed as: In the formula, P3(p, d) represents the penalty parameter P3 of pixel p at disparity d, d′ represents the disparity of the upper-level pyramid, d″ represents the disparity transferred to the lower-level pyramid, and the value of k ranges from 0 to 1; Mini is set to a value less than kP1, P1 is the normalized average value of the matching costs of all pixels, and the disparity d represents the disparity between the two epipolar line images of the main matching object.
7. The method according to claim 1, wherein: The trinocular epipolar image consistency detection by forming a closed-loop projection mapping between three underlying disparity maps in step S3 is to obtain the back-projected coordinates by projecting the pixel coordinates and disparity values of a pixel in one underlying disparity map back to the original underlying disparity map through two forward projections and one backward projection with the other two underlying disparity maps, and then compare the back-projected coordinates with the original pixel coordinates to complete the trinocular epipolar image consistency verification; Specifically as follows: If a, b, and c represent the three underlying disparity maps, the coordinates and disparity value of pixel p in the underlying disparity map a are used as the initial observations, and the pixel p is forward projected to the underlying disparity map b through the cross-view geometric projection coordinate transformation based on the trinocular epipolar geometric constraint to obtain the pixel coordinates and disparity value of the corresponding point in the underlying disparity map b; The pixel coordinates and disparity value of the corresponding point in the underlying disparity map b are forward projected to the underlying disparity map c to obtain the pixel coordinates and disparity value of the corresponding point in the underlying disparity map c; The pixel coordinates and disparity value of the corresponding point in the underlying disparity map c are back-projected to the underlying disparity map a to obtain the back-projected coordinates; The distance deviation between the pixel coordinates of pixel p in the underlying disparity map a and the back-projected coordinates is verified by a threshold. If it is less than the threshold, it is considered that the trinocular epipolar image consistency verification is satisfied at the corresponding pixel p in the underlying disparity map a; otherwise, it is considered that the trinocular epipolar image consistency verification is not satisfied.
8. The method according to claim 7, wherein: The process of disparity optimization by forming a closed-loop projection mapping between three underlying disparity maps in step S3 is an adaptive disparity optimization based on the trinocular epipolar constraint; Among them, for pixel p in the underlying disparity map a that does not satisfy the trinocular epipolar image consistency verification, the disparity values of the corresponding points of pixel p in the underlying disparity maps b and c are geometrically corrected using the disparity scale factor to obtain the disparity values projected to the underlying disparity map a, and then three disparity values including the original disparity value of pixel p in the underlying disparity map a are obtained; If the three disparity values are the same, the original disparity value is retained; If two of the disparity values are the same, based on the principle of disparity continuity in the occluded area, the most consistent disparity value is selected as the true value, and the other disparity values are updated using the disparity scale factor; If all three are different, the disparity value closest to the background depth is selected as the true value according to the background priority criterion, and the other disparity values are updated using the disparity scale factor.
9. A matching system based on the method according to any one of claims 1-8, characterized in that: Including: A disparity scale factor determination module, configured to obtain the trinocular epipolar image and generate a pyramid image, and select corresponding points to calculate the disparity scale factor between each image in the trinocular epipolar image. The trinocular epipolar image includes a rear view, a downward view, and a front view image, and the front view image is selected as the left image, the downward view image is selected as the middle image, and the rear view image is selected as the right image; A cost calculation and cost aggregation module, configured to match corresponding points on the trinocular epipolar image in the epipolar direction, use the disparity scale factor to calculate the trinocular epipolar image matching cost layer by layer on the pyramid image, and introduce cost aggregation to optimize the disparity result to obtain the underlying disparity maps between any two of the left, middle, and right images; The consistency detection optimization module is used to perform trinocular epipolar image consistency detection and optimization through the formation of a closed-loop projection mapping among three underlying disparity maps, and output the optimized underlying disparity maps for generating subsequent point cloud and DSM products.
10. A computer device, characterized in that: It includes: One or more processors; A memory storing one or more computer programs; Wherein, the processor calls the computer program to implement: The steps of the method according to any one of claims 1-8.