3D Model Construction Method, Device and Equipment
By constructing the target graph structure in large-scale scenarios and using the cascading hash matching method, the problem of slow construction of three-dimensional models in the existing technology is solved, and a more efficient construction of three-dimensional models is achieved.
Patent Information
- Application Number
- CN202111226009.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-10-21
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2041-10-21
AI Technical Summary
The problem of slow construction of three-dimensional models in large-scale scenarios in the prior art.
By obtaining the image collection, the vertices and edges in the target graph structure are determined, and the edges represent the correlation relationship and matching information between the two scene pictures. The matching point pairs are quickly calculated using the cascading hash matching method to build a three-dimensional model of the target scene.
The calculation efficiency of matching relationships between scene pictures is improved, and the construction efficiency of three-dimensional models is significantly improved.
Smart Images

Figure CN114140575B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer technology, and particularly to a method, apparatus, and device for constructing a three-dimensional model. Background Art
[0002] With the continuous development and popularization of image processing technology, three-dimensional model construction is widely used in fields such as medicine, architecture, and art. Three-dimensional model construction is to construct a point cloud of a target object or a target scene based on a series of pictures in a picture set.
[0003] Currently, most three-dimensional model construction algorithms obtain matching point pairs from single-view pictures, multi-view picture sets, or key frames of videos. Then, based on these matching point pairs in the pictures, they are stitched together to construct a sparse point cloud of the target object or the target scene.
[0004] However, in the construction of three-dimensional models in large-scale scenes in the prior art, there is a problem of slow construction speed. Summary of the Invention
[0005] The present application provides a method, apparatus, and device for constructing a three-dimensional model to solve the problem of slow construction speed of three-dimensional models in large-scale scenes in the prior art.
[0006] In a first aspect, the present application provides a method for constructing a three-dimensional model, including:
[0007] Obtain a picture set, where the picture set includes at least one scene picture of a target scene and picture information of each scene picture;
[0008] According to the scene pictures in the picture set and the picture information of each scene picture, determine a target graph structure, where the vertices in the target graph structure are the scene pictures in the picture set, and the edges in the target graph structure are used to indicate the matching information between the two scene pictures connected by the edge;
[0009] Construct a three-dimensional model of the target scene according to the target graph structure.
[0010] Optionally, the picture information includes the shooting coordinates when shooting the target scene. The determining of the target graph structure according to the scene pictures in the picture set and the picture information of each scene picture includes:
[0011] According to the shooting coordinates of each scene picture in the picture set, determine the association relationship between the scene pictures in the picture set in pairs;
[0012] Construct the connection relationship of the target graph structure according to the association relationship between the scene pictures in the picture set, where the edges in the target graph structure are used to indicate the existence of an association relationship between the two scene pictures connected by the edge;
[0013] Determine the matching information corresponding to each edge in the target graph structure according to the connection relationship of the target graph structure, where the matching information includes the matching point pairs between the two scene pictures corresponding to the edge.
[0014] Optionally, the determining the matching information corresponding to each edge in the target graph structure according to the connection relationship of the target graph structure includes:
[0015] Extract the key points of the first picture from the points of the first picture and the key points of the second picture from the points of the second picture according to the feature point extraction algorithm, where the first picture and the second picture are the two scene pictures connected by the edge;
[0016] Determine the matching key point pairs in the first picture and the second picture according to the key points of the first picture and the key points of the second picture;
[0017] Estimate the fundamental matrix for indicating the transformation relationship between the first picture and the second picture according to the key point pairs;
[0018] Determine the matching information between the first picture and the second picture according to the points in the first picture, the points in the second picture, and the fundamental matrix, where the matching information includes the matching point pairs between the first picture and the second picture and the number of the matching point pairs.
[0019] Optionally, the determining the matching information between the first picture and the second picture according to the points in the first picture, the points in the second picture, and the fundamental matrix includes:
[0020] Perform rough matching on the points to be matched in the first picture and the points in the second picture according to the first preset hashing algorithm and the fundamental matrix to obtain the first point set that matches the points to be matched in the second picture;
[0021] Perform fine matching on the points to be matched and the points in the first point set according to the second preset hashing algorithm and the fundamental matrix to obtain the second point set that matches the points to be matched in the first point set, where the number of points in the second point set is less than the number of points in the first point set;
[0022] According to a preset matching algorithm and the basic matrix, match the points to be matched and the points in the second point set one by one to determine a target matching point that matches the point to be matched, where the target matching point is a point in the second point set.
[0023] Optionally, determining the association relationship between every two of the scene pictures in the picture set according to the shooting coordinates of each of the scene pictures in the picture set includes:
[0024] Determine the distance between every two of the scene pictures according to the shooting coordinates of each of the scene pictures in the picture set;
[0025] When the distance between every two of the scene pictures is less than or equal to a distance threshold, determine that there is an association relationship between every two of the scene pictures.
[0026] Optionally, the picture information includes the camera internal parameters when shooting the target scene, and constructing the three-dimensional model of the target scene according to the target graph structure includes:
[0027] Obtain an edge in the target graph structure and two scene pictures corresponding to the edge;
[0028] Construct the three-dimensional models of the two scene pictures according to the camera internal parameters, the basic matrix, and the matching information of the edge.
[0029] In a second aspect, the present application provides a three-dimensional model construction device, including:
[0030] An acquisition module, configured to acquire a picture set, where the picture set includes at least one scene picture of a target scene and the picture information of each of the scene pictures;
[0031] A processing module, configured to determine a target graph structure according to the scene pictures in the picture set and the picture information of each of the scene pictures, where the vertices in the target graph structure are the scene pictures in the picture set, and the edges in the target graph structure are used to indicate the matching information between the two scene pictures connected by the edge; construct the three-dimensional model of the target scene according to the target graph structure.
[0032] Optionally, the processing module is specifically configured to:
[0033] Determine the association relationship between every two of the scene pictures in the picture set according to the shooting coordinates of each of the scene pictures in the picture set;
[0034] Construct the connection relationship of the target graph structure according to the association relationship between every two of the scene pictures in the picture set, where the edges in the target graph structure are used to indicate that there is an association relationship between the two scene pictures connected by the edge;
[0035] Determine the matching information corresponding to each edge in the target structure diagram according to the connection relationship of the target structure diagram, where the matching information includes the matching point pairs between the two scene pictures corresponding to the edge.
[0036] Optionally, the processing module is specifically configured to:
[0037] Extract the key points of the first picture from the points of the first picture and the key points of the second picture from the points of the second picture according to the feature point extraction algorithm, where the first picture and the second picture are the two scene pictures connected by the edge;
[0038] Determine the matching key point pairs in the first picture and the second picture according to the key points of the first picture and the key points of the second picture;
[0039] Estimate the fundamental matrix for indicating the transformation relationship between the first picture and the second picture according to the key point pairs;
[0040] Determine the matching information between the first picture and the second picture according to the points in the first picture, the points in the second picture, and the fundamental matrix, where the matching information includes the matching point pairs between the first picture and the second picture and the number of the matching point pairs.
[0041] Optionally, the processing module is specifically configured to:
[0042] Perform rough matching on the points to be matched in the first picture and the points in the second picture according to the first preset hash algorithm and the fundamental matrix to obtain the first point set that matches the points to be matched in the second picture;
[0043] Perform fine matching on the points to be matched and the points in the first point set according to the second preset hash algorithm and the fundamental matrix to obtain the second point set that matches the points to be matched in the first point set, where the number of points in the second point set is less than the number of points in the first point set;
[0044] Perform one-by-one matching on the points to be matched and the points in the second point set according to the preset matching algorithm and the fundamental matrix to determine the target matching point that matches the points to be matched, where the target matching point is the point in the second point set.
[0045] Optionally, the processing module is specifically configured to:
[0046] Determine the distance between every two of the scene pictures according to the shooting coordinates of each scene picture in the picture set;
[0047] When the distance between any two of the scene pictures is less than or equal to the distance threshold, it is determined that there is an association relationship between any two of the scene pictures.
[0048] Optionally, the processing module is specifically configured to:
[0049] Obtain an edge in the target graph structure and two scene pictures corresponding to the edge;
[0050] Construct three-dimensional models of the two scene pictures according to the camera internal parameters, the fundamental matrix, and the matching information of the edge.
[0051] In a third aspect, the present application provides an electronic device, including: a memory and a processor;
[0052] The memory is used to store a computer program; the processor is configured to execute the three-dimensional model construction method in the first aspect and any possible design of the first aspect according to the computer program stored in the memory.
[0053] In a fourth aspect, the present application provides a readable storage medium, in which a computer program is stored. When at least one processor of the electronic device executes the computer program, the electronic device executes the three-dimensional model construction method in the first aspect and any possible design of the first aspect.
[0054] In a fifth aspect, the present application provides a computer program product, the computer program product includes a computer program. When at least one processor of the electronic device executes the computer program, the electronic device executes the three-dimensional model construction method in the first aspect and any possible design of the first aspect.
[0055] The three-dimensional model construction method provided by the present application obtains a picture set from a drone or other terminal device; determines vertices in the target graph structure according to the scene pictures in the picture set; determines the association relationship between any two scene pictures according to the picture information of each scene picture in the picture set; when there is an association relationship between any two scene pictures, add an edge for the two scene pictures in the target graph structure to obtain the target graph structure; determine the matching information between the two scene pictures connected by the edge; and constructs a three-dimensional model of the target scene according to the target graph structure, thereby achieving the effect of improving the calculation efficiency of the matching relationship between scene pictures and improving the construction efficiency of the three-dimensional model. Description of the Drawings
[0056] To more clearly illustrate the technical solutions in the present application or the prior art, the following will briefly introduce the accompanying drawings required for the description of the embodiments or the prior art. Obviously, the accompanying drawings in the following description are some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can also be obtained based on these drawings.
[0057] Figure 1 Schematic diagram of a scenario for three-dimensional model construction provided by an embodiment of the present application;
[0058] Figure 2 Flowchart of a three-dimensional model construction method provided by an embodiment of the present application;
[0059] Figure 3 Schematic diagram of a scenario picture provided by an embodiment of the present application;
[0060] Figure 4 Schematic diagram of epipolar constraint provided by an embodiment of the present application;
[0061] Figure 5 Flowchart of a three-dimensional model construction method provided by an embodiment of the present application;
[0062] Figure 6 Schematic diagram of a data scheduling method in the feature point matching stage provided by an embodiment of the present application
[0063] Figure 7 Schematic diagram of an optimization scheme for data transmission between child nodes provided by an embodiment of the present application;
[0064] Figure 8 Flowchart of a three-dimensional model construction method provided by an embodiment of the present application;
[0065] Figure 9 Schematic diagram of the structure of a three-dimensional model construction device provided by an embodiment of the present application;
[0066] Figure 10 Schematic diagram of the hardware structure of a server provided by an embodiment of the present application;
[0067] Figure 11 Schematic diagram of the structure of a distributed server cluster provided by an embodiment of the present application. Detailed implementation manners
[0068] To make the objectives, technical solutions, and advantages of this application clearer, the following will clearly and completely describe the technical solutions in this application in conjunction with the accompanying drawings in this application. Obviously, the described embodiments are part of the embodiments of this application, rather than all of them. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in this application without creative efforts shall fall within the scope of protection of this application.
[0069] With the continuous development and popularization of image processing technology, three-dimensional (3D) model construction has been widely applied in fields such as healthcare, architecture, and art. 3D model construction is to construct the point cloud of a target object or target scene based on a series of images in an image set. Currently, the research results of 3D models have been widely applied in fields such as healthcare, architecture, and art. For example, 3D models have been applied in application scenarios such as virtual surgery, map reconstruction, geomorphic terrain reconstruction, ancient cultural relic protection, and game design. With the continuous development of imaging technology, the difficulty of obtaining high-definition multi-view image sets has been greatly reduced. For example, the aerial image set captured by a drone has advantages such as safety, stability, speed, low cost, and large scale. Moreover, the aerial image set captured by a drone can be taken from multiple angles and the clarity can be adjusted according to actual needs to enrich the texture. Therefore, using the aerial image set captured by a drone as the input image set for constructing a 3D model can make the point cloud of the target object or target scene for 3D model construction more complete and accurate.
[0070] In the prior art, most 3D model construction algorithms obtain matching point pairs from single-view images, multi-view image sets, or key frames of videos. Then, based on these matching point pairs in the images, they are stitched together to construct the sparse point cloud of the target object or target scene. These 3D model construction algorithms are mostly applicable to indoor scenes or small scenes. In the construction of 3D models for large-scale scenes, the image set will include many more scene images than indoor scenes or small scenes. If the existing 3D model construction algorithms are directly used to construct the sparse point cloud of the target object or target scene based on the matching point pairs, it will take a very long time, and there is a problem of slow 3D model construction speed.
[0071] In view of the above problems, the present application proposes a three-dimensional model construction method. In the present application, after the server obtains a set of pictures of the target scene, it can screen the association relationship of the pictures in the set of pictures according to the GPS information of each scene picture in the set of pictures. When the server determines that the distance between two scene pictures is less than the distance threshold, there is an association relationship between the two scene pictures. The server can generate a target structure diagram according to the association relationship between every two scene pictures in the set of pictures. In the target structure diagram, each vertex is a picture in the set of pictures. The edges in the target structure diagram are used to indicate that there is an association relationship between the two scene pictures connected by the edge. After determining the association relationship between every two scene pictures in the set of pictures, for two scene pictures with an association relationship, the server can use the cascaded hash matching method to quickly determine the matching information between the two scene pictures. The matching information may include pairs of matching points of the two scene pictures. The server can use the pair of matching points as the information of the edge. The server can successively obtain the two scene pictures connected by the edge according to the target graph structure, and merge the three-dimensional scenes constructed by the two scene pictures into the three-dimensional scene of the target scene. By using the cascaded hash matching method, the present application achieves the effect of improving the model construction efficiency.
[0072] Figure 1 Fig. 4 shows a schematic diagram of a scene for constructing a three-dimensional model provided by an embodiment of the present application. After a drone captures scene pictures of a target scene, these scene pictures are input into a server. After the server obtains the set of pictures composed of these scene pictures, it constructs a three-dimensional model of the target scene according to the three-dimensional model construction method in the following embodiments. The server outputs the constructed three-dimensional model to other terminal devices to implement the application of the three-dimensional model. For example, the three-dimensional model is applied to scenarios such as virtual surgery, map reconstruction, landform reconstruction, ancient cultural relic protection, and game design.
[0073] Among them, the drone can also send the captured target scene to the first terminal device in real time. The first terminal device is used to organize the scene pictures into a set of pictures. The first terminal device can send the organized set of pictures to the server after completing the shooting of the target scene. Alternatively, the first terminal device can also divide the target scene into multiple regions. When the first terminal device obtains all the scene pictures captured by the drone in one region, the first terminal device can send the partial scene pictures to the server.
[0074] Among them, the server can be a server for implementing the following three-dimensional model construction method. Alternatively, the server can also be the main server in a server cluster for implementing the following three-dimensional model construction method.
[0075] In this application, the server is used as the execution entity to execute the 3D model construction method of the following embodiments. Specifically, the execution entity can be the hardware device of the server, or the software application in the server that implements the following embodiments, or the computer-readable storage medium installed with the software application that implements the following embodiments.
[0076] Figure 2 The flowchart of a 3D model construction method provided by an embodiment of the present application is shown. Based on the embodiment shown in Figure 1 On the basis of the embodiment shown, as shown in Figure 2 shown, with the server as the execution entity, the method of this embodiment may include the following steps:
[0077] S101. Obtain a set of pictures, where the set of pictures includes at least one scene picture of the target scene and the picture information of each scene picture.
[0078] In this embodiment, the server obtains the set of pictures from a drone or other terminal device. The set of pictures includes at least one scene picture and the picture information corresponding to each scene picture. All the scene pictures in the set of pictures can cover every detail in the target scene. The picture information of a scene picture may include the camera internal parameters for taking the scene picture, the shooting coordinates for taking the scene picture, etc.
[0079] Among them, the target scene can be a large-scale area, such as a certain island, a certain farm, a certain forest farm, etc. Or, the target scene can also be a large-scale target object, such as a ship.
[0080] S102. Determine a target graph structure according to the scene pictures in the set of pictures and the picture information of each scene picture. The vertices in the target graph structure are the scene pictures in the set of pictures, and the edges in the target graph structure are used to indicate the matching information between the two scene pictures connected by the edge.
[0081] In this embodiment, the server can determine the vertices in the target graph structure according to the scene pictures in the set of pictures. Among them, each scene picture corresponds to a vertex. The server can also determine the association relationship between two scene pictures according to the picture information of each scene picture in the set of pictures. When there is an association relationship between two scene pictures, the server can add an edge in the target graph structure for the two scene pictures. The server can also determine the matching information between the two scene pictures according to the two scene pictures connected by the edge.
[0082] In one example, the process of determining the above target graph structure can be summarized by the following steps, and its specific steps may include:
[0083] Step 1: Determine the association relationship between pairwise scene pictures in the picture set according to the shooting coordinates of each scene picture in the picture set.
[0084] In this step, as Figure 3 shown are two scene pictures with an association relationship. There are partial matching features in these two scene pictures. For example, Figure 3 the rivers and roads in the left picture and the right picture are matching features. In 3D model construction, these matching features will be overlapped in the form of matching point pairs. Therefore, scene pictures with an association relationship can be considered as scene pictures with matching features. As Figure 3 shown, connect the matching point pairs in the left picture and the right picture with straight lines.
[0085] In the prior art, using the traditional method to calculate the matching point pairs between two scene pictures can accurately determine the association relationship between the two scene pictures. However, in the construction of a 3D model of a large-scale scene, the picture set includes a large number of scene pictures, and these scene pictures are usually high-definition pictures. In such a picture set, if it is necessary to use the traditional method to calculate the matching features between pairwise scene pictures to determine the association relationship, its time complexity is at the square level of the number of pictures, which is time-consuming and inefficient. Moreover, in a large-scale scene, there is not necessarily an association relationship between the scene pictures collected by drone aerial photography. Therefore, before using the traditional method to calculate the matching features of two scene pictures, it is obvious that a preliminary screening of whether there may be matching features between pairwise scene pictures can greatly improve the construction efficiency of the 3D model.
[0086] Actually, during the actual shooting process, the camera of the drone is fixed. Therefore, when there are matching features in two scene pictures, the shooting positions of these two scene pictures must be relatively close. Generally, there are no meaningful matching features between two scene pictures taken at a relatively far distance. Therefore, in this application, the server can roughly estimate whether there are matching features between pairwise scene pictures by using the shooting coordinates. When there are matching features between two scene pictures, it can be considered that there is an association relationship between these two scene pictures. Otherwise, there is no association relationship between these two scene pictures. Specifically, the process for the server to determine the association relationship between pairwise scene pictures according to the shooting coordinates of the scene pictures may include the following steps:
[0087] Step 1.1: Determine the distance between pairwise scene pictures according to the shooting coordinates of each scene picture in the picture set.
[0088] In this step, the shooting coordinates are used to indicate the position of the drone in the three-dimensional space when shooting this scene picture. The server can calculate the distance between these two scene pictures in this three-dimensional space according to the shooting coordinates of the two scene pictures.
[0089] In one implementation, the shooting coordinates can be GPS coordinates. The GPS coordinates can include longitude, latitude, and altitude. A GPS module can also be installed in the drone that shoots the scene picture. The GPS module is used to generate the GPS coordinates of the scene picture correspondingly when the drone shoots the scene picture. The GPS coordinates are stored together with the scene picture as picture information.
[0090] In the actual calculation process, a Cartesian three-dimensional model is required to correctly calculate the Euclidean distance between two scene pictures. The longitude, latitude, and altitude of the GPS coordinates use a spherical model. And if the GPS coordinates are directly used to calculate the distance between two scene pictures, the geoid deviation e will affect the calculation result. Therefore, in order to improve the calculation efficiency and accuracy, this application first needs to convert the shooting coordinates from the geocentric coordinate system (World Geodetic System - 1984 Coordinate System, WGS-84) to the Earth-centered, Earth-fixed Coordinate System (ECEF). The specific conversion formula is as follows:
[0091]
[0092] where P E is the shooting coordinate in the ECEF coordinate system. N is the radius of the Earth's curvature. e is the eccentricity of the Earth. h is the height of the camera when shooting the scene picture. is the latitude of the camera when shooting the scene picture. λ is the longitude of the camera when shooting the scene picture.
[0093] The calculation formula for the eccentricity e of the Earth is:
[0094]
[0095] The calculation formula for the radius of the Earth's curvature N is:
[0096]
[0097] where a is the length of the semi-major axis of the Earth, its value is 6378137, and the unit is meters. f is the flattening of the Earth, and its value is 1 / 298.257223563. b is the length of the semi-minor axis, and the unit is meters. The calculation formula for b is:
[0098] b = a(1 - f)
[0099] The server can directly use the shooting coordinate P in the ECEF coordinate system ECalculate the distance between two scene pictures. This calculation process does not consume excessive computing resources. Denote the shooting coordinates of the two scene pictures as P1(x1, y1, z1) and P2(x2, y2, z2) respectively. The server can calculate the Euclidean distance between the two scene pictures in the three-dimensional space through the following formula:
[0100]
[0101] The server needs to calculate the distance between every two scene pictures in the picture set. When the picture set includes N pictures, the server needs to calculate a total of the distance between scene pictures.
[0102] Step 1.2: When the distance between every two scene pictures is less than or equal to the distance threshold, determine that there is an association relationship between every two scene pictures.
[0103] In this step, a distance threshold can be preset in the server. The server can compare the distance between two scene pictures with the distance threshold. When the distance between two scene pictures is less than the distance threshold, the server can determine that there is an association relationship between the two scene pictures. Otherwise, when the distance between two scene pictures is greater than or equal to the distance threshold, the server can determine that there is no association relationship between the two scene pictures. The distance threshold can be an empirical value. Or, the distance threshold can be calculated by the server according to all the distances in the picture set. For example, when the picture set includes N pictures, the server determines the mean value of the distances as the distance threshold, or the server can determine the median value of the distances as the distance threshold.
[0104] In one implementation, since the distance between each picture and other pictures may vary, using a single distance threshold may result in a one-size-fits-all situation, leading to inaccurate determination of the association relationship. Therefore, in this implementation, the server can determine the distance threshold for each scene picture according to the distance between each scene picture and other pictures. The server determines the association relationship between the scene picture and other pictures according to the distance threshold of the scene picture. Specifically, the process of determining the association relationship can include:
[0105] Step 1.2.1: The server determines a scene picture in the picture set as the picture to be associated. The distance is calculated between the picture to be associated and each other scene picture in the picture set. For example, when the picture set includes N pictures, N - 1 distances can be determined according to the picture to be associated.
[0106] Step 1.2.2: The server determines the maximum distance maxD and the minimum distance minD among them according to the distances between the to-be-associated picture and each of the other scene pictures in the picture set.
[0107] Step 1.2.3: The server determines the distance threshold of the to-be-associated picture according to the following formula:
[0108]
[0109] where D φ is the distance threshold of the to-be-associated picture. The subscript φ of the distance threshold D φ represents the distance threshold of the φ-th scene picture in the picture set. N is the number of pictures in the picture set. δ l is a scale constant, 10 ≤ δ l ≤ 20.
[0110] Step 1.2.4: The server compares the distances between the to-be-associated picture and each of the other scene pictures in the picture set with the distance threshold. When the distance between a scene picture and the to-be-associated picture is less than the distance threshold, there is an association relationship between the scene picture and the to-be-associated picture. Otherwise, when the distance between a scene picture and the to-be-associated picture is greater than or equal to the distance threshold, there is no association relationship between the scene picture and the to-be-associated picture.
[0111] It should be noted that there may be a situation when calculating using this implementation method: when the first scene picture is used as the to-be-associated picture, the distance between the first scene picture and the second scene picture is less than the distance threshold of the first scene picture. That is, there is an association relationship between the first scene picture and the second scene picture. When the second scene picture is used as the to-be-associated picture, the distance between the second scene picture and the first scene picture is greater than the distance threshold of the second scene picture. That is, there is no association relationship between the second scene picture and the first scene picture. For this situation, it can be considered that there is an association relationship between the first scene picture and the second scene picture to avoid missing two scene pictures with an association relationship during the calculation process, which may lead to abnormalities in the three-dimensional model construction.
[0112] The server completes the first screening of the scene pictures in the picture set by determining whether there is an association relationship between two pictures. In subsequent calculations, the server only needs to perform matching calculations for two scene pictures with an association relationship, and no longer needs to calculate the matching information for pairwise scene pictures in the picture set, greatly reducing unnecessary calculation amounts.
[0113] Step 2: According to the correlation relationships between pairwise scene images in the image set, construct the connection relationships of the target graph structure, where the edges in the target graph structure are used to indicate that there are correlation relationships between the two scene images connected by the edge.
[0114] In this step, the server takes each scene image in the image set as a vertex in the target graph structure. When there is a correlation relationship between two scene images, the server adds an edge between the vertices corresponding to the two scene images to determine the connection relationship between the two vertices. That is, the edges in the target graph structure are used to indicate that there are correlation relationships between the two scene images connected by the edge. When a scene image has no correlation relationship with other scene images, the server can delete the vertex corresponding to the scene image from the target graph structure.
[0115] Among them, the target graph structure can be expressed as G=(V, E). Among them, the vertex v i ∈V. c i represents the i-th scene image in the image set. The edge e ij ∈E. e ij represents the edge connecting the i-th scene image and the j-th scene image. At the same time, the existence of the edge e ij indicates that there is a matching relationship between the i-th scene image and the j-th scene image.
[0116] Step 3: According to the connection relationships of the target graph structure, determine the matching information corresponding to each edge in the target graph structure. The matching information includes the matching point pairs between the two scene images corresponding to the edge.
[0117] In this step, in the target graph structure, the two scene images with a connection relationship are the two scene images preliminarily determined to have matching features. The server can use the matching information to describe the matching relationship between the two scene images connected by an edge. The matching information can be represented by the matching point pairs corresponding to the matching features between the two scene images. For the convenience of description, the two scene images connected by an edge are respectively named the first image and the second image. Since there are matching features in the first image and the second image. Therefore, determine a feature point in the matching features as X. For example, the feature point X can be as Figure 4 shown.
[0118] As Figure 4 shown, O L is the center of the camera when taking the first picture. O R is the center of the camera when taking the second picture. X L is the mapping point of the feature point X in the first image. X R is the mapping point of the feature point X in the second image. The feature point X, the camera centers O L and OR can form a polar plane (XO L O R ). There can be a polar constraint for this polar plane: The feature point X of the first picture L The matching point in the second picture must be on the epipolar line X R e R . This epipolar line X R e R can be l R . Assume that the mapped point X L The pixel point coordinates in the first image are x1. The fundamental matrix F can be used to represent the mapping relationship from the pixel point coordinates x1 in the first picture to the epipolar line l R in the second picture. Its specific formula can be expressed as:
[0119] l R = Fx1
[0120] The mapped point X in the first picture L and the mapped point X in the second picture R are a pair of matching points. Assume that the pixel coordinate point corresponding to the mapped point X R in the second picture is x2. Then x2 must be on the epipolar line l R . Therefore, the above formula can be transformed into:
[0121]
[0122] When the pair of matching points (X L , X R ) is known, the server can estimate the fundamental matrix F based on the pair of matching points. Among them, the specific steps to implement the fundamental matrix estimation can include:
[0123] Step 3.1: According to the feature point extraction algorithm, extract the key points of the first picture from the points of the first picture, and, extract the key points of the second picture from the points of the second picture. The first picture and the second picture are two scene pictures connected by edges.
[0124] In this step, the server can use the feature point extraction algorithm to perform feature extraction on the first picture and the second picture respectively. Specifically, the server can input the first picture into the feature point extraction algorithm. After the calculation of the feature point extraction algorithm, the server can determine that some pixel points in the pixel points of the first picture are key points. And, the server can input the second picture into the feature point extraction algorithm. After the calculation of the feature point extraction algorithm, the server can determine that some pixel points in the pixel points of the second picture are key points.
[0125] In one implementation, the feature point extraction algorithm may be the Scale-invariant feature transform (SIFT) algorithm. The SIFT algorithm can calculate the SIFT features of each pixel point in the first image and the second image. The SIFT feature of each pixel point is used to describe the scale of the pixel point. The SIFT feature is 128-dimensional.
[0126] The server can calculate the scale of each pixel point based on the SIFT feature of each pixel point in the first image. The server can sort these pixel points according to the scale of each pixel point in the first image. The server can select the first preset number of pixel points with the largest scale from these pixel points as the key points of the first image. Among them, the first preset number can be a fixed number, such as 1000, 100, etc. Or, the first preset number can also be the first 20% of all the pixel points of the first image.
[0127] The server can calculate the scale of each pixel point based on the SIFT feature of each pixel point in the second image. The server can sort these pixel points according to the scale of each pixel point in the second image. The server can select the first preset number of pixel points with the largest scale from these pixel points as the key points of the second image. Among them, the first preset number can be a fixed number, such as 1000, 100, etc. Or, the first preset number can also be the first 20% of all the pixel points of the second image.
[0128] For example, when the first image and the second image each include 4000 pixel points, the server can determine the first 20% of the pixel points with the largest scale among them as the key points according to the scale of each pixel point. That is, 800 key points can be screened from the first image and the second image respectively.
[0129] Step 3.2: Determine the matching key point pairs in the first image and the second image according to the key points of the first image and the key points of the second image.
[0130] In this step, after the server determines the key points of the first image and the key points of the second image, it can pair these key points according to the epipolar constraint to obtain the key point pairs. For example, as Figure 4 shown, (X L , X R ) is a key point pair. For example, when the server screens 800 key points from the first image and the second image respectively, the server can determine that there are 200 key point pairs among them through the epipolar constraint.
[0131] Step 3.3. Estimate the fundamental matrix for indicating the transformation relationship between the first picture and the second picture according to the key point pairs.
[0132] In this step, the server can estimate the fundamental matrix based on these key point pairs to obtain the final fundamental matrix. This fundamental matrix can be used to indicate the transformation process between two points in the matching point pairs of the first picture and the second picture. Specifically, the process of the server estimating the fundamental matrix may include:
[0133] Step 3.3.1. The server randomly selects a second preset number of key point pairs from these key point pairs. The second preset number can be an empirical value. For example, 8.
[0134] Step 3.3.2. The server estimates the fundamental matrix using a preset algorithm according to the second preset number of key point pairs. The preset algorithm can be an existing algorithm or an improved algorithm. In this estimation process, an initial fundamental matrix can be set in the server. The server uses the second preset number of key point pairs to optimize the initial fundamental matrix to obtain an optimized fundamental matrix.
[0135] Step 3.3.3. The server calculates the matching relationship of these key point pairs using the fundamental matrix calculated in step S3.3.2 and determines the percentage of the matching key point pairs among these key point pairs. These key point pairs are actually the point pairs that have been determined to be matched. Therefore, the larger the percentage of the matching key point pairs calculated in step S3.3.2, the more accurate the fundamental matrix.
[0136] Step 3.3.3. The server determines whether the percentage of the matching key point pairs is greater than a preset ratio. When the percentage of the matching key point pairs is less than or equal to the preset ratio, the server can return to S3.3.1 to optimize the fundamental matrix. Since the selection process of these key point pairs is random and the number of key point pairs is usually much larger than the second preset number. Therefore, the probability that the key point pairs randomly selected again in step 3.3.1 are the same key point pairs is extremely small. For example, the server can randomly select 8 key point pairs from 200 key point pairs. In this case, step 3.3.2 will use the fundamental matrix calculated in the previous time as the initial matrix for the next iteration. When the percentage of the matching key point pairs is greater than the preset ratio, the server ends this iteration. The preset ratio is an empirical value. For example, 95%.
[0137] Step 3.4. Determine the matching information between the first picture and the second picture according to the points in the first picture, the points in the second picture, and the fundamental matrix. The matching information includes the matching point pairs between the first picture and the second picture and the number of matching point pairs.
[0138] In this step, there are matching features between the first image and the second image, and these matching features can be described in the form of matching point pairs. Among them, there are also non-matching features in the first image and the second image in addition to the matching features. Therefore, there are some pixel points in the first image that match some pixel points in the second image. At the same time, there are still some pixel points in the first image that do not match the remaining pixel points in the second image. And since the first preset number of pixel points in the first image have been matched in the above steps 3.1 and 3.2, in this step, the server only needs to match the remaining pixel points.
[0139] When the server completes the matching of all pixel points in the first image, the server can obtain multiple matching point pairs. The matching point pairs include the key point pairs determined in the above step 3.2. The server can store these matching points in the set M ij . M ij represents the set of matching point pairs between the i-th scene image and the j-th scene image. This set M ij can be used as the edge connecting the first image and the second image, that is, the edge e ij between the i-th scene image and the j-th scene image. The matching information can also include the weight w ij . w ij represents the number of matching point pairs between the i-th scene image and the j-th scene image. That is, the number of pairs in the set M ij . It can be expressed as w ij = |M ij |.
[0140] In one implementation, the server can obtain the remaining pixel points one by one through traversal. When the server obtains a pixel point in the first image as the point to be matched, the server can traverse all the pixel points in the second image and match the pixel points in the second image with the point to be matched one by one. The server can determine the target matching point in the second image corresponding to the point to be matched through this matching method. When a pixel point in the second image is matched with a pixel point in the first image to form a matching point pair, this pixel point in the second image will no longer be able to match with other pixel points in the first image. Therefore, in order to improve efficiency, after determining the point to be matched, the server can find the target matching point by matching the unmatched pixel points in the second image one by one.
[0141] In another implementation, after the server obtains a pixel point in the first image as the point to be matched, it uses the epipolar constraint to filter the pixel points in the second image. The server matches the point to be matched with the filtered pixel points one by one to determine the target matching point among them. In the construction of a large-scale 3D model, the first image and the second image are usually high-definition images. Using the epipolar constraint to filter the pixel points in the second image can effectively reduce the number of point pairs that need to be matched and calculated, improving the matching efficiency.
[0142] Among them, the process of the server filtering pixel points may include:
[0143] Assume that the point to be matched in the first image can be expressed as p = (x p , y p , 1), and the target matching point in the second image can be expressed as According to the epipolar constraint, the target matching point in the second image is on the epipolar line l q = (a q , b q, , c q, ). Therefore, the candidate point set of the target matching point q s can be expressed as:
[0144] C = {q|dist(q, l q ) ≤ d}
[0145] Among them, dist(q, l q ) is used to represent the distance from the candidate point q in this candidate point set to the epipolar line l q . That is, when the distance from the pixel point in the second image to the epipolar line is less than or equal to the first threshold d, it can be determined that this pixel point is a candidate point. Otherwise, this pixel point is not a candidate point. Among them, the first threshold d can be an empirical value. For example, the first threshold d is 0.1. Among them, the specific calculation formula of dist(q, l q ) can be as follows:
[0146]
[0147] Among them, the candidate point q = (x q , y q , 1).
[0148] In another implementation, after determining the points to be matched, the server can also screen the pixel points in the second image through hash matching. During the hash matching process, the server can allocate the original data to different buckets according to the hash values. Here, a bucket refers to the storage space corresponding to a hash value. A series of data stored in a bucket can be regarded as a data set. For example, in common hash algorithms, the server can calculate the hash value by means of taking the remainder, modulo, etc. Different hash values can correspond to different buckets. The same original data will necessarily correspond to the same hash value. The same hash value can correspond to different original data.
[0149] Among them, the specific process for the server to screen and match points can include:
[0150] Step 3.4.1: Coarsely match the points to be matched in the first image and the points in the second image according to the first preset hash algorithm and the fundamental matrix, and obtain the first set of points in the second image that match the points to be matched.
[0151] In this step, the server can calculate according to the point p to be matched and the fundamental matrix to obtain the mapped point p' after the point p to be matched is mapped to the second image. Among them, p'=(x p′ , y p′ , 1). The server can use the first preset hash algorithm to determine the hash values of the mapped point p' and each pixel point q in the second image. Among them, q=(x q , y q , 1). Among them, the first preset hash algorithm can be a locality-sensitive hashing algorithm. For example, the hash value of the mapped point p' can be 10100010. The hash value of a pixel point q in the second image can be 10100010. The server can calculate the Hamming distance between the hash value of the mapped point p' and the hash values of each pixel point q in the second image. For example, when there are 2000 pixel points in the second image, the server can calculate 2000 Hamming distances. The server can map each pixel in the second image to different buckets according to the Hamming distance. Each bucket corresponds to a Hamming distance.
[0152] In the most ideal state, when the mapped point p' matches the pixel point q, these two points coincide. Therefore, in the ideal state, the Hamming distance between the mapped point p' and the pixel point q is 0. Therefore, the server can select multiple buckets with hash values less than the second threshold according to the second threshold. The server can form the first set of points from the pixel points q in these multiple buckets. Among them, the second threshold is an empirical value. For example, the second threshold can be 3.
[0153] Step 3.4.2: According to the second preset hash algorithm and the base matrix, perform fine matching on the point to be matched and the points in the first point set to obtain a second point set that matches the point to be matched in the first point set, and the number of points in the second point set is less than the number of points in the first point set.
[0154] In this step, the server can use the second preset hash algorithm to calculate the hash values of the mapped point p' of the point to be matched p and each point q in the first point set. Among them, the second preset hash algorithm can be a locality-sensitive hashing algorithm. Since this hash algorithm will be used to more carefully screen out pixel points closer to the mapped point p' from the first point set. Therefore, the second preset hash algorithm in this step can use a longer hash value to increase the accuracy of the data. For example, the first preset hash algorithm can calculate an 8-bit hash value. The second preset hash algorithm can calculate a 16-bit hash value. For example, the hash value of the mapped point p' is 1001010001001010. The hash value of a pixel point q is 1001010001001011, and their Hamming distance is 1.
[0155] In the most ideal state, when the mapped point p' matches the point q in the first point set, these two points coincide. Therefore, in the ideal state, the distance between the mapped point p' and the point q in the first point set is 0. Therefore, the server can select multiple buckets with hash values less than the third threshold according to the third threshold. The server can form the points q in the first point set in these multiple buckets into a second point set. Among them, the third threshold is an empirical value. For example, the third threshold can be 3.
[0156] Step 3.4.3: According to the preset matching algorithm and the base matrix, perform one-by-one matching on the point to be matched and the points in the second point set to determine the target matching point that matches the point to be matched, and the target matching point is a point in the second point set.
[0157] In this step, the server can use the preset matching algorithm to perform one-by-one matching on the point to be matched and the points in the second point set. When a point in the second point set is successfully matched with the point to be matched, determine this point as the target matching point. When none of the points in the second point set can be successfully matched with the point to be matched, determine that there is no target matching point for the point to be matched. Among them, the preset matching algorithm can be an existing algorithm or an improved algorithm.
[0158] S103: Construct a three-dimensional model of the target scene according to the target graph structure.
[0159] In this embodiment, after the calculation of the target graph structure is completed, the server can construct a three-dimensional model of the target scene according to the target graph structure. Specifically, the server can sequentially obtain the edges in the target graph structure. After the server obtains an edge in the target graph structure, the server can splice the two scene pictures corresponding to the edge together and construct a three-dimensional model of the two scene pictures. After the three-dimensional models of the two pictures are constructed, the server can also add the three-dimensional models constructed from the two pictures to the three-dimensional model of the target scene picture to gradually improve the three-dimensional model of the target scene.
[0160] Among them, after the server obtains an edge in the target graph structure, the specific process of the server constructing the three-dimensional model may include:
[0161] Step 1: Obtain an edge in the target graph structure and the two scene pictures corresponding to the edge.
[0162] Step 2: Construct a three-dimensional model of the two scene pictures according to the camera internal parameters, the fundamental matrix, and the matching information of the edge.
[0163] In the three-dimensional model construction method provided by this application, the server obtains a picture set from a drone or other terminal devices. The server can determine the vertices in the target graph structure according to the scene pictures in the picture set. The server can also determine the association relationship between two scene pictures according to the picture information of each scene picture in the picture set. When there is an association relationship between two scene pictures, the server can add an edge for the two scene pictures in the target graph structure to obtain the target graph structure. The server can also determine the matching information between the two scene pictures according to the two scene pictures connected by the edge. The server constructs a three-dimensional model of the target scene according to the target graph structure. In this application, by constructing the target graph structure, the rapid construction of the three-dimensional model of the target scene is realized. This application also greatly improves the calculation efficiency of the matching point pairs between two scene pictures by using the hash algorithm, and improves the construction efficiency of the three-dimensional model.
[0164] Figure 5 The flowchart of a three-dimensional model construction method provided by an embodiment of this application is shown. On the basis of the above-mentioned embodiment in the figure, this embodiment can also implement the steps of constructing a three-dimensional model of the target scene in a distributed server cluster to improve the construction efficiency of the three-dimensional model in a large-scale scene. As Figure 5 shown, with the main server of the distributed server as the execution entity, the process of constructing a three-dimensional model of the target scene in this embodiment may include the following steps:
[0165] S201: The main server obtains the target graph structure.
[0166] In this embodiment, the main server may beFigure 2 The server in the illustrated embodiment. The main server can directly read the target graph structure from the memory. Or, the main server and Figure 2 the servers in the illustrated embodiment are different servers. The main server can obtain the target graph structure from Figure 2 the servers in the illustrated embodiment.
[0167] S202. The main server randomly reads an unread edge from the target graph structure. According to the task execution situation of the sub - servers, the main server sends this edge to a sub - server.
[0168] In this embodiment, the main server obtains the task processing situation of each sub - server. Among them, a task is used to indicate the construction situation of the three - dimensional model fragments of two scene pictures corresponding to an edge. When the sub - server completes the construction of the three - dimensional model fragments of two scene pictures corresponding to an edge and sends the constructed three - dimensional model fragments to the main server, the task is completed. After the sub - server processes a task, the main server can obtain an unread edge again and send this edge to the sub - server.
[0169] In one implementation, when the main server connects to the sub - server for the first time, it sends a third preset number of edges to the sub - server. In subsequent processing, the main server can determine whether to send new edges to the sub - server according to the task completion situation of the sub - server. Among them, the third preset number is a positive integer.
[0170] In one implementation, when the third preset number is 1, each sub - server only constructs the three - dimensional model fragment corresponding to the currently obtained edge. When the sub - server completes the construction of a three - dimensional model fragment and feeds it back to the main server, the sub - server enters the idle state. The sub - server needs to wait to receive a new edge sent by the main server and then enters the working state again.
[0171] In another implementation, when the third preset number is greater than 1, each sub - server can store multiple edges. For example, when the third preset number is 3, the sub - server stores 3 edges. Among them, the sub - server can process these three edges in order according to the writing order. When the sub - server completes the construction of a three - dimensional model fragment and feeds it back to the main server, the sub - server can obtain the next edge from the main server so that the number of edges to be processed in the sub - server remains 2, and there is always one edge in the processing state.
[0172] S203. The sub - server determines the groups where the two scene pictures are located according to the edge sent by the main server and writes these two groups into the memory.
[0173] In this embodiment, all the scene pictures in the picture set are stored in the disks of each sub-server. The pictures in the picture set are divided into m groups and stored in the disks of the sub-servers. Each group includes n scene pictures, and each scene picture is a block. The storage of the scene picture in the hard disk of the sub-server can be as Figure 6 shown. Among them, in addition to storing the scene picture, the hard disk can also store the sift features of each pixel point in the scene picture and picture information.
[0174] In one implementation, when the sub-server obtains the edge sent by the main server, it reads the two scene pictures corresponding to the edge from the disk.
[0175] Among them, each scene picture can correspond to a number. The sub-server can determine the groups where the two scene pictures are located according to this number. For example, as Figure 6 shown, the two scene pictures corresponding to the edge are located in group i and group j respectively. The sub-server reads group i and group j from the disk into the memory.
[0176] The GPU of the sub-server can read block a and block b from group i and group j in the memory respectively. Block a and block b respectively correspond to the two scene pictures of the edge. The GPU constructs a three-dimensional model segment of the two scene pictures according to the matching information stored in the edge.
[0177] In another implementation, as Figure 7 shown, the sub-server can also obtain the groups where the scene pictures corresponding to the edge are located from other sub-servers. Specifically, taking sub-server A and sub-server B as an example, the process can include:
[0178] Step 1: When sub-server A determines the picture numbers of the two scene pictures according to the edge, sub-server A can determine the groups where the two scene pictures are located. For example, the two scene pictures are in group i and group j respectively.
[0179] Step 2: Sub-server A sends a query request to the main server. The query request is used to request the main server to query whether group i and group j have been read into the memory by other sub-servers.
[0180] Step 3: The main server receives the query request sent by sub-server A and conducts a query. For example, as Figure 7As shown, the main server includes a table for marking group numbers and positions. This table is used to indicate that a certain group has been read by a certain child node. For example, the first row in this table is used to illustrate that group 11 has been read into the memory of child server C. The second row is used to illustrate that group 127 has been read into the memory of child server D. The fourth row is used to illustrate that group 239 has not been read into the memory of the child server yet. The main server can query from this table whether group i and group j have been read into the memory of a certain child server.
[0181] Step 3: The main server feeds back the result to child server A. According to this result, child server A obtains group i and group j. For example, when child server A determines that no other child servers have read group i into the memory, the child server can read this group i from the disk. When child server A determines that other child server B has read group j into the memory, child server A requests child server B to obtain this group j. Since the reading efficiency from the disk is lower than the data transmission efficiency between child servers, when child server B has read group j into the memory, the efficiency of sending group j from the memory of child server B to the memory of child server A is higher than the efficiency of child server A directly reading this group j from the disk.
[0182] Step 4: Child server A uploads the information that its memory stores group i and group j to the main server. The main server records child server A corresponding to group i and group j in the table.
[0183] S204: The GPU of the child server reads block a and block b from group i and group j in the memory. The child server constructs a three-dimensional model fragment according to block a and block b. The child server sends this three-dimensional model fragment to the main server.
[0184] S205: The main server obtains the three-dimensional model fragments processed by the child server and splices these three-dimensional model fragments.
[0185] In this step, the main server can obtain from each child server the three-dimensional model fragments spliced by the child server according to the edges in the data block information and the corresponding scene pictures. The main server can add this three-dimensional model fragment to the three-dimensional model of the target scene according to the position of this data block in the target graph structure.
[0186] In one implementation, when the main server splices multiple three-dimensional model fragments together, it can also use the method based on Similarity Transformation to splice the three-dimensional model of the target scene obtained from these three-dimensional model fragments. For example, when one three-dimensional model fragment includes scene pictures A and C, and another three-dimensional model fragment includes C and D, the main server can use similarity transformation to splice the feature points corresponding to scene picture C to obtain a three-dimensional model fragment including A, C, and D.
[0187] The three-dimensional model construction method provided by the present application is that the main server obtains the target graph structure. The main server can divide the target graph structure into multiple data blocks according to the data processing capabilities of the sub-servers. The main server randomly reads a data block that has not been read from the target graph structure. The main server can send the data block information corresponding to the data block to the corresponding sub-server according to the task processing status of multiple sub-servers. The main server obtains the three-dimensional model fragments processed by each sub-server. The three-dimensional model fragment includes part of the three-dimensional model of the target scene. The main server can add the three-dimensional model fragment to the three-dimensional model of the target scene according to the position of the data block in the target graph structure. In the present application, by using distributed computing, the construction process of the three-dimensional model is distributed to each sub-server for calculation. The main server only needs to combine the three-dimensional model segments constructed by the sub-servers to obtain the three-dimensional model of the target scene, which greatly improves the calculation efficiency and the construction efficiency of the three-dimensional model.
[0188] Figure 8 The flowchart of a three-dimensional model construction method provided by an embodiment of the present application is shown. Based on the above embodiment, this embodiment can also realize the calculation of matching information in the target graph structure in a distributed server cluster to improve the construction efficiency of the three-dimensional model in large-scale scenes. Figure 8 As shown, with the main server of the distributed server as the execution subject, the process of constructing the three-dimensional model of the target scene in this embodiment may include the following steps:
[0189] S301. The main server obtains a target graph structure including only association relationships.
[0190] In this embodiment, the target graph structure is Figure 2 The target graph structure shown in step 2 of S102 in FIG. The target graph structure does not include matching information for indicating the matching relationship between the two scene images corresponding to the edge. Each vertex is a scene image in the image set. Each edge is used to indicate that there is a relationship between the two scene images corresponding to the edge.
[0191] The main server can be Figure 2 The main server can directly read the target graph structure from the memory. Alternatively, the main server and Figure 2 The servers in the illustrated embodiment are different servers. The main server can be Figure 2 The target graph structure is obtained from the server of the illustrated embodiment.
[0192] S302: The primary server randomly reads a data block that has not been read from the target graph structure. The data block may include at least one edge.
[0193] In this embodiment, the main server may divide the target graph structure into multiple data blocks according to the data processing capabilities of the sub-servers. Each data block may include at least one edge. For example, a data block may include one edge, three edges, etc. The number of edges included in each data block is equal. Among them, when a data block includes multiple edges, these edges may be adjacent edges, or there may be a same fixed point between any two edges. For example, when the data block includes three edges, the three edges may be respectively connected to Graph A and Graph B, Graph A and Graph C, and Graph A and Graph D. Another example is that the three edges may be respectively connected to Graph A and Graph B, Graph B and Graph C, and Graph C and Graph D.
[0194] The server may use the method of labels to label the data blocks that have been read to distinguish the read data blocks from the unread data blocks. Alternatively, the server may also distinguish the read data blocks from the unread data blocks by storing the read data blocks and the unread data blocks in different storage areas. Alternatively, the server may distinguish the read data blocks from the unread data blocks by storing the indexes of the read data blocks and the indexes of the unread data blocks in two different storage areas.
[0195] S303. The main server may send the data block information corresponding to the data block to the corresponding sub-server according to the task processing conditions of multiple sub-servers.
[0196] In this embodiment, the main server obtains the task processing conditions of each sub-server. Among them, an arbitrary one is used to indicate the processing of a data block information. When the sub-server completes the processing of a data block information and sends the matching information corresponding to each edge in the data block obtained by calculation to the main server, the task is completed. When the sub-server finishes processing a task, the main server may send a new data block information to the sub-server. For example, when the main server connects to the sub-server for the first time, it sends the fourth preset number of data block information to the sub-server. In subsequent processing, the main server may determine whether to send new data block information to the sub-server according to the task completion situation of the sub-server. Among them, the fourth preset number is a positive integer.
[0197] In one implementation manner, when the fourth preset number is 1, each sub-server only stores the data block information currently being processed. When the sub-server completes the processing of a data block information and feeds back the matching information corresponding to each edge in the data block information obtained by calculation to the main server, the sub-server enters the idle state. The sub-server needs to wait to receive the new data block information sent by the main server and then enters the working state again.
[0198] In another implementation, when the fourth preset quantity is greater than 1, multiple data block information can be stored in each sub-server. For example, when the fourth preset quantity is 3, 3 data block information are stored in the sub-server. One of the data block information can be in a computing state, and the other two data block information are stored in the sub-server. After the data block information in the computing state is processed, the sub-server can read the next data block information for calculation according to the writing order. When the main server receives the matching information corresponding to each edge in the data block information sent by the sub-server, the main server can send a new data block information to the sub-server so that the number of data block information in the sub-server remains 3.
[0199] Among them, when the sub-server completes the calculation of the data in a data block information, the sub-server can delete the content corresponding to the data block information from the memory to avoid abnormal storage of the sub-server caused by excessive data volume. Similarly, when the sub-server sends the matching information corresponding to each edge in the data block information to the main server, the sub-server can delete the matching information corresponding to each edge in the data block information so that the sub-server retains more free memory space.
[0200] In one implementation, before sending the data block information to the sub-server, the main server can also obtain the scene pictures corresponding to the edges according to the edges in the data block. The main server can add these scene pictures to the database information so that these scene pictures can be sent to the sub-server together with the edges in the data block.
[0201] In another implementation, the main server sends the data block information including only the edges to the sub-server. When the sub-server reads the data block information from the memory, the sub-server requests the main server to obtain the scene picture information corresponding to the edge according to the edge in the data block information. When the data block information includes multiple edges, the sub-server needs to process these edges and the two scene pictures corresponding to the edges one by one. Therefore, when processing the first edge, the sub-server can request the main server for the two scene pictures corresponding to the second edge, thereby improving the processing efficiency of the second edge.
[0202] S304. The main server obtains the matching information corresponding to each edge in the data block information processed by each sub-server.
[0203] In this step, the main server can obtain from each sub-server the matching information corresponding to each edge in the data block information calculated by the sub-server according to the edge in the data block information and the scene picture corresponding to the edge. The main server can add the matching information of each edge to the edge.
[0204] The 3D model construction method provided by this application is such that the main server obtains a target graph structure that only includes association relationships. The target graph structure does not include matching information for indicating the matching relationship between two scene pictures corresponding to an edge. The main server can divide the target graph structure into multiple data blocks according to the data processing capabilities of the sub-servers. The main server randomly reads an unread data block from the target graph structure. The main server can send the data block information corresponding to the data block to the corresponding sub-server according to the task processing conditions of multiple sub-servers. The main server obtains the matching information corresponding to each edge in the data block information processed by each sub-server. The main server can add the matching information of each edge to the edge. In this application, by using distributed computing, the matching information between two scene pictures corresponding to an edge in the scene picture is distributed to each sub-server for calculation, greatly improving the calculation efficiency and the construction efficiency of the 3D model.
[0205] Figure 9 FIG. shows a schematic structural diagram of a 3D model construction device provided by an embodiment of this application. As Figure 9 shown, the 3D model construction device 10 in this embodiment is used to implement the operations corresponding to the server in any of the above method embodiments. The 3D model construction device 10 in this embodiment further includes:
[0206] An acquisition module 11, configured to acquire a picture set, where the picture set includes at least one scene picture of a target scene and the picture information of each scene picture.
[0207] A processing module 12, configured to determine a target graph structure according to the scene pictures in the picture set and the picture information of each scene picture. The vertices in the target graph structure are the scene pictures in the picture set, and the edges in the target graph structure are used to indicate the matching information between two scene pictures connected by the edge. Construct a 3D model of the target scene according to the target graph structure.
[0208] In an example, the processing module 12 is specifically configured to determine the association relationship between scene pictures in the picture set according to the shooting coordinates of each scene picture in the picture set. Construct the connection relationship of the target graph structure according to the association relationship between scene pictures in the picture set. The edges in the target graph structure are used to indicate that there is an association relationship between two scene pictures connected by the edge. Determine the matching information corresponding to each edge in the target graph according to the connection relationship of the target graph structure. The matching information includes the matching point pairs between two scene pictures corresponding to the edge.
[0209] In one example, the processing module 12 is specifically configured to extract the key points of the first picture from the points of the first picture and the key points of the second picture from the points of the second picture according to the feature point extraction algorithm. The first picture and the second picture are two scene pictures connected by sides. According to the key points of the first picture and the key points of the second picture, determine the matching key point pairs in the first picture and the second picture. According to the key point pairs, estimate the fundamental matrix for indicating the transformation relationship between the first picture and the second picture. According to the points in the first picture, the points in the second picture, and the fundamental matrix, determine the matching information between the first picture and the second picture. The matching information includes the matching point pairs between the first picture and the second picture and the number of matching point pairs.
[0210] In one example, the processing module 12 is specifically configured to perform rough matching on the points to be matched in the first picture and the points in the second picture according to the first preset hashing algorithm and the fundamental matrix, and obtain the first point set that matches the points to be matched in the second picture. According to the second preset hashing algorithm and the fundamental matrix, perform fine matching on the points to be matched and the points in the first point set, and obtain the second point set that matches the points to be matched in the first point set. The number of points in the second point set is less than the number of points in the first point set. According to the preset matching algorithm and the fundamental matrix, perform one-by-one matching on the points to be matched and the points in the second point set, and determine the target matching point that matches the points to be matched. The target matching point is the point in the second point set.
[0211] In one example, the processing module 12 is specifically configured to determine the distance between each pair of scene pictures according to the shooting coordinates of each scene picture in the picture set. When the distance between each pair of scene pictures is less than or equal to the distance threshold, determine that there is an association relationship between each pair of scene pictures.
[0212] In one example, the processing module 12 is specifically configured to obtain an edge in the target graph structure and the two scene pictures corresponding to the edge. According to the camera internal parameters, the fundamental matrix, and the matching information of the edge, construct the three-dimensional models of the two scene pictures.
[0213] The three-dimensional model construction device 10 provided in the embodiments of the present application can execute the above method embodiments. For the specific implementation principles and technical effects, reference can be made to the above method embodiments, which will not be elaborated here in this embodiment.
[0214] Figure 10 The hardware structure diagram of a server provided in the embodiments of the present application is shown. As Figure 10 shown, the server 20 is used to implement the operations corresponding to the server in any of the above method embodiments. The server 20 in this embodiment may include: a memory 21, a processor 22, and a communication interface 24.
[0215] A memory 21 for storing a computer program. A processor 22 for executing the computer program stored in the memory to implement the three-dimensional model construction method in the above embodiments. Optionally, the memory 21 can be either independent or integrated with the processor 22.
[0216] When the memory 21 is a device independent of the processor 22, the electronic device 20 may further include a bus 23. The bus 23 is used to connect the memory 21 and the processor 22.
[0217] The communication interface 24 can be connected to the processor 21 through the bus 23. The communication interface 24 is used to obtain a picture set. The picture set can be directly uploaded to the server 20 by the communication interface 24 after being taken by a drone. Or, the picture set can also be obtained by other terminal devices after organizing the pictures taken by the camera and then uploaded to the server 20. The communication interface 24 is also used to output the three-dimensional model of the constructed target scene. The processor 21 can output the three-dimensional model to the display of the server 20 through the communication interface 24. Or the processor 21 can also output the three-dimensional model to other terminal devices through the communication interface 24.
[0218] The server provided in this embodiment can be used to execute the above three-dimensional model construction method, and its implementation manner and technical effects are similar, which will not be elaborated here in this embodiment.
[0219] Figure 11 Shows a schematic structural diagram of a distributed server cluster provided by an embodiment of the present application. As Figure 11 shown, the distributed server cluster 30 is used to implement the operations corresponding to the distributed server cluster in any of the above method embodiments. The distributed server cluster 30 in this embodiment may include: a main server 31 and at least one sub-server 32.
[0220] The main server 31 is the server as Figure 10 shown. The main server 31 is used to divide the target graph structure into multiple data blocks according to the embodiments as Figure 7 and Figure 8 shown, and distribute the data blocks to the sub-servers for calculation.
[0221] The sub-server 32 is used to receive the data blocks sent by the main server and calculate the information of each data block according to a preset algorithm. The sub-server 32 is also used to feedback the calculation results to the main server so that the main server can merge the calculation results.
[0222] The server provided in this embodiment can be used to execute the above three-dimensional model construction method, and its implementation manner and technical effects are similar, which will not be elaborated here in this embodiment.
[0223] The present application also provides a computer-readable storage medium, in which a computer program is stored. When the computer program is executed by a processor, it is used to implement the methods provided by the above various embodiments.
[0224] The present application also provides a program product, which includes execution instructions stored in a computer-readable storage medium. At least one processor of the device can read the execution instructions from the computer-readable storage medium, and the execution of the execution instructions by at least one processor causes the device to implement the methods provided by the above various embodiments.
[0225] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the various embodiments of the present application.
Claims
1. A three-dimensional model construction method, characterized in that, The method includes: Obtaining a set of pictures, where the set of pictures includes at least one scene picture of a target scene and picture information of each of the scene pictures; Determining a target graph structure according to the scene pictures in the set of pictures and the picture information of each of the scene pictures, where the vertices in the target graph structure are the scene pictures in the set of pictures, and the edges in the target graph structure are used to indicate the matching information between two scene pictures connected by the edge; Constructing a three-dimensional model of the target scene according to the target graph structure; The picture information includes shooting coordinates when shooting the target scene. Determining the target graph structure according to the scene pictures in the set of pictures and the picture information of each of the scene pictures includes: Determining the distance between every two of the scene pictures according to the shooting coordinates of each of the scene pictures in the set of pictures; When the distance between every two of the scene pictures is less than or equal to a distance threshold, determining that there is an association relationship between every two of the scene pictures; Wherein, the distance threshold is determined according to the following formula: Among them, is the distance threshold for each scene picture, is the maximum distance between each scene picture and every other scene picture in the picture set, is the minimum distance between each scene picture and every other scene picture in the picture set, and the distance threshold subscript of indicates the distance threshold of the th scene picture in the picture set, is the number of pictures in the picture set, scale constant, ; Constructing the connection relationship of the target graph structure according to the association relationship between every two of the scene pictures in the set of pictures, and the edges in the target graph structure are used to indicate that there is an association relationship between two scene pictures connected by the edge; Determining the matching information corresponding to each edge in the target graph structure according to the connection relationship of the target graph structure, where the matching information includes the matching point pairs between the two scene pictures corresponding to the edge; Determining the matching information corresponding to each edge in the target graph structure according to the connection relationship of the target graph structure includes: Extracting key points of a first picture from the points of the first picture and extracting key points of a second picture from the points of the second picture according to a feature point extraction algorithm, where the first picture and the second picture are two scene pictures connected by the edge; Determining the matching key point pairs in the first picture and the second picture according to the key points of the first picture and the key points of the second picture; Estimating a fundamental matrix for indicating the transformation relationship between the first picture and the second picture according to the key point pairs; Determining the matching information between the first picture and the second picture according to the points in the first picture, the points in the second picture and the fundamental matrix, where the matching information includes the matching point pairs between the first picture and the second picture and the number of the matching point pairs; Determining the matching information between the first picture and the second picture according to the points in the first picture, the points in the second picture and the fundamental matrix includes: After obtaining a pixel point in the first picture as a point to be matched, by traversing all pixel points in the second picture, matching each pixel point in the second picture with the point to be matched one by one to determine the target matching point in the second picture corresponding to the point to be matched.
2. The method according to claim 1, wherein Determining the matching information between the first picture and the second picture according to the points in the first picture, the points in the second picture and the fundamental matrix includes: Coarsely match the points to be matched in the first picture and the points in the second picture according to the first preset hash algorithm and the base matrix, and obtain a first set of points in the second picture that match the points to be matched; Fine-match the points to be matched and the points in the first set of points according to the second preset hash algorithm and the base matrix, and obtain a second set of points in the first set of points that match the points to be matched, where the number of points in the second set of points is less than the number of points in the first set of points; Match the points to be matched and the points in the second set of points one by one according to the preset matching algorithm and the base matrix, and determine a target matching point that matches the points to be matched, where the target matching point is a point in the second set of points.
3. The method according to claim 1 or 2, characterized in that, The picture information includes the camera internal parameters when the target scene is photographed. The constructing of the three-dimensional model of the target scene according to the target graph structure includes: Obtain an edge in the target graph structure and two scene pictures corresponding to the edge; Construct the three-dimensional models of the two scene pictures according to the camera internal parameters, the base matrix, and the matching information of the edge.
4. A three-dimensional model construction device, characterized in that, The device includes: A sending module, configured to obtain a picture set, where the picture set includes at least one scene picture of the target scene and the picture information of each scene picture; A processing module, configured to determine a target graph structure according to the scene pictures in the picture set and the picture information of each scene picture, where the vertices in the target graph structure are the scene pictures in the picture set, and the edges in the target graph structure are used to indicate the matching information between the two scene pictures connected by the edge; construct the three-dimensional model of the target scene according to the target graph structure; Specifically, the processing module is configured to determine the distance between every two of the scene pictures according to the shooting coordinates of each scene picture in the picture set; When the distance between every two of the scene pictures is less than or equal to a distance threshold, determine that there is an association relationship between every two of the scene pictures; Wherein, the distance threshold is determined according to the following formula: Among them, is the distance threshold for each scene picture, is the maximum distance between each scene picture and every other scene picture in the picture set, is the minimum distance between each scene picture and every other scene picture in the picture set, and the distance threshold subscript of represents the distance threshold of the th scene picture in the picture set, is the number of pictures in the picture set, is the scale constant, ; Construct the connection relationship of the target graph structure according to the association relationship between every two of the scene pictures in the picture set, and the edges in the target graph structure are used to indicate that there is an association relationship between the two scene pictures connected by the edge; Determine the matching information corresponding to each edge in the target graph structure according to the connection relationship of the target graph structure, where the matching information includes the matching point pairs between the two scene pictures corresponding to the edge; Specifically, the processing module is configured to extract the key points of the first picture from the points of the first picture and extract the key points of the second picture from the points of the second picture according to the feature point extraction algorithm, where the first picture and the second picture are two scene pictures connected by the edge; Determine the matching key point pairs in the first picture and the second picture according to the key points of the first picture and the key points of the second picture; Estimate the base matrix for indicating the transformation relationship between the first picture and the second picture according to the key point pairs. Determine the matching information between the first picture and the second picture according to the points in the first picture, the points in the second picture, and the fundamental matrix, where the matching information includes the matching point pairs between the first picture and the second picture and the number of the matching point pairs; The processing module is specifically configured to, after obtaining a pixel point in the first picture as a point to be matched, match the pixel points in the second picture with the point to be matched one by one by traversing all the pixel points in the second picture, and determine the target matching point in the second picture corresponding to the point to be matched.
5. A server, characterized in that, The server includes: a memory, a processor; The memory is used to store a computer program; the processor is used to implement the three-dimensional model construction method according to any one of claims 1 to 3 based on the computer program stored in the memory.
6. A computer-readable storage medium, characterized in that, A computer program is stored in the computer-readable storage medium, and when the computer program is executed by a processor, it is used to implement the three-dimensional model construction method according to any one of claims 1 to 3.
7. A computer program product, characterized in that, The computer program product includes a computer program, and when the computer program is executed by a processor, it implements the three-dimensional model construction method according to any one of claims 1 to 3.
Citation Information
Patent Citations
Three-dimensional scene construction method and device, storage medium and electronic equipment
CN112270755A