AI-assisted scene white model mapping method and device, and storage medium
By using an AI-assisted white model mapping method, aerial imagery and sparse reconstruction technology are employed to generate multi-angle views, solving the consistency and detail issues of traditional white model mapping and achieving efficient and accurate texture mapping.
Patent Information
- Application Number
- CN202511460838.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-14
- Publication Date
- 2026-02-03
- Estimated Expiration
- 2045-10-14
AI Technical Summary
Traditional white model texturing processes lack consistency and detail, are inefficient, and are easily affected by subjective human operation, resulting in large errors.
Using an AI-assisted approach, aerial images are acquired to establish a real-world coordinate system. Sparse reconstruction and coordinate transformation are then performed to train an AI model that generates multi-angle views. The best view is then selected for texture mapping.
It improves the efficiency and accuracy of white model texture mapping, reduces human subjective error, and ensures the consistency and fineness of texture mapping.
Smart Images

Figure CN120931802B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, and in particular to an AI-assisted method, apparatus and storage medium for scene white model mapping. Background Technology
[0002] With the continuous advancement of urban digitalization and intelligentization, the importance of 3D scenes in fields such as urban planning, geographic information systems, and virtual reality is becoming increasingly prominent. In these applications, large-scale scene white models are gradually becoming a key tool due to their efficient and realistic representation capabilities.
[0003] White models, also known as untextured models, have their texturing process directly impacting the realism and usability of the final 3D model. Traditional white model texturing techniques rely heavily on manual operation, requiring the capture of numerous images from different perspectives and the stitching of these images onto the white model using complex algorithms. This process is not only time-consuming and labor-intensive, but also often results in inconsistent texture quality and a lack of detail due to the subjectivity of manual operation.
[0004] In stereo photogrammetry, aerial triangulation is the commonly used method. However, when dealing with large-scale scenes, the process of texturing the white model after aerial triangulation is often limited by computational resources, algorithm efficiency, and the algorithm itself, making it difficult to achieve high-quality textures. Furthermore, relying solely on aerial triangulation to obtain aerial images and then constructing texture maps results in textures that lack consistency and detail. Additionally, texture map construction is inefficient and subject to significant subjective errors from manual operations. Summary of the Invention
[0005] This application provides an AI-assisted scene white model mapping method, apparatus, and storage medium, which solves the technical problems in the prior art where the white model mapping process lacks consistency and detail, and the white model mapping process is inefficient and easily affected by subjective human operation, thus improving the efficiency and accuracy of the white model mapping process, reducing errors caused by subjective human operation, and ensuring the consistency and detail of texture mapping.
[0006] In a first aspect, embodiments of this application provide an AI-assisted method for scene white model texturing, comprising: acquiring aerial images of white model units within a scanned area and establishing a real scene coordinate system for the scanned area; performing sparse reconstruction on the aerial images to obtain sparse reconstruction results, and inputting the sparse reconstruction results into the real scene coordinate system to obtain coordinates of the white model units; training an AI model based on the coordinates of the white model units to obtain a virtual view model; constructing a viewpoint for the white model units and generating multi-angle views of the white model units based on the virtual view model; selecting the best view based on the multi-angle views, calculating the texture coordinates of the white model units, determining the maximum texture region, and texturing the white model units in the best view based on the maximum texture region.
[0007] In conjunction with the first aspect, in one possible implementation, the sparse reconstruction of aerial images to obtain sparse reconstruction results includes: using aerial triangulation technology to sparsely reconstruct the aerial images to determine the relative pose relationships between each aerial image, thereby obtaining sparse reconstruction results; the sparse reconstruction results are presented in coordinate form.
[0008] In conjunction with the first possible implementation of the first aspect, in the second possible implementation, the step of inputting the sparse reconstruction result into the real scene coordinate system to obtain the coordinates of the white model unit includes: acquiring GPS information based on aerial imagery and converting the GPS information into projected coordinates; establishing a sparse reconstruction coordinate system using aerial triangulation technology and corresponding the sparse reconstruction result with the projected coordinates to obtain a transformation matrix between the sparse reconstruction coordinate system and the real scene coordinate system; inputting the sparse reconstruction result into the real scene coordinate system and converting the sparse reconstruction result into the coordinates of the white model unit based on the transformation matrix; the transformation matrix includes: Where s represents the scale factor, R represents the rotation matrix, and T represents the translation vector.
[0009] In conjunction with the second possible implementation of the first aspect, in the third possible implementation, after inputting the sparse reconstruction result into the real scene coordinate system to obtain the coordinates of the white model unit, the method further includes: obtaining the back projection error of the white model unit; obtaining the back projection error of the white model unit based on the coordinates of the white model unit and the projection coordinates includes: obtaining multiple 3D sparse point clouds based on the sparse reconstruction result, inputting the 3D sparse point clouds into the projection coordinates to obtain multiple projection positions; converting the 3D sparse point clouds into pixel coordinates based on intrinsic and extrinsic parameters; calculating the error of the 3D sparse point clouds based on the projection positions and pixel coordinates; calculating the average of the overall error of the 3D sparse point clouds to obtain the back projection error of the white model unit; the calculation method of the error of the 3D sparse point clouds includes: ;in, The x-coordinate represents the projection position. The x-coordinate of the pixel coordinate. The ordinate represents the projected position. The ordinate represents the pixel coordinate; the method for calculating the mean of the overall error of the 3D sparse point cloud includes: Where m represents the number of 3D sparse point clouds.
[0010] In conjunction with the third possible implementation of the first aspect, in the fourth possible implementation, the step of training the AI model based on the coordinates of the white model unit to obtain the virtual view model includes: setting an error threshold for the coordinates of the white model unit, and filtering out back projection errors based on the error threshold to ensure the stability of the model, thereby obtaining the filtered coordinates of the white model unit; training the AI model based on the filtered coordinates of the white model unit to obtain the virtual view model; the step of filtering out the back projection errors of the white model unit based on the error threshold includes: filtering out back projection errors with an error range greater than the error threshold, and retaining other back projection errors.
[0011] In conjunction with the first aspect, in the fifth possible implementation, the construction of the perspective of the white model unit and the generation of multi-angle views of the white model unit based on the virtual view model include: generating multiple virtual perspectives of the white model unit according to the geometric structure and specific location of the white model unit; inputting the virtual perspectives into the virtual view model to obtain multiple virtual views of the white model unit; and combining the multiple virtual views to obtain the multi-angle views of the white model unit.
[0012] In conjunction with the first aspect, in the sixth possible implementation, the step of selecting the best view based on multi-angle views, calculating the texture coordinates of the white model unit, and determining the texture region includes: obtaining multiple vertex back projections of the white model unit based on the relationship between the normal of the facet and the pose of the white model unit; inputting the vertex back projections of the facet to the corresponding views of the multi-angle views to obtain the texture coordinates of the white model unit; obtaining the initial texture region based on the position of the vertex back projections in the texture coordinates, and selecting the initial texture region with the largest back projection area based on the multi-angle views to obtain the texture region.
[0013] Secondly, embodiments of this application provide an apparatus for performing an AI-assisted scene white model texturing method, comprising: an image acquisition module for acquiring aerial images of a scanned area to obtain aerial images of a single white model within the scanned area, and recording the geographical location information of the aerial images; a data processing module for establishing a real scene coordinate system based on the geographical location information, and performing sparse reconstruction of the aerial images based on aerial triangulation technology to obtain sparse reconstruction results; a viewpoint generation module for training an AI model based on the sparse reconstruction results to obtain a virtual view model, and obtaining multi-angle views of the single white model based on the virtual view model, determining the best view among the multi-angle views, calculating texture coordinates, and determining the texture region of the single white model; and a texturing module for performing texture mapping on the single white model based on the texture region.
[0014] Thirdly, embodiments of this application provide an apparatus for executing an AI-assisted scene white model texture mapping method, the apparatus comprising: a processor; a memory for storing processor-executable instructions; wherein, when the processor executes the executable instructions, it implements the method as described in the first aspect or any possible implementation of the first aspect.
[0015] Fourthly, embodiments of this application provide a non-volatile computer-readable storage medium, the non-volatile computer-readable storage medium including storage for storing a computer program or instructions that, when executed, cause the method described in the first aspect or any possible implementation of the first aspect to be implemented.
[0016] One or more technical solutions provided in the embodiments of this application have at least the following technical effects or advantages:
[0017] This application embodiment employs a technique of acquiring large-scale aerial images within a scanning area, establishing a real scene coordinate system within that area, performing sparse reconstruction and coordinate transformation on the aerial images, training an AI model based on the sparse reconstruction results, obtaining multi-angle views of the white model unit using the trained model, combining texture coordinates with the multi-angle views to obtain texture regions, and then performing texture mapping on the white model unit based on the texture regions. This effectively solves the technical problem that the traditional white model mapping process is limited by computing resources, algorithm efficiency, and the algorithm itself, resulting in inconsistent and less detailed mapping results and low mapping efficiency. Thus, it achieves the technical effect of improving the efficiency and progress of the white model mapping process and reducing errors caused by subjective human operation. Attached Figure Description
[0018] To more clearly illustrate the technical solutions of the embodiments of this application, the drawings used in the description of the embodiments of this application or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0019] Figure 1 A flowchart of an AI-assisted scene white model texture mapping method provided in the embodiments of this application;
[0020] Figure 2 A flowchart illustrating the method for obtaining the coordinates of a white model unit as provided in this application embodiment;
[0021] Figure 3 This is a flowchart illustrating a method for obtaining the back projection error of a white model unit according to an embodiment of this application.
[0022] Figure 4This is a structural diagram of an AI-assisted scene white model texture mapping device provided in an embodiment of this application. Detailed Implementation
[0023] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.
[0024] Figure 1 This is a flowchart of an AI-assisted scene white model texture mapping method provided in the embodiments of this application, including steps S1 to S5.
[0025] Step S1: Acquire aerial images of the white model unit within the scanning area and establish a real scene coordinate system for the scanning area.
[0026] Step S2: Perform sparse reconstruction on the aerial imagery to obtain the sparse reconstruction result, and input the sparse reconstruction result into the real scene coordinate system to obtain the coordinates of the white model unit.
[0027] Step S3: Train the AI model based on the coordinates of the white model unit to obtain the virtual view model.
[0028] Step S4: Construct the perspective of the white model unit and generate multi-angle views of the white model unit based on the virtual view model.
[0029] Step S5: Select the best view based on the multi-angle view, calculate the texture coordinates of the white model unit, determine the maximum texture area, and apply texture to the white model unit in the best view based on the maximum texture area.
[0030] This application embodiment employs a technique of acquiring large-scale aerial images within a scanning area, establishing a real scene coordinate system within that area, performing sparse reconstruction and coordinate transformation on the aerial images, training an AI model based on the sparse reconstruction results, obtaining multi-angle views of the white model unit using the trained model, combining texture coordinates with the multi-angle views to obtain texture regions, and then performing texture mapping on the white model unit based on the texture regions. This effectively solves the technical problems of traditional white model mapping processes being limited by computing resources, algorithm efficiency, and the algorithm itself, resulting in inconsistent and less detailed mapping results and low mapping efficiency. Thus, it achieves the technical effect of improving the efficiency and progress of the white model mapping process and reducing errors caused by subjective human operation.
[0031] The textures of high-precision city white models vary, requiring texture data acquisition to construct realistic textures. The image acquisition method for large-scale scene white model mapping based on AI assistance provided in this application eliminates the tedious process of acquiring each individual white model, instead utilizing large-scale aerial imagery. This image acquisition method not only ensures that each individual white model is completely captured but also simplifies the workflow, shortens operation time, and improves the efficiency of white model image acquisition.
[0032] First, this application embodiment selects suitable drones and imaging equipment. High-resolution cameras are used to accurately capture high-precision images over a large area. By planning the drone's flight path, it is ensured that the entire target area is covered. Multi-flight, multi-angle oblique photography techniques are typically employed to obtain comprehensive aerial imagery.
[0033] During the flight, the drone flies along a predetermined route, acquiring multiple aerial images. Each image records corresponding GPS information to ensure the accuracy of subsequent processing.
[0034] In this embodiment of the application, the process of sparse reconstruction of aerial images to obtain sparse reconstruction results includes:
[0035] Aerial triangulation is used to sparsely reconstruct aerial images to determine the relative pose relationships between different aerial images and obtain sparse reconstruction results. Specifically, the sparse reconstruction results are presented in coordinate form.
[0036] For example, after importing aerial images into specialized software for processing, a sparse point cloud model can be constructed. This point cloud model is used to verify the results of aerial triangulation and lay the foundation for subsequent processing.
[0037] Secondly, the results obtained from aerial triangulation can also provide the intrinsic and extrinsic perspective parameters for each aerial image.
[0038] Figure 2 A flowchart illustrating the method for obtaining the coordinates of a single white model unit. (For example...) Figure 2 As shown, the sparse reconstruction results are input into the real scene coordinate system to obtain the coordinates of the white model unit. The specific process includes:
[0039] Step 201: Obtain GPS information based on aerial imagery and convert the GPS information into projected coordinates.
[0040] Step 202: Establish a sparse reconstruction coordinate system using aerial triangulation technology, and match the sparse reconstruction results with the projected coordinates to obtain the transformation matrix between the sparse reconstruction coordinate system and the real scene coordinate system.
[0041] Step 203: Input the sparse reconstruction results into the real scene coordinate system, and convert the sparse reconstruction results into white model unit coordinates based on the transformation matrix.
[0042] The transformation matrix is: , where s represents the scale factor, R represents the rotation matrix, and T represents the translation vector.
[0043] For example, the detailed process of mapping sparse reconstruction results to projected coordinates includes:
[0044] The transformation matrix has 7 degrees of freedom. When performing aerial triangulation, at least three point pairs (projected coordinates of the aerial image and sparse reconstruction results from the aerial triangulation) need to be calculated.
[0045] However, to prevent large errors during the solution process, this embodiment uses the least squares method to construct constraints far exceeding three point pairs, finding the transformation matrix with the minimum error. This matrix is the transformation matrix provided in this embodiment. .
[0046] Specifically, the coordinates (x, y) in the projected coordinate system (converting the GPS information from the image acquisition process, i.e., latitude and longitude, into a coordinate system, which is the projected coordinate system) are obtained using the GPS information of each aerial image. Here, x represents the abscissa of each aerial image in the projected coordinate system, and y represents the ordinate of each aerial image in the projected coordinate system. A one-to-one correspondence is established between the projected coordinates (x, y) of each aerial image and the flight altitude z, and the sparse reconstructed coordinate system obtained after aerial triangulation. The transformation matrix with the minimum error is obtained based on the least squares method.
[0047] Based on the transformation matrix, the sparse reconstruction result is converted into white model unit coordinates, which are equivalent to the coordinates of the real scene established based on the real scene coordinate system.
[0048] This transformation process ensures that the textures generated in the virtual viewpoint can be accurately matched to each individual white model, making the subsequent texture mapping process more accurate and efficient. This step guarantees an accurate correspondence between the model and the actual scene through mathematical and geometric transformations, enabling the texture construction of the white model to achieve high precision and realism.
[0049] Figure 3 This is a flowchart of a method for obtaining the back projection error of a white model unit according to an embodiment of this application. After obtaining the coordinates of the white model unit, the method further includes obtaining the back projection error of the white model unit, as detailed below. Figure 3 As shown, it includes:
[0050] Step 204: Based on the sparse reconstruction results, obtain multiple 3D sparse point clouds, input the 3D sparse point clouds into the projection coordinates, and obtain multiple projection positions.
[0051] Step 205: Convert the 3D sparse point cloud into pixel coordinates based on intrinsic and extrinsic parameters. These intrinsic and extrinsic parameters are obtained through aerial triangulation.
[0052] Step 206: Calculate the error of the 3D sparse point cloud based on the projection position and pixel coordinates.
[0053] Step 207: Calculate the mean of the errors of all 3D sparse point clouds to obtain the back projection error of the white model unit.
[0054] The calculation method for the error of the 3D sparse point cloud includes:
[0055] .
[0056] in, The x-coordinate represents the projection position. The x-coordinate of the pixel coordinate. The ordinate represents the projected position. The vertical coordinate represents the pixel coordinate.
[0057] The calculation methods for the mean error of all 3D sparse point clouds include:
[0058] .
[0059] Where m represents the number of 3D sparse point clouds.
[0060] In this embodiment of the application, an AI model is trained based on the coordinates of the white model unit to obtain a virtual view model, specifically including:
[0061] An error threshold is set for the coordinates of individual white model units, and backprojection errors are filtered out based on this threshold to ensure model stability. The specific filtering process includes: filtering out backprojection errors with an error range greater than the error threshold, and retaining other backprojection errors. The resulting white model unit coordinates are then obtained.
[0062] The AI model is trained based on the coordinates of the filtered white model units to obtain the virtual view model.
[0063] It should be noted that the AI model can be selected from models such as NeRF and 3D Gaussian Splatting. In this embodiment, the NeRF model is selected as the basic model for implementing the training process. The NeRF model is a technique for synthesizing new views by implicitly representing the radiation field of a scene. It can achieve significant results in high-quality view synthesis. By training on a large number of aerial images, the NeRF model can generate high-quality virtual views from any angle.
[0064] Specifically, the NeRF model maps the input aerial imagery into a high-dimensional neural network and generates new views by optimizing the network parameters.
[0065] In this embodiment of the application, a perspective construction is performed on the white model unit, and multi-angle views of the white model unit are generated based on the virtual view model, specifically including:
[0066] Based on the geometry and specific location of the white model unit, multiple virtual perspectives are generated for the white model unit. These virtual perspectives should cover the main appearance features of the white model unit to ensure that detailed texture information of the white model can be captured from different angles.
[0067] Input the virtual perspective into the virtual view model to obtain multiple virtual views of the white model unit.
[0068] By combining multiple virtual views, a multi-angle view of the white model unit can be obtained.
[0069] For example, in the embodiments of this application, two ideas for designing virtual perspectives are proposed.
[0070] The first type is for simple white mold monomers, as detailed below:
[0071] First, determine the bounding box of the white model unit, that is, the smallest cuboid structure that can enclose the entire white model unit.
[0072] Finally, four virtual perspectives in a cross shape are designed around the bounding box, and a downward virtual perspective is designed at the very center of the top of the bounding box, for a total of five virtual perspectives. Based on these five virtual perspectives, all texture details of the white model unit can be covered.
[0073] The second approach targets complex white model units. This involves not only designing virtual viewpoints around the boundaries but also within the white model unit itself. For example, depending on the complexity of the white model, multiple virtual viewpoints can be set at different heights and positions to ensure sufficient virtual view coverage of the complex details of the white model unit. The optimal virtual viewpoint is automatically calculated using algorithms. Uniform distribution and viewpoint optimization algorithms can be employed to ensure no texture details are missed. Spherical or hemispherical distribution methods can be used to enhance the coverage.
[0074] Specifically, the virtual perspective is designed for simple white model units using the following method.
[0075] Determine the bounding box (box) of the white model unit and design the resolution (r) of the virtual viewpoint, where resolution represents the actual distance represented by each pixel.
[0076] Taking a virtual viewpoint at the very top center as an example, the method for constructing the viewpoint intrinsic parameters includes:
[0077] The image dimensions of the aerial image are calculated as box.W / r and box.L / r, where box.W represents the width of the bounding box and box.L represents the length of the bounding box.
[0078] The intrinsic viewpoint parameter `camera.K` is a 3x3 matrix. `camera.K[0,0]` and `camera.K[1,1]` are 1 / r; `camera.K[0,2]` and `camera.K[1,2]` are `image.width / 2` and `image.height / 2` respectively; `camera.K[2,2]` is 1, and all other elements are 0. Here, `image.width` represents the width of the image; `image.height` represents the height of the image.
[0079] Methods for constructing extrinsic parameters of a viewpoint include:
[0080] Rotate the viewpoint to a virtual viewpoint at the top center, and use this virtual viewpoint as the main viewpoint to obtain the rotation matrix R: Translate the rotation matrix to obtain the matrix. Where box.center.x is the horizontal axis coordinate of the main view, box.center.y is the vertical axis coordinate of the main view, box.center.z is the vertical axis coordinate of the main view, and box.H is the height of the bounding box.
[0081] Based on the above method, the intrinsic and extrinsic parameters of the virtual viewpoint at the top center are constructed, and the intrinsic and extrinsic parameters of the remaining virtual viewpoints are constructed in the same way.
[0082] For constructing virtual viewpoints for complex white model units, the simple method described above cannot fully observe all textures. It is necessary to judge based on the normals of the facets. When adding a new virtual viewpoint, the rotation matrix R of the virtual viewpoint is constructed based on the given facet normals. Specifically, faces that do not cover the texture are classified. Connected regions with normals less than a certain threshold (e.g., 20°) are classified into one category. Facets with large difference in normal angles or not within connected regions can be reconstructed into a new viewpoint.
[0083] For a certain angle, the virtual viewpoint position is constructed along the normal direction of the bounding box of this type of patch at this angle. The rotation matrix R is directly opposite to the average normal direction of this type of patch. Based on R and the bounding box of the patch at this angle, the image size in the virtual viewpoint at this angle is determined. The rest is the same as the above model parameter construction method.
[0084] Based on the virtual perspectives obtained above and the trained virtual view model, a virtual view of each white model unit under the corresponding perspective is constructed. By aggregating the virtual views under all perspectives, a multi-angle view is obtained.
[0085] In this embodiment of the application, the optimal view is selected based on multiple angle views, the texture coordinates of the white model unit are calculated, and the texture region is determined, specifically including:
[0086] Based on the relationship between the surface normal and the pose of the white model unit, the back projections of multiple vertices of the white model unit are obtained.
[0087] The vertices of the face are back-projected into the corresponding virtual view of the multi-angle view to obtain the texture coordinates of the white model unit.
[0088] The initial texture region is obtained by backprojecting the vertices into texture coordinates. Then, based on multiple viewpoints, the initial texture region with the largest backprojected area is selected to obtain the final texture region. This method can maximize the clarity of the texture and avoid texture blurring caused by inconsistent texture quality from different viewpoints.
[0089] Specifically, the process of solving for texture coordinates is as follows:
[0090] Based on the patch normal, the texture coordinates corresponding to the three vertices of the patch are calculated:
[0091] First, the texture coordinates are converted from the real coordinate system to the sparse reconstruction coordinate system:
[0092] .
[0093] in, Represents the coordinates of the vertices of the face in the real coordinate system. Represents the coordinates in the sparsely reconstructed coordinate system, where [R,T] is the extrinsic parameter of the virtual viewpoint.
[0094] Secondly, the coordinates in the sparse reconstruction coordinate system are converted to the texture coordinate system:
[0095] .
[0096] Where K is the intrinsic viewpoint parameter constructed from the virtual viewpoint, u represents the x-coordinate, v represents the y-coordinate, and Z represents the y-coordinate. c f represents the depth value in the sparse reconstruction coordinate system.x f represents the focal length of the camera along the x-axis. y c represents the focal length of the camera on the y-axis. x and c y Represents the principal point of the image plane.
[0097] While this application provides the method operation steps as described in the embodiments or flowcharts, more or fewer operation steps may be included based on conventional or non-inventive labor. The order of steps listed in this embodiment is merely one possible execution order among many and does not represent the only execution order. In actual device or client product execution, the methods shown in this embodiment or the accompanying drawings can be executed sequentially or in parallel (e.g., in a parallel processor or multi-threaded processing environment).
[0098] like Figure 4 As shown in the illustration, this application also provides an apparatus 500 for performing an AI-assisted large-scale scene white model mapping method. The apparatus includes:
[0099] The image acquisition module 501 is used to acquire aerial images of the scanning area, obtain aerial images of the white model unit within the scanning area, and record the geographical location information of the aerial images.
[0100] The data processing module 502 is used to establish a real scene coordinate system based on geographical location information and to perform sparse reconstruction of aerial images based on aerial triangulation technology to obtain sparse reconstruction results.
[0101] The view generation module 503 is used to train the AI model based on the sparse reconstruction results, obtain a virtual view model, obtain multi-angle views of the white model unit based on the virtual view model, determine the best view among the multi-angle views, calculate the texture coordinates, and determine the texture region of the white model unit.
[0102] The texture mapping module 504 is used to perform texture mapping on the white model unit based on the texture region.
[0103] Some modules in the apparatus described in this application can be described in the general context of computer-executable instructions that are executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, classes, etc., that perform a specific task or implement a specific abstract data type. This application can also be practiced in distributed computing environments where tasks are performed by remote processing devices connected via a communication network. In distributed computing environments, program modules can reside in local and remote computer storage media, including storage devices.
[0104] The apparatus or module described in the above embodiments can be implemented by a computer chip or physical entity, or by a product with a certain function. For ease of description, the above apparatus is described by dividing it into various modules according to their functions. When implementing the embodiments of this application, the functions of each module can be implemented in one or more software and / or hardware. Of course, a module that implements a certain function can also be implemented by combining multiple sub-modules or sub-units.
[0105] The methods, apparatus, or modules described in this application can be implemented in a computer-readable program code manner. The controller can be implemented in any suitable manner, such as a microprocessor or processor and a computer-readable medium storing computer-readable program code (e.g., software or firmware) executable by the (micro)processor, logic gates, switches, application-specific integrated circuits (ASICs), programmable logic controllers, and embedded microcontrollers. Examples of controllers include, but are not limited to, the following microcontrollers: ARC 625D, Atmel AT91SAM, Microchip PIC18F26K20, and Silicon Labs C8051F320. A memory controller can also be implemented as part of the control logic of a memory. Those skilled in the art will also recognize that, in addition to implementing the controller in purely computer-readable program code manner, the same functionality can be achieved by logically programming the method steps to make the controller take the form of logic gates, switches, application-specific integrated circuits, programmable logic controllers, and embedded microcontrollers. Therefore, such a controller can be considered a hardware component, and the means included within it for implementing various functions can also be considered as structures within the hardware component. Alternatively, the device used to implement various functions can be viewed as either a software module that implements the method or a structure within a hardware component.
[0106] This application also provides an apparatus, the apparatus comprising: a processor; a memory for storing processor-executable instructions; wherein, when the processor executes the executable instructions, it implements the method described in this application.
[0107] This application also provides a non-volatile computer-readable storage medium storing a computer program or instructions thereon, which, when executed, enables the method described in this application embodiment to be implemented.
[0108] Furthermore, in the various embodiments of the present invention, each functional module can be integrated into a processing module, or each module can exist independently, or two or more modules can be integrated into a single module.
[0109] The aforementioned storage media include, but are not limited to, Random Access Memory (RAM), Read-Only Memory (ROM), Cache, Hard Disk Drive (HDD), or Memory Card. The memory can be used to store computer program instructions.
[0110] As can be seen from the above description of the embodiments, those skilled in the art can clearly understand that this application can be implemented by means of software plus necessary hardware. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product, or it can be embodied in the process of data migration. The computer software product can be stored in a storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, mobile terminal, server, or network device, etc.) to execute the methods described in various embodiments or some parts of the embodiments of this application.
[0111] The various embodiments described in this specification are presented in a progressive manner. Similar or identical parts between embodiments can be referred to interchangeably. Each embodiment focuses on its differences from other embodiments. All or part of this application can be used in numerous general-purpose or special-purpose computer system environments or configurations. Examples include: personal computers, server computers, handheld or portable devices, tablet devices, mobile communication terminals, multiprocessor systems, microprocessor-based systems, programmable electronic devices, network PCs, minicomputers, mainframe computers, and distributed computing environments including any of the above systems or devices, etc.
[0112] The above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit this application. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features therein. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of this application.
Claims
1. A method for scene white model texturing based on AI assistance, characterized in that, include: Acquire aerial images of the white model unit within the scanned area and establish a real scene coordinate system for the scanned area; Sparse reconstruction is performed on aerial images to obtain sparse reconstruction results, and the sparse reconstruction results are input into the real scene coordinate system to obtain the coordinates of the white model unit. Based on the coordinates of the white model unit, the back projection error is filtered out by setting an error threshold to obtain the filtered coordinates of the white model unit. The AI model is then trained based on the filtered coordinates of the white model unit to obtain the virtual view model. The perspective of the white model unit is constructed, and multi-angle views of the white model unit are generated based on the virtual view model; Based on the multi-angle view, the best view is selected, the texture coordinates of the white model unit are calculated, and the back projections of multiple vertices of the white model unit are obtained according to the relationship between the normal of the facet and the pose of the white model unit. The back projections of the facet vertices are input into the corresponding views of the multi-angle view to obtain the texture coordinates of the white model unit. The initial texture region is obtained according to the position of the vertex back projection in the texture coordinates, and the initial texture region with the largest back projection area is selected based on the multi-angle view to obtain the maximum texture region. The white model unit is textured in the best view based on the maximum texture region.
2. The method according to claim 1, characterized in that, The sparse reconstruction of aerial images, resulting in sparse reconstruction results, includes: Aerial triangulation is used to sparsely reconstruct aerial images in order to determine the relative pose relationship between each aerial image and obtain sparse reconstruction results. The sparse reconstruction results are presented in coordinate form.
3. The method according to claim 2, characterized in that, The step of inputting the sparse reconstruction result into the real scene coordinate system to obtain the coordinates of the white model unit includes: GPS information is obtained from aerial imagery and then converted into projected coordinates. A sparse reconstruction coordinate system is established using aerial triangulation technology, and the sparse reconstruction results are correlated with the projected coordinates to obtain the transformation matrix between the sparse reconstruction coordinate system and the real scene coordinate system. The sparse reconstruction results are input into the real scene coordinate system, and the sparse reconstruction results are converted into white model unit coordinates based on the transformation matrix. The transformation matrix includes: ; Where s represents the scale factor, R represents the rotation matrix, and T represents the translation vector.
4. The method according to claim 3, characterized in that, After inputting the sparse reconstruction result into the real scene coordinate system to obtain the coordinates of the white model unit, the process further includes: Obtain the back projection error of the white model unit; The process of obtaining the back projection error of the white model unit includes: Multiple 3D sparse point clouds are obtained based on the sparse reconstruction results. The 3D sparse point clouds are then input into the projection coordinates to obtain multiple projection positions. Convert 3D sparse point clouds into pixel coordinates based on intrinsic and extrinsic parameters; Errors in calculating 3D sparse point clouds based on projection position and pixel coordinates; Calculate the mean of the overall error of the 3D sparse point cloud to obtain the back projection error of the individual white model; The calculation method for the error of the 3D sparse point cloud includes: ; in, The x-coordinate represents the projection position. The x-coordinate of the pixel coordinate. The ordinate represents the projected position. The ordinate representing the pixel coordinate; The method for calculating the mean of the overall error of the 3D sparse point cloud includes: ; Where m represents the number of 3D sparse point clouds.
5. The method according to claim 4, characterized in that, The process of training an AI model based on the coordinates of a white model unit to obtain a virtual view model includes: Set an error threshold for the coordinates of individual white model units, and filter out back projection errors based on the error threshold to ensure the stability of the model, and obtain the coordinates of individual white model units after filtering. The AI model is trained based on the coordinates of the filtered white model units to obtain the virtual view model; The filtering of back projection errors of the white model unit based on an error threshold includes: Back projection errors with an error range greater than the error threshold are filtered out, while other back projection errors are retained.
6. The method according to claim 1, characterized in that, The process of constructing the perspective of the white model unit and generating multi-angle views of the white model unit based on the virtual view model includes: Based on the geometric structure and specific location of the white model unit, multiple virtual perspectives of the white model unit are generated; Input the virtual perspective into the virtual view model to obtain multiple virtual views of the single white model; By combining multiple virtual views, a multi-angle view of the white model unit can be obtained.
7. An apparatus for performing an AI-assisted scene white model mapping method, characterized in that, include: The image acquisition module is used to acquire aerial images of the scanning area, obtain aerial images of the white model unit within the scanning area, and record the geographical location information of the aerial images. The data processing module is used to establish a real scene coordinate system based on geographic location information and to perform sparse reconstruction of aerial images based on aerial triangulation technology to obtain sparse reconstruction results. The viewpoint generation module is used to filter out backprojection errors based on sparse reconstruction results by setting an error threshold, obtaining the coordinates of the filtered white model unit. An AI model is then trained based on these coordinates to obtain a virtual view model. Multiple views of the white model unit are generated from the virtual view model, determining the optimal view among the multiple views. Texture coordinates are calculated, and multiple vertex backprojections of the white model unit are obtained based on the relationship between the facet normals and the pose of the white model unit. The vertex backprojections of the facets are input into the corresponding views of the multiple views to obtain the texture coordinates of the white model unit. An initial texture region is obtained based on the position of the vertex backprojections in the texture coordinates, and the initial texture region with the largest backprojection area is selected based on the multiple views to obtain the maximum texture region. The texture mapping module is used to apply texture mapping to individual white model units based on texture regions.
8. A device for performing an AI-assisted scene white model texture mapping method, characterized in that, include: processor; Memory used to store processor-executable instructions; When the processor executes the executable instructions, it implements the method as described in any one of claims 1 to 6.
9. A non-volatile computer-readable storage medium, characterized in that, Includes storage of computer programs or instructions that, when executed, cause the method as described in any one of claims 1 to 6 to be implemented.
Citation Information
Patent Citations
Texture map generation method, device and equipment of 3D white mold and medium
CN117218266A
Urban white pattern production method and device based on AI prediction
CN118135102A