Methods and apparatus for modeling building complexes
By acquiring video and remote sensing images of the building complex, and using camera parameters and neural network models to generate a target 3D model, the problem of low accuracy of virtual 3D models in existing technologies is solved, and a higher precision modeling effect is achieved.
Patent Information
- Application Number
- CN202210138921.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-02-15
- Publication Date
- 2025-10-31
- Estimated Expiration
- 2042-02-15
AI Technical Summary
Existing technologies for creating virtual 3D models have low accuracy and cannot effectively solve the accuracy problem of building complex modeling.
By acquiring video and remote sensing images of the building complex, the camera parameters of the cameras are used to determine the virtual camera parameters in the virtual scene. Combined with a neural network model, a target 3D model is generated, and texture data is added to optimize the model generation process.
It improved the accuracy of the virtual 3D model of the building complex, achieving a more precise modeling effect.
Smart Images

Figure CN114565917B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of computers, and more specifically, to a method and apparatus for modeling building complexes. Background Technology
[0002] In existing technologies, in certain scenarios, it is often necessary to model building complexes and obtain virtual 3D models of the complexes. However, the accuracy of virtual 3D models created by existing methods is relatively low.
[0003] There is currently no effective solution to the above problems. Summary of the Invention
[0004] This invention provides a method and apparatus for modeling building complexes, which at least solves the technical problem of low accuracy in creating virtual three-dimensional models of building complexes.
[0005] According to one aspect of the present invention, a method for modeling building complexes is provided, comprising:
[0006] Acquire video and remote sensing images of the building complex to be modeled; determine the virtual camera parameters of a virtual camera matching the camera in the virtual scene based on the first camera parameters of the camera that captured the video; determine the target location of the building complex to be modeled in the virtual scene based on the correspondence between the first camera parameters and the virtual camera parameters; generate a target 3D model of the building complex to be modeled based on the remote sensing images; place the target 3D model at the target location; and add texture data from the video images to the surface of the target 3D model.
[0007] According to another aspect of the present invention, a building complex modeling apparatus is also provided, comprising: an acquisition unit for acquiring video images and remote sensing images of a building complex to be modeled; a first determining unit for determining virtual camera parameters of a virtual camera matching the camera in a virtual scene based on first camera parameters of the camera that captured the video images; a second determining unit for determining the target position of the building complex to be modeled in the virtual scene based on the correspondence between the first camera parameters and the virtual camera parameters; a generation unit for generating a target three-dimensional model of the building complex to be modeled based on the remote sensing images; a processing unit for placing the target three-dimensional model at the target position; and an adding unit for adding texture data from the video images to the surface of the target three-dimensional model.
[0008] As an optional example, the above-mentioned generation unit includes: a first input module for inputting the remote sensing image into a target neural network model, wherein the target neural network model is used to model the remote sensing image to obtain the target three-dimensional model; a third acquisition module for acquiring the modeling result output by the target neural network model; and a second determination module for determining the modeling result as the target three-dimensional model.
[0009] As an optional example, the above-mentioned generation unit further includes: a fourth acquisition module, used to acquire sample remote sensing images and sample modeling results of the sample remote sensing images before inputting the remote sensing images into the target neural network model; an extraction module, used to extract multiple sample pixel blocks from the sample remote sensing images; a deletion module, used to delete pixel blocks with a sharpness lower than a first threshold from the multiple sample pixel blocks; a processing module, used to perform at least one of blurring, rotation, magnification, and reduction processing on each of the remaining sample pixel blocks to obtain a first pixel block; a second input module, used to input the first pixel block as a training sample into the original neural network model to obtain the original modeling result output by the original neural network model; and an adjustment module, used to adjust the model parameters of the original neural network model according to the comparison result between the original modeling result and the sample modeling result, until the target neural network model is obtained.
[0010] As an optional example, the second input module includes: an acquisition submodule for acquiring multiple intermediate models of different sizes obtained by the original neural network model recognizing the first pixel block; and a scaling submodule for scaling the multiple intermediate models to obtain the original modeling result of a target size.
[0011] As an optional example, the above adjustment module includes: a first determining submodule, used to determine the cross-entropy loss and boundary loss of the above original modeling result and the above sample modeling result; a second determining submodule, used to determine the weighted summation result of the above cross-entropy loss and the above boundary loss; an adjustment submodule, used to adjust the above model parameters when the above weighted summation result is less than a second threshold; and a third determining submodule, used to determine the above original neural network model as the above target neural network model when the above weighted summation result is greater than or equal to the above second threshold.
[0012] According to another aspect of the present invention, a computer-readable storage medium is also provided, wherein a computer program is stored therein, wherein the computer program is configured to execute the above-described building complex modeling method at runtime.
[0013] According to another aspect of the present invention, an electronic device is also provided, including a memory and a processor, wherein the memory stores a computer program and the processor is configured to execute the above-described building complex modeling method through the computer program.
[0014] In this embodiment of the invention, a method is employed that involves acquiring video and remote sensing images of a building complex to be modeled; determining the virtual camera parameters of a virtual camera matching the camera in the virtual scene based on the first camera parameters of the camera that captured the video; determining the target position of the building complex in the virtual scene based on the correspondence between the first camera parameters and the virtual camera parameters; generating a target 3D model of the building complex based on the remote sensing images; placing the target 3D model at the target position; and adding texture data from the video images to the surface of the target 3D model. Because this method can acquire video and remote sensing images of the building complex, it allows for determining the position of the building complex in the virtual scene based on the video images and determining the target 3D model of the building complex based on the remote sensing images. This allows for placing the target 3D model at the target position in the virtual scene and adding texture data to the target 3D model, thereby improving the accuracy of the established virtual 3D model and solving the technical problem of low accuracy in establishing virtual 3D models of building complexes. Attached Figure Description
[0015] The accompanying drawings, which are included to provide a further understanding of the invention and form part of this application, illustrate exemplary embodiments of the invention and, together with their description, serve to explain the invention and do not constitute an undue limitation thereof. In the drawings:
[0016] Figure 1 This is a schematic diagram of the application environment of an optional building group modeling method according to an embodiment of the present invention;
[0017] Figure 2 This is a schematic diagram of the application environment of another optional building group modeling method according to an embodiment of the present invention;
[0018] Figure 3 This is a schematic diagram of the flow of an optional building complex modeling method according to an embodiment of the present invention;
[0019] Figure 4 This is a system schematic diagram of an optional building complex modeling method according to an embodiment of the present invention;
[0020] Figure 5 This is a flowchart illustrating an optional building complex modeling method according to an embodiment of the present invention;
[0021] Figure 6This is a schematic diagram of an optional building complex modeling device according to an embodiment of the present invention. Detailed Implementation
[0022] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.
[0023] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0024] According to one aspect of the present invention, a method for modeling building complexes is provided. Optionally, as an alternative implementation, the above-described method for modeling building complexes can be applied to, but is not limited to, [examples of other methods]. Figure 1 In the environment shown.
[0025] like Figure 1 As shown, the terminal device 102 includes a memory 104 for storing various data generated during its operation, a processor 106 for processing and calculating the aforementioned data, and a display 108 for displaying the modeling results. The terminal device 102 can interact with the server 112 via a network 110. The server 112 includes a database 114 for storing various data and a processing engine 116 for processing the aforementioned data. In steps S102 to S106, the terminal device 102 acquires video and remote sensing images of the building complex to be modeled, sends the video and remote sensing images to the server 112, which performs modeling and returns the modeling results.
[0026] As an optional implementation, the above-described building complex modeling method can be applied, but is not limited to, to applications such as... Figure 2 In the environment shown.
[0027] like Figure 2 As shown, the terminal device 202 includes a memory 204 for storing various data generated during the operation of the terminal device 202, a processor 206 for processing and calculating the aforementioned data, and a display 208 for displaying the modeling results. The terminal device 202 can execute steps S202 to S212 to achieve modeling and display the modeling results.
[0028] Optionally, in this embodiment, the terminal device can be a terminal device configured with a target client, which may include, but is not limited to, at least one of the following: mobile phone (such as Android phone, iOS phone, etc.), laptop computer, tablet computer, PDA, MID (Mobile Internet Devices), PAD, desktop computer, smart TV, etc. The target client may be a video client, instant messaging client, browser client, educational client, etc. The network may include, but is not limited to, wired network and wireless network, wherein the wired network includes: local area network, metropolitan area network and wide area network, and the wireless network includes: Bluetooth, WIFI and other networks that enable wireless communication. The server may be a single server, a server cluster composed of multiple servers, or a cloud server. The above is only an example, and no limitation is made in this embodiment.
[0029] Alternatively, as an alternative implementation method, such as Figure 3 As shown, the above-mentioned building complex modeling methods include:
[0030] S302, acquire video and remote sensing images of the building complex to be modeled;
[0031] S304, Based on the first camera parameters of the camera that captures video images, determine the virtual camera parameters of the virtual camera that matches the camera in the virtual scene;
[0032] S306. Based on the correspondence between the parameters of the first camera and the parameters of the virtual camera, determine the target position of the building complex to be modeled in the virtual scene.
[0033] S308: Generate a target 3D model of the building complex to be modeled based on remote sensing imagery.
[0034] S310, Place the target 3D model at the target location;
[0035] S312 adds texture data from the video image to the surface of the target 3D model.
[0036] Optionally, the above modeling method can be applied to the process of modeling building complexes. The video images mentioned above can be images captured by a camera. By capturing video images with a camera and acquiring remote sensing images, the position of the target 3D model can be determined from the video images, and the target 3D model can be determined from the remote sensing images. Then, the target 3D model can be placed at the target position, and texture data can be added to the target 3D model, thus achieving the effect of accurately building the target 3D model.
[0037] As an example, the virtual camera parameters for determining a virtual camera that matches the camera in a virtual scene, based on the first camera parameters of the camera capturing video images, include:
[0038] Multiple pairs of matching points are obtained from video frames in the video image and virtual scene, wherein each pair of matching points includes a two-dimensional point in the video frame and a three-dimensional point in the virtual scene;
[0039] Obtain the angle and position information from the first camera parameters;
[0040] Based on the mapping relationship from two-dimensional points to three-dimensional points, angle and position information are mapped onto the virtual scene to obtain virtual camera parameters.
[0041] Optionally, in this embodiment, virtual cameras matching the physical camera can be created in the virtual scene. The number, angle, and position of the virtual cameras in the virtual scene are the same as the number, angle, and position of the physical cameras in the real scene.
[0042] The matching points mentioned above can be points that correspond to each other in the virtual scene and the real scene.
[0043] By using matching points, cameras in real-world scenes can be matched to virtual scenes, thus completing the matching between cameras and virtual cameras.
[0044] As an example, based on the correspondence between the parameters of the first camera and the parameters of the virtual camera, the target location of the building complex to be modeled in the virtual scene is determined, including:
[0045] Determine the orientation of the building complex to be modeled relative to the camera and its distance from the camera;
[0046] Based on the correspondence between the parameters of the first camera and the parameters of the virtual camera, the direction and distance are mapped onto the virtual scene to determine the target position.
[0047] In this embodiment, since a virtual camera matching the camera can be determined in the virtual scene, the image captured by the camera can be matched to the virtual scene accordingly. Based on information such as the object distance in the image, the target position is determined by matching the virtual scene according to the corresponding ratio.
[0048] As an example, generating a target 3D model of a building complex to be modeled based on remote sensing imagery includes:
[0049] The remote sensing image is input into the target neural network model, which is used to model the remote sensing image to obtain a three-dimensional model of the target.
[0050] Obtain the modeling results output by the target neural network model;
[0051] The modeling results are used as the target 3D model.
[0052] Optionally, in this embodiment, remote sensing images can be input into a target neural network model, and the target neural network model can identify the remote sensing images to generate a target 3D model.
[0053] As an example, the method further includes the following steps before inputting remotely sensed imagery into the target neural network model:
[0054] Obtain sample remote sensing images and sample modeling results from sample remote sensing images;
[0055] Extract multiple sample pixel blocks from sample remote sensing images;
[0056] Delete pixel blocks with a resolution lower than the first threshold from multiple sample pixel blocks;
[0057] Perform at least one of the following processing methods on each of the remaining sample pixel blocks: blurring, rotation, magnification, and reduction, to obtain the first pixel block;
[0058] The first pixel block is used as a training sample and input into the original neural network model to obtain the original modeling result output by the original neural network model.
[0059] Based on the comparison between the original modeling results and the sample modeling results, the model parameters of the original neural network model are adjusted until the target neural network model is obtained.
[0060] In other words, the target neural network model in this embodiment is a model pre-trained using sample remote sensing images. During the training process, multiple sample pixel blocks need to be obtained from the sample remote sensing images. Then, sample pixel blocks with low resolution are deleted. Next, the remaining sample pixel blocks are subjected to at least one of the following processing methods: blurring, rotation, magnification, and reduction, to obtain the first pixel block. The first pixel block is used to train the original neural network model to obtain the trained target neural network model.
[0061] As an example, by inputting the first pixel block as a training sample into the original neural network model, the original modeling results output by the original neural network model include:
[0062] Obtain multiple intermediate models of different sizes from the original neural network model that recognize the first pixel block;
[0063] Multiple intermediate models are scaled to obtain an original modeling result of a target size.
[0064] Optionally, during the training of the original neural network model, multiple intermediate models of different sizes can be obtained by recognizing the first pixel block. The intermediate models can be virtual modeling models in the recognition process. Since the intermediate models are of different sizes, they need to be scaled to obtain an original modeling model of one size.
[0065] As an example, based on the comparison between the original modeling results and the sample modeling results, the model parameters of the original neural network model are adjusted until the target neural network model is obtained, including:
[0066] Determine the cross-entropy loss and boundary loss of the original modeling results and the sample modeling results;
[0067] Determine the weighted sum of the cross-entropy loss and the boundary loss;
[0068] If the weighted summation result is less than the second threshold, adjust the model parameters;
[0069] If the weighted summation result is greater than or equal to the second threshold, the original neural network model is determined as the target neural network model.
[0070] By applying the aforementioned loss constraints, the original neural network model is trained, ultimately resulting in a target neural network model with high recognition accuracy.
[0071] The following example illustrates this. In this embodiment, the system consists of three layers: data acquisition, intelligent analysis, and business application. The system framework is as follows: Figure 4 As shown. The data acquisition layer includes satellite remote sensing image data, UAV remote sensing image data, and high-altitude and ground video surveillance data; the intelligent analysis layer is supported by remote sensing image intelligent analysis, 3D video fusion, and 3D modeling algorithms, and outputs a wide-area spatial 3D model containing multi-dimensional information, which is the target 3D model; it can be applied to urban infrastructure review, illegal construction investigation and handling, ancient building protection and other businesses.
[0072] 3D modeling technology is achieved through aerial survey data acquisition combined with post-processing algorithms. Aerial survey data acquisition mainly relies on remote sensing and UAV platforms, using oblique photogrammetry to simultaneously obtain multiple high-resolution images of the same location from different angles, i.e., acquiring video images and remote sensing images, and collecting rich information on the side textures and locations of ground features. Then, based on the acquired images and location information, aerial triangulation data is obtained, the images are densely matched, and a 3D irregular triangular mesh (TIN) model is constructed based on the matching results. Finally, the textures are mapped onto the model to complete the modeling.
[0073] 3D video fusion technology allows for interactive adjustment of the internal and external parameters of a virtual camera, ensuring that the monitoring scene of the virtual camera in the 3D model matches the real video monitoring scene. Then, the targets detected in the real video monitoring scene are mapped to their corresponding positions in the 3D model, achieving real-time stereoscopic fusion display of the 3D model and the monitoring video.
[0074] Intelligent remote sensing image analysis technology, leveraging algorithms from image processing, machine learning, and deep learning, enables semantic segmentation, target detection, and change detection for multi-level, multi-view, and multi-domain remote sensing observation systems, thereby improving the efficiency of spatial information transmission and practical application capabilities. Specifically, the change detection algorithm based on remote sensing images outputs the differences between different time-phase data of the same location, quickly identifying land cover change processes; the semantic segmentation algorithm based on remote sensing images achieves pixel-level classification of wide-area regions, rapidly acquiring land cover and utilization information, and is widely applied in agriculture, urban planning, environmental protection, and land resource management; and the target detection algorithm based on remote sensing images automatically determines and identifies the categories and locations of multiple targets, quickly obtaining target distribution information within a wide area.
[0075] To address the characteristics of large-size remote sensing images and class imbalance in changed areas, a preprocessing method based on multi-channel fusion for image stretching and normalization is proposed to resolve issues such as significant differences in ground representation. Regarding model structure, the performance of FCN and DeepLab frameworks in remote sensing change detection tasks is evaluated. Building upon the U-Net algorithm framework, a Tversky Loss function is proposed to alleviate class imbalance. Furthermore, R2-Unet / Attention-Unet, combining recurrent convolution and residual networks, is introduced to construct a multimodal architecture, significantly improving both precision and recall.
[0076] To address the characteristics of large remote sensing images, background interference, and small target size, a data augmentation strategy based on appearance consistency heatmaps is proposed for data preprocessing to solve the problem of imbalanced distribution of target samples of specific categories. In terms of model structure, optimizations are made in three aspects: feature representation, network output, and equalization, based on the ROITransformer algorithm framework.
[0077] In the specific process, this embodiment first utilizes UAV oblique photogrammetry technology to achieve high-precision 3D modeling of a wide-area space (a specific administrative region of a city or the entire city, which may include building clusters and the natural environment). Addressing the issues of long modeling cycles and complex schemes in wide-area 3D modeling, before planning UAV flight routes, remote sensing imagery is used to determine the aerial survey range, understand the aerial survey terrain, rationally allocate flight sorties, optimize the aerial photography scheme, and improve operational efficiency. In other words, the scope of the building clusters to be modeled is first determined based on remote sensing imagery, and then UAVs are deployed for filming. Finally, 3D video fusion technology is used to project real-time video monitoring information within the area onto the corresponding scene of the 3D model.
[0078] In this process, to further enhance the information representation capability of the wide-area spatial 3D model, this embodiment proposes a multi-layer 3D model by fusing intelligent analysis results of remote sensing images. The model identifies remote sensing images into four layers: Layer 1 is a 3D model that integrates real-time video surveillance information; Layer 2 is the result of a change detection algorithm, which obtains the location of areas that have changed between two time phases within the corresponding remote sensing image area, intuitively displaying the changes in urban space; Layer 3 is the result of a semantic segmentation algorithm, which performs pixel-level classification on the remote sensing image area corresponding to the 3D model, intuitively displaying the land use type of the space; and Layer 4 is the result of a target detection algorithm, which obtains information such as the location and quantity of targets of interest within the corresponding remote sensing image area, intuitively displaying the distribution status of targets in urban space.
[0079] Furthermore, this embodiment proposes using intelligent remote sensing image analysis to guide 3D model updates. Specifically, based on the results of the change detection algorithm, the location, area, and magnitude of areas where changes have occurred between the current time phase and the modeling time phase within the remote sensing image region corresponding to the 3D model are obtained. If the degree of change meets the model update settings, a UAV aerial photography plan is developed for the changed area, the model is reconstructed, and the original model is locally updated. This local update strategy based on the results of intelligent remote sensing image analysis can effectively solve the problems of inefficiency and high resource consumption encountered in model reconstruction.
[0080] Figure 5 This is a flowchart of this embodiment. Figure 5 As shown, S502 to S510,
[0081] In the video image fusion step, specifically the 3D video fusion step, the 3D model is first fused with the real scene through camera calibration, coordinate transformation, and texture mapping. This technology first calibrates the intrinsic parameters of the real-world camera; then, through interaction between the camera image and the virtual 3D scene, multiple pairs of matching points are obtained. Using these multiple 3D-to-2D matching points, the camera's extrinsic parameters are solved by minimizing the reprojection error. These extrinsic parameters include the camera's 3D spatial position, yaw angle, pitch angle, viewpoint position, and field of view, ensuring that the virtual camera's monitoring scene in the 3D model coincides with the real video monitoring scene; Formula 1:
[0082]
[0083] in These are camera-intra-camera parameters. These are camera extrinsic parameters. To monitor the position of objects in the scene, s represents the position of the object in the virtual 3D model, and s is the scale factor.
[0084] Then, using Formula 1 above, the target detected in the real video surveillance scene is projected to the corresponding position in the 3D model, and the size of the real scene target in the 3D model is estimated according to the scale factor; finally, the texture captured by the camera is converted into the virtual 3D model according to a fixed position, so that the model is bound to the corresponding texture, and the real-time stereoscopic fusion display of the 3D model and the surveillance video is realized.
[0085] Due to the complex distribution, diverse boundaries, and significant differences in area size of segmented targets in remote sensing images, these factors severely limit the accuracy of semantic segmentation algorithms for remote sensing images, thus restricting the large-scale application of high-resolution remote sensing images. Therefore, deep learning-based semantic segmentation methods for remote sensing images exhibit a significant accuracy advantage over traditional methods, but they heavily rely on the quantity and quality of labeled data, as well as the performance of training resources such as GPUs and storage space. Therefore, data augmentation methods, data transfer methods, the development of high-precision models, and the optimization of training methods have become the main development directions in the field of semantic segmentation algorithms for remote sensing images. In this embodiment, the semantic segmentation problem of remote sensing images is divided into binary segmentation and multi-class segmentation according to the segmentation target category, and the algorithm advantages are illustrated using automatic extraction of remote sensing buildings and subdivision of building types as examples:
[0086] 1) Based on the characteristics of remote sensing images, targeted data preprocessing was performed. The original remote sensing data had pixel sizes in the tens of thousands, which were too large for deep learning networks. Therefore, the original data was cropped to a target size of 512x512 pixels with a step size of 256. Cropping yielded a large number of smaller 512x512 images. Data cleaning was then performed, removing images with poor image quality from the smaller images. For the remaining images, data augmentation techniques such as blurring, scaling, and rotation were applied to increase the data volume. Additionally, ground truth images of building edges could be appropriately enhanced to improve the learning ability of building edges.
[0087] 2) Select the simple and efficient U-net network structure, and add branch convolutional networks to extract building edges, thereby improving the learning ability of the entire basic network to buildings and effectively segmenting buildings.
[0088] 3) Integrate and improve multiple loss functions to jointly optimize the final segmentation result. This scheme initially attempted to use the traditional cross-entropy loss function, but later improved it by incorporating boundary loss and online hard sample mining (OHEM) techniques to enhance the feedback of the loss function on the learning ability.
[0089] 4) Multi-scale prediction and morphological post-processing methods are employed to effectively fuse and improve prediction results, forming the final building segmentation result. During the network's forward pass, the predicted image is scaled down to 3 to 5 different sizes. The network predicts result images at different scaling rates. Finally, these results at different resolutions are scaled back to the original image size and fused to obtain multi-scale prediction results. Then, opening and closing operations are used to remove scattered small pixels and fill gaps in buildings, forming a complete building mask. Connected component extraction is used to assign numbers to all buildings, forming the final building segmentation result.
[0090] Land use classification extraction - multi-class segmentation
[0091] 1) This solution employs an advanced deep learning network model that balances accuracy and real-time performance requirements. It utilizes HRNet, which improves the accuracy and computational power of the network by modifying the base network and using a high-precision, low-complexity base network. Furthermore, it improves accuracy by modifying the loss function and using OHEM and IoU loss to guide the learning direction of the entire network.
[0092] 2) By leveraging generative adversarial network technology, an effective data augmentation and style transfer method was constructed. Based on the existing StandardGAN technology, this case normalizes remote sensing datasets from different sources (i.e., data with different distributions) to a unified data distribution, so that the feature space of data with different styles remains consistent during training and testing, effectively improving the model's prediction accuracy.
[0093] 3) Detailed surveys and classifications of land use types have greatly enhanced the model's compatibility.
[0094] This embodiment proposes a remote sensing + video fusion technology, which can observe the real-time dynamic changes of the region of interest from macro to micro perspectives. Furthermore, through remote sensing + UAV oblique photography + 3D modeling technology, it can reconstruct the target 3D model after regional changes in a timely and efficient manner.
[0095] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that the present invention is not limited to the described order of actions, because according to the present invention, some steps can be performed in other orders or simultaneously. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are preferred embodiments, and the actions and modules involved are not necessarily essential to the present invention.
[0096] According to another aspect of the present invention, a building complex modeling apparatus for implementing the above-described building complex modeling method is also provided. For example... Figure 6 As shown, the device includes:
[0097] Acquisition unit 602 is used to acquire video and remote sensing images of the building complex to be modeled.
[0098] The first determining unit 604 is used to determine the virtual camera parameters of the virtual camera that matches the camera in the virtual scene based on the first camera parameters of the camera that captures video images.
[0099] The second determining unit 606 is used to determine the target position of the building complex to be modeled in the virtual scene based on the correspondence between the first camera parameters and the virtual camera parameters.
[0100] The generation unit 608 is used to generate a target 3D model of the building complex to be modeled based on remote sensing images.
[0101] Processing unit 610 is used to place the target 3D model at the target location;
[0102] Add unit 612 to add texture data from video images to the surface of the target 3D model.
[0103] As an optional example, the first determining unit mentioned above includes:
[0104] The first acquisition module is used to acquire multiple pairs of matching points from video frames in the video image and virtual scene, wherein each pair of matching points includes a two-dimensional point in the video frame and a three-dimensional point in the virtual scene;
[0105] The second acquisition module is used to acquire the angle information and position information from the first camera parameters;
[0106] The first mapping module is used to map angle and position information onto the virtual scene according to the mapping relationship from two-dimensional points to three-dimensional points, so as to obtain virtual camera parameters.
[0107] As an optional example, the second determining unit described above includes:
[0108] The first determining module is used to determine the direction of the building complex to be modeled relative to the camera and its distance from the camera;
[0109] The second mapping module is used to map the direction and distance to the virtual scene according to the correspondence between the first camera parameters and the virtual camera parameters, thereby determining the target position.
[0110] As an optional example, the above-mentioned generation unit includes:
[0111] The first input module is used to input remote sensing images into the target neural network model, wherein the target neural network model is used to model the remote sensing images to obtain a target three-dimensional model;
[0112] The third acquisition module is used to acquire the modeling results output by the target neural network model;
[0113] The second determination module is used to determine the modeling results as the target 3D model.
[0114] As an optional example, the above-mentioned generation unit also includes:
[0115] The fourth acquisition module is used to acquire sample remote sensing images and sample modeling results of sample remote sensing images before inputting remote sensing images into the target neural network model.
[0116] The extraction module is used to extract multiple sample pixel blocks from sample remote sensing images;
[0117] The deletion module is used to delete pixel blocks with a resolution lower than a first threshold from multiple sample pixel blocks.
[0118] The processing module is used to perform at least one of blurring, rotation, magnification and reduction processing on each of the remaining sample pixel blocks to obtain the first pixel block;
[0119] The second input module is used to input the first pixel block as a training sample into the original neural network model to obtain the original modeling result output by the original neural network model.
[0120] The adjustment module is used to adjust the model parameters of the original neural network model based on the comparison results between the original modeling results and the sample modeling results, until the target neural network model is obtained.
[0121] As an optional example, the second input module described above includes:
[0122] The acquisition submodule is used to acquire multiple intermediate models of different sizes obtained by the original neural network model recognizing the first pixel block;
[0123] The scaling submodule is used to scale multiple intermediate models to obtain a raw modeling result of a target size.
[0124] As an optional example, the above adjustment module includes:
[0125] The first determination submodule is used to determine the cross-entropy loss and boundary loss of the original modeling results and the sample modeling results;
[0126] The second determination submodule is used to determine the weighted sum of the cross-entropy loss and the boundary loss;
[0127] The adjustment submodule is used to adjust the model parameters when the weighted summation result is less than the second threshold.
[0128] The third determination submodule is used to determine the original neural network model as the target neural network model if the weighted summation result is greater than or equal to the second threshold.
[0129] For other examples of this embodiment, please refer to the above embodiments, which will not be repeated here.
[0130] According to another aspect of the present invention, an electronic device for implementing the above-described building complex modeling method is also provided. The electronic device includes a memory and a processor, the memory storing a computer program, and the processor being configured to execute the steps of any of the above-described method embodiments via the computer program.
[0131] According to another aspect of the present invention, a computer-readable storage medium is also provided, wherein a computer program is stored therein, wherein the computer program is configured to perform the steps in any of the above method embodiments when it is run.
[0132] Optionally, in this embodiment, those skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be implemented by a program instructing the hardware related to the terminal device. The program can be stored in a computer-readable storage medium, which may include: flash drive, read-only memory (ROM), random access memory (RAM), disk or optical disk, etc.
[0133] The sequence numbers of the above embodiments of the present invention are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.
[0134] If the integrated units in the above embodiments are implemented as software functional units and sold or used as independent products, they can be stored in the aforementioned computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause one or more computer devices (which may be personal computers, servers, or network devices, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention.
[0135] In the above embodiments of the present invention, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments.
[0136] In the several embodiments provided in this application, it should be understood that the disclosed client can be implemented in other ways. The device embodiments described above are merely illustrative; for example, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces, indirect coupling or communication connection between units or modules, and may be electrical or other forms.
[0137] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0138] Furthermore, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0139] The above description is only a preferred embodiment of the present invention. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principle of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention.
Claims
1. A method for modeling building complexes, characterized in that, include: Acquire video and remote sensing images of the building complex to be modeled; Based on the first camera parameters of the camera that captured the video image, determine the virtual camera parameters of the virtual camera that matches the camera in the virtual scene; Based on the correspondence between the first camera parameters and the virtual camera parameters, the target location of the building complex to be modeled in the virtual scene is determined; Obtain the sample remote sensing image and the sample modeling results of the sample remote sensing image; Extract multiple sample pixel blocks from the sample remote sensing image; Delete pixel blocks whose clarity is lower than the first threshold from among the multiple sample pixel blocks; Perform at least one of blurring, rotation, magnification, and reduction processing on each of the remaining sample pixel blocks to obtain the first pixel block; The first pixel block is used as a training sample and input into the original neural network model to obtain the original modeling result output by the original neural network model. Based on the comparison between the original modeling results and the sample modeling results, the model parameters of the original neural network model are adjusted until the target neural network model is obtained; The remote sensing image is input into the target neural network model, wherein the target neural network model is used to model the remote sensing image to obtain a target three-dimensional model; Obtain the modeling results output by the target neural network model; The modeling result is determined as the target 3D model; Place the target 3D model at the target location; The texture data from the video image is added to the surface of the target 3D model.
2. The method according to claim 1, characterized in that, The step of determining the virtual camera parameters of the virtual camera matching the camera in the virtual scene based on the first camera parameters of the camera that captured the video image includes: Multiple pairs of matching points are obtained from video frames in the video image and the virtual scene, wherein each pair of matching points includes a two-dimensional point in the video frame and a three-dimensional point in the virtual scene; Obtain the angle and position information from the first camera parameters; According to the mapping relationship from the two-dimensional point to the three-dimensional point, the angle information and position information are mapped onto the virtual scene to obtain the virtual camera parameters.
3. The method according to claim 1, characterized in that, Determining the target location of the building complex to be modeled in the virtual scene based on the correspondence between the first camera parameters and the virtual camera parameters includes: Determine the orientation of the building complex to be modeled relative to the camera and its distance from the camera; According to the correspondence between the first camera parameters and the virtual camera parameters, the direction and distance are mapped onto the virtual scene to determine the target position.
4. The method according to claim 1, characterized in that, The step of inputting the first pixel block as a training sample into the original neural network model to obtain the original modeling result output by the original neural network model includes: Obtain multiple intermediate models of different sizes obtained by recognizing the first pixel block using the original neural network model; The intermediate models are scaled to obtain the original modeling result of a target size.
5. The method according to claim 1, characterized in that, The step of adjusting the model parameters of the original neural network model based on the comparison results between the original modeling results and the sample modeling results until the target neural network model is obtained includes: Determine the cross-entropy loss and boundary loss of the original modeling result and the sample modeling result; Determine the weighted sum of the cross-entropy loss and the boundary loss; If the weighted summation result is less than the second threshold, adjust the model parameters; If the weighted summation result is greater than or equal to the second threshold, the original neural network model is determined as the target neural network model.
6. A building complex modeling device, characterized in that, include: The acquisition unit is used to acquire video and remote sensing images of the building complex to be modeled. The first determining unit is used to determine the virtual camera parameters of a virtual camera that matches the camera in the virtual scene based on the first camera parameters of the camera that captures the video image. The second determining unit is used to determine the target position of the building complex to be modeled in the virtual scene based on the correspondence between the first camera parameters and the virtual camera parameters. The generation unit is used to generate a target 3D model of the building complex to be modeled based on the remote sensing image. A processing unit is used to place the target 3D model at the target location; An adding unit is used to add texture data from the video image to the surface of the target 3D model; The generation unit further includes: The first input module is used to input the remote sensing image into the target neural network model, wherein the target neural network model is used to model the remote sensing image to obtain the target three-dimensional model; The third acquisition module is used to acquire the modeling results output by the target neural network model; The second determining module is used to determine the modeling result as the target three-dimensional model; The fourth acquisition module is used to acquire sample remote sensing images and sample modeling results of the sample remote sensing images before inputting the remote sensing images into the target neural network model. The extraction module is used to extract multiple sample pixel blocks from the sample remote sensing image; The deletion module is used to delete pixel blocks whose clarity is lower than a first threshold from a plurality of sample pixel blocks; The processing module is used to perform at least one of blurring, rotation, magnification and reduction processing on each of the remaining sample pixel blocks to obtain a first pixel block; The second input module is used to input the first pixel block as a training sample into the original neural network model to obtain the original modeling result output by the original neural network model. The adjustment module is used to adjust the model parameters of the original neural network model based on the comparison results between the original modeling results and the sample modeling results, until the target neural network model is obtained.
7. The apparatus according to claim 6, characterized in that, The first determining unit includes: The first acquisition module is used to acquire multiple pairs of matching points from video frames in the video image and the virtual scene, wherein each pair of matching points includes a two-dimensional point in the video frame and a three-dimensional point in the virtual scene; The second acquisition module is used to acquire the angle information and position information in the first camera parameters; The first mapping module is used to map the angle information and position information into the virtual scene according to the mapping relationship from the two-dimensional point to the three-dimensional point, so as to obtain the virtual camera parameters.
8. The apparatus according to claim 6, characterized in that, The second determining unit includes: The first determining module is used to determine the direction of the building complex to be modeled relative to the camera and the distance to the camera; The second mapping module is used to map the direction and distance to the virtual scene according to the correspondence between the first camera parameters and the virtual camera parameters, thereby determining the target position.
Citation Information
Patent Citations
Method and system for three-dimensional video monitor
CN101931790A
Video augmented reality method and device
CN111080704A
Substation three-dimensional model reconstruction method
CN111985161A
Video fusion system
CN112584060A