Computer program products used for 3D modeling
The computer program product for 3D modeling addresses the challenge of moving objects in UAV-based 3D modeling by identifying and removing moving regions using feature point detection and transformation matrices, ensuring accurate and user-friendly 3D model creation.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- DELTA ELECTRONICS INC(CN)
- Filing Date
- 2025-07-07
- Publication Date
- 2026-05-01
AI Technical Summary
Conventional 3D modeling using unmanned aerial vehicles (UAVs) is challenged by the introduction of noise in the 3D model due to moving objects, which are difficult to remove manually and current noise reduction methods are ineffective, especially for unusual appearances, leading to inaccurate models.
A computer program product and method for 3D modeling that includes feature point detection, region segmentation, feature point matching, and transformation matrix calculation to identify and replace moving object regions with featureless areas, ensuring accurate 3D modeling by removing moving objects before constructing the model.
The method effectively removes moving objects from 3D models, maintaining model integrity and accuracy by utilizing spatial and temporal recognition, and providing user-interactive functions for flexible object removal, compatible with commercial modeling software.
Smart Images

Figure 0007854554000013 
Figure 0007854554000014 
Figure 0007854554000015
Abstract
Description
[Technical Field]
[0001] This invention relates to 3D modeling technology, and more particularly to a computer program product used for 3D modeling and a method for removing its moving object. [Background technology]
[0002] Unmanned aerial vehicles (UAVs) have the ability to overcome terrain constraints and perform missions, and can also assist professional personnel in carrying out tasks, effectively improving efficiency, which has led to a booming market for UAVs in recent years. They are known to be rapidly developing in fields such as construction / engineering / mining, energy / transportation / public infrastructure, and precision agriculture, and in particular, 3D modeling of terrain, landscapes, and buildings using UAVs has become one of the mainstream applications in recent years.
[0003] Conventional 3D modeling of unmanned aerial vehicles (UAVs) involves using an UAV equipped with a high-resolution camera to take a series of images of a stationary object or scenery from different positions and angles. The 3D coordinates of the UAV are estimated using feature points from each image, and then the measured 3D coordinates are inversely estimated to create a 3D model.
[0004] Figure 1 is a schematic diagram of 3D modeling using an unmanned aerial vehicle. As shown in Figure 1, the unmanned aerial vehicle flies from a first position 100 in the air along a flight path 102 to a second position 104, and then along a flight path 106 to a third position 108. Images are taken of the target object 110 at these three positions, and the first image 120, second image 130, and third image 140, which are used for modeling, are obtained. Feature points A112 and B114 on the target object 110 are projected by multiple rays 116 onto points a1122 and b1114 in the first image 120, points a2132 and b2134 in the second image 130, and points a3142 and b3144 in the third image 140, respectively. Next, by calculating the characteristic similarity of each image and back-projecting the multi-angle light rays into 3D space, the 3D coordinates of feature points A112 and B114 can be estimated and represented in the form of a 3D point cloud.
[0005] The 3D coordinate calculation described above is based on the principle that the projections of the conjugate points of images taken at different angles intersect at the same point in 3D space. However, since the above shooting process needs to be sustained for a certain period of time (for example, a 30-meter-long bridge requires approximately 10 minutes of shooting), if a moving object (e.g., a train or high-speed rail) passes by during shooting, the ray projections of these moving objects in different frame images cannot intersect at the exact location, resulting in noise in the constructed 3D model.
[0006] Figure 2 illustrates the source of noise in 3D modeling. As shown in Figure 2, the unmanned aerial vehicle (UAV) flies along a flight path 202 from one position 200 in the air to another position 204, taking images of the same object at both positions to acquire a first image 214 and a second image 224 for modeling, respectively. During this period, the object moves from point A 210 to point A' 220. Before the UAV moves, point A 210 is projected onto point a 216 on the first image 214 via ray 212. Before the UAV moves, point A' 220 is projected onto point a' 226 on the second image 224 via ray 222. Based on the 3D modeling principle described above, the image properties of points a 216 and a' 226 are similar, so when they are back-projected into 3D space, they intersect at point A* 230. However, since point A*230 is not at the same position as point A210 or point A'220, noise is generated.
[0007] In reality, since it's impossible to restrict the entry and exit of objects from the shooting scene when using unmanned aerial vehicles (UAVs), ensuring the accuracy of the final 3D model requires manually removing noise one by one after 3D modeling. However, this removal process is tedious and time-consuming. Currently, there are methods to perform noise reduction on 3D models after modeling, but these depend on the projection relationships obtained in the initial 3D modeling, and therefore fail in certain situations. For example, if the input image contains moving objects with unusual appearances (such as being too large), initial modeling errors may occur, potentially rendering subsequent noise reduction ineffective for the model.
[0008] As the demand for various 3D modeling applications using unmanned aerial vehicles is expected to increase in the future, designing moving object removal solutions that can be used in 3D modeling is becoming an increasingly important challenge. [Overview of the Initiative] [Problems that the invention aims to solve]
[0009] The purpose of this disclosure is to provide a computer program product for use in 3D modeling and a method for removing the moving object thereof. [Means for solving the problem]
[0010] The moving object removal method according to this disclosure is a moving object removal method used for 3D modeling, and includes the steps of: detecting a plurality of feature points for each of the original images in the original image sequence by a feature point detection process; dividing each of the original images into a plurality of regions by a region segmentation process; identifying a target region and one or more non-target regions from the plurality of regions in each of the original images; determining whether each of the non-target regions of the original images in two consecutive frames is a moving region by a feature point matching process based on a plurality of feature points in the original images of two consecutive frames of the original image sequence; and obtaining a still image sequence by replacing the regions determined to be moving regions in each of the non-target regions of the original images with featureless regions.
[0011] One embodiment of the moving object removal method according to the present disclosure includes the steps of: determining whether each of the non-target regions of the original images of two consecutive frames is a moving region by a feature point matching process based on a plurality of feature points in the original images of two consecutive frames of the original image sequence; obtaining a plurality of matching pairs of feature points by a feature point matching process based on a plurality of feature points in the original images of two consecutive frames; estimating a transformation matrix based on a plurality of matching pairs within the target region of the original images of two consecutive frames; calculating the average projection error for each of the non-target regions of the original images of two consecutive frames based on a plurality of matching pairs within the non-target region and the transformation matrix; and determining whether each of the non-target regions is a moving region by comparing the average projection error with a threshold for each of the non-target regions of the original images of two consecutive frames.
[0012] One embodiment of the moving object removal method according to the present disclosure further includes, before the step of estimating a transformation matrix, a step of examining the homography relationships of the target region in the original images of two consecutive frames based on the number of matching pairs within the target region in the original images of two consecutive frames.
[0013] One embodiment of the moving object removal method according to the present disclosure includes the steps of: calculating the average projection error for each non-target region of the original images of two consecutive frames, based on a plurality of matching pairs and a transformation matrix within the non-target region; for each matching pair within the non-target region, projecting the first feature point in the matching pair onto the projection point using the transformation matrix; calculating the Euclidean distance between the projection point and the second feature point in the matching pair; and using the average of the Euclidean distances of the matching pairs as the average projection error.
[0014] One embodiment of the moving object removal method according to the present disclosure further includes, before the step of comparing the mean projection error with a threshold, the steps of calculating the mean projection error of a target region and the standard deviation corresponding to the target region based on matching pairs and transformation matrices of the target region of two consecutive frames of original images, and determining a threshold based on the mean projection error and standard deviation of the target region of the two consecutive frames of original images.
[0015] In one embodiment of the moving object removal method according to this disclosure, each feature point has a feature descriptor.
[0016] One embodiment of the moving object removal method according to the present disclosure includes a feature point matching procedure which involves comparing these feature descriptors in the original images of two consecutive frames to obtain a plurality of matching pairs for feature points.
[0017] One embodiment of the moving object removal method according to this disclosure includes a feature point detection process that uses a scale-invariant feature transformation (SIFT) algorithm.
[0018] One embodiment of the moving object removal method according to this disclosure includes a region segmentation process that uses a vision transformer adapter (ViT-Adapter).
[0019] One embodiment of the moving object removal method according to this disclosure includes a feature point matching process that uses a brute-force matcher (BFMatcher).
[0020] One embodiment of the present disclosure further provides a computer program product for 3D modeling, including a user interface module, a 3D modeling module, and a moving object removal module. The user interface module provides a user interface. The 3D modeling module creates a 3D model. The moving object removal module executes a moving object removal method. After being loaded into a computer, the computer program product can execute the following steps: obtaining an original image sequence; when receiving a moving object removal instruction from the user interface, calling the moving object removal module to execute the moving object removal method, and driving the 3D modeling module to create a 3D model using a series of still images obtained by executing the moving object removal method; and when receiving a direct modeling instruction from the user interface, driving the 3D modeling module to create a 3D model using the original image sequence.
[0021] One embodiment of the computer program product according to the present disclosure is that the user interface is a graphical user interface (GUI) that presents the region division result of the region division process and allows the user of the computer to select a target region on the region division result.
[0022] One embodiment of the computer program product according to the present disclosure is that the moving object removal module further detects the area ratio of the moving regions in the original image sequence, checks the number determined to not include moving regions in the original image sequence, and in response to the area ratio of the moving regions in the original image sequence exceeding a first specified threshold and the number determined to not include moving regions in the original image sequence not reaching a second specified threshold, the moving object removal module notifies the user interface module to present an exception message to the user interface.
[0023] Another embodiment of the present disclosure further provides a computer program product for 3D modeling. After being loaded by a computer, the computer program product can provide a graphical user interface (GUI). The graphical user interface includes an image import section, a moving object removal section, and a modeling section. The image import section is for a user to input a specified path to import an original image sequence. The moving object removal section is for a user to input a removal command. The modeling section is for a user to input a modeling command. In response to the received removal command, the computer program product causes the computer to execute a moving object removal method on the original image sequence and obtain a static image sequence. In response to the received modeling command, the computer program product causes the computer to select and use one of the original image sequence and the static image sequence based on the removal command to create a 3D model.
[0024] One embodiment of the computer program product according to the present disclosure is that the graphical user interface further includes a removal progress display section for presenting the processing progress status of the moving object removal method.
[0025] One embodiment of the computer program product according to the present disclosure is that the graphical user interface further includes a target area selection section for presenting a region division result for a user to select a target area on the region division result.
[0026] One embodiment of the computer program product according to the present disclosure is that the graphical user interface further includes an image display section for presenting at least one of the original image sequence and the static image sequence.
[0027] In one embodiment, the graphical user interface further includes a removal result display unit. In response to the area ratio of the moved region in the original image sequence exceeding a first specified threshold, and the number of images determined not to contain the moved region in the original image sequence not reaching a second specified threshold, the removal result display unit presents an exception message. In response to the acquired still image sequence, the removal result display unit presents a success message.
[0028] In one embodiment, after an exception message is presented, the graphical user interface further includes a manual removal section for the user to manually remove the moving object.
[0029] In one embodiment, the exception message is configured to guide the user to add more original images to the original image sequence.
[0030] The solutions used to remove moving objects from 3D models provided by various embodiments of this disclosure simultaneously possess spatial (object) and temporal (motion) recognition capabilities, ensuring the integrity and accuracy of the 3D model by removing only moving objects in the image while retaining information on stationary objects. Furthermore, the user interface of this disclosure provides flexibility in enabling or disabling moving object removal, and offers user-interactive functions different from conventional 3D modeling, such as allowing the user to select a target area on the region segmentation results and view the removal results. Compared to some denoising processes performed after modeling, removing objects before modeling can further prevent noise from being introduced into the 3D model. The removed images are compatible with various commercial modeling software and have market applicability.
[0031] This disclosure can be better understood from the following description of exemplary embodiments and the accompanying drawings. Furthermore, it should be understood that in the flowcharts of this disclosure, the execution order of each block may be changed, and / or some blocks may be modified, deleted, or combined. [Brief explanation of the drawing]
[0032] [Figure 1] Figure 1 is a schematic diagram of 3D modeling using an unmanned aerial vehicle. [Figure 2] Figure 2 shows the causes of noise in 3D modeling. [Figure 3] Figure 3 is a flowchart illustrating a method for removing a moving object for 3D modeling according to one embodiment of the present disclosure. [Figure 4] Figure 4 is a schematic diagram showing the original images of two frames captured by an unmanned aerial vehicle flying from right to left, according to one embodiment of the present disclosure. [Figure 5] Figure 5 is a schematic diagram showing the feature point detection results of the original image according to one embodiment of the present disclosure. [Figure 6] Figure 6 is a schematic diagram showing the region segmentation result of the original image according to one embodiment of the present disclosure. [Figure 7] Figure 7 is a schematic diagram showing a matching pair of feature points according to one embodiment of the present disclosure. [Figure 8] Figure 8 is a flowchart showing more detailed steps for determining the moving region according to one embodiment of the present disclosure. [Figure 9A] Figure 9A is a conceptual diagram illustrating a method for determining a movement region using a transformation matrix of a target region, according to one embodiment of the present disclosure. [Figure 9B] Figure 9B is a conceptual diagram illustrating a method for determining a movement region using a transformation matrix of a target region, according to one embodiment of the present disclosure. [Figure 10] Figure 10 shows an example of a still image consisting of two frames according to one embodiment of the present disclosure. [Figure 11] Figure 11 is a system block diagram showing a computer system for 3D modeling according to one embodiment of the present disclosure. [Figure 12] Figure 12 is a flowchart illustrating the basic operation of a 3D modeling program according to one embodiment of the present disclosure. [Figure 13]Figure 13 is an interface block diagram showing a user interface that can be provided by a 3D modeling program after being loaded by a computer, according to one embodiment of the present disclosure. [Figure 14] Figure 14 shows a schematic example of the user interface provided by the 3D modeling program when the above-mentioned exception event is triggered. [Modes for carrying out the invention]
[0033] The following description enumerates various embodiments of the present invention, but is not intended to limit the scope of the invention. The actual scope of the invention will be determined by the scope of the patent application.
[0034] In each of the embodiments listed below, identical or similar elements or components are represented by the same reference numeral.
[0035] The ordinal terms such as "first," "second," and "third" used in the patent claims are used solely for explanatory purposes and do not imply any priority relationship between them.
[0036] The following description relating to embodiments of an apparatus or system also applies to embodiments of a method, and vice versa.
[0037] Figure 3 is a flowchart of a moving object removal method 300 for 3D modeling according to one embodiment of the present disclosure. As shown in Figure 3, the method 300 may include steps S302 to S310.
[0038] In step S302, the feature point detection process detects multiple feature points within each original image in the original image sequence.
[0039] The source image sequence may be a series of images taken from different angles by a camera mounted on an unmanned aerial vehicle, in which moving objects have not been removed. This disclosure does not limit the type of unmanned aerial vehicle or camera, or the shooting scene.
[0040] Feature points can be understood as the most identifiable locations or regions within an image, and are typically points where characteristics such as changes in brightness, texture, edges, color, and shape are prominent (e.g., possessing local extrema). It should be understood that changes in the field of view (or the position of the unmanned aerial vehicle) result in different location distributions of feature points within different source images in the source image sequence, regardless of whether the captured scene is completely still. In one embodiment, each feature point has a feature descriptor, which can be represented in the form of a feature vector symbolizing the gradient, direction, and feature intensity around the feature point.
[0041] Figure 4 is a schematic diagram showing two frames of original images, namely the first original image 400 and the second original image 410, taken by an unmanned aerial vehicle flying from right to left according to one embodiment of the present disclosure. As shown in Figure 4, the first original image 400 includes a train 402, a bridge 404, and the ground 406, and the second original image 410 includes a train 412, a bridge 414, and the ground 416. In this example, it is assumed that trains 402 and 412 are the same moving train photographed at different time points, and that the bridge and the ground are stationary. From the first original image 400 and the second original image 410, it can be observed that the appearance of train 412 is significantly different from that of electric train 402.
[0042] Figure 5 is a schematic diagram showing the first feature point detection result 500 and the second feature point detection result 510 obtained after step S302 of the first source image 400 and the second source image 410 according to one embodiment of the present disclosure. As shown in Figure 5, the upper left portion of the first feature point detection result 500 has feature points on a stationary bridge (such as feature point 502) and feature points located on a moving train (such as feature point 504). However, in the second feature point detection result 510, the upper left portion no longer has feature points on a stationary bridge, and instead has feature points on a moving train, such as feature point 512. That is, the feature points on a stationary bridge are partially obscured by the moving train.
[0043] The feature point detection process in step S302 can be carried out using various feature detection algorithms or corner detectors, such as the Speeded Up Robust Features (SURF) algorithm, the Accelerated-KAZE algorithm, the Harris corner detector, Features from Accelerated Segment Test (FAST), Binary Robust Invariant Scalable Keypoints (BRISK), etc. In one embodiment, the feature point detection process is carried out using the Scale Invariant Feature Transform (SIFT) algorithm. This concept involves performing a convolution operation on the image using Gaussian filters of different scales to obtain Gaussian variance based on a continuous Gaussian blur image, thereby obtaining feature points that are invariant in scaling and rotation. Each feature point has a feature descriptor, which can be represented in the form of a feature vector, and can also be represented in the form of a feature vector that symbolizes the gradient, direction, and feature intensity around the feature point. In actual operation, approximately 50,000 SIFT feature points can be detected from a single 3840x2160 image, and each SIFT feature point contains 128-dimensional information representing the magnitude and direction of the gradient within the region surrounding the feature point.
[0044] Referring again to Figure 3, in step S304, each original image is divided into multiple regions by the region segmentation process.
[0045] More specifically, the region segmentation process assigns an index value to each pixel in the original image, and pixels with the same index value form a single region. For example, pixels with index value "1" constitute the first region, pixels with index value "2" constitute the second region, and so on. In other words, pixels in the first region have index value "1", pixels in the second region have index value "2", and so on. This allows subsequent steps to identify each region based on its index value and perform the corresponding processing.
[0046] Figure 6 is a schematic diagram showing the first region division result 600 and the second region division result 610 of the first original image 400 and the second original image 410, respectively, obtained after step S304, according to one embodiment of the present disclosure. As shown in Figure 6, the first region division result 600 includes a train region 602, a bridge region 604, and a ground region 606, and the second region division result 610 includes a train region 612, a bridge region 614, and a ground region 616. The terms “train region,” “bridge region,” and “ground region” used above are named to make the correspondence between the region division result and the original image easier for the reader to understand, and it should be noted that in reality these regions are identified by index values. The region division process does not specify or identify what kind of object these regions belong to, such as trains, bridges, or ground.
[0047] The region segmentation process in step S304 is carried out using algorithms used for region segmentation, such as simple linear iterative clustering (SLIC), graph-based segmentation (Felzenszwalb's Graph-Based Segmentation), Quick Shift, or Region Growing. Alternatively, a region segmentation process built using a convolutional neural network (CNN)-based machine learning model such as U-Net, SegNet, or Fully Convolutional Network (FCN) may be used. The training process may include acquiring labeled data, selecting a loss function, and constructing an optimization algorithm, and various conventional methods may be employed, but are not limited to these. Furthermore, the region segmentation model may be trained locally, or it may be trained first on another computer device (e.g., a server) and then the trained region segmentation model may be obtained via a network (e.g., downloaded from the cloud), storage medium (e.g., external hardware), or other communication interface (e.g., USB), but are not limited to these. In one embodiment, the domain segmentation process is performed using a Visual Transformer Adapter (ViT-Adapter). The Visual Transformer Adapter is a deep learning-based classifier whose concept is to improve the performance and versatility of a model by adding small neural network modules called "adapters" to different layers of a pre-trained Visual Tramp model.
[0048] Referring again to Figure 3, in step S306, the target region and one or more non-target regions (i.e., regions other than the target region) are determined from these regions of each original image.
[0049] In one embodiment, the target region can be selected and defined by the user from the region segmentation results presented in a graphical user interface (GUI). In another embodiment, a machine learning model for identifying the target region, i.e., a target region identification model, can be pre-trained, and then in step S306, the trained target region identification model can be used to identify the target region from these regions in each original image. The target region identification model can be a convolutional neural network (CNN) based model, and the training process can include, but is not limited to, various conventional methods such as acquiring label data, selecting a loss function, and constructing an optimization algorithm. Furthermore, the target region identification model can also be trained locally, and may be trained first on another computer device (e.g., a server) and then obtained via a network (e.g., downloaded from the cloud), storage medium (e.g., external hardware), or other communication interface (e.g., USB), but is not limited to these.
[0050] In a single modeling task, the target object to be modeled is known in advance. For example, in the modeling of a bridge, the target is the bridge itself and the surrounding terrain. Taking Figure 6 as an example, the bridge region 604 can be selected as the target region. Since the original images in the original image sequence have a high degree of overlap with each other (generally 50% or more), the index value of the original bridge region 604 is further extended to the bridge region 614 of the next original image based on the overlap ratio (for example, only regions that overlap by 50% or more are considered), and the bridge region 614 is considered the target region of the second original image 410. In other words, by performing step S306 once, the target regions of the remaining original images in the original image sequence can be determined by inference.
[0051] Referring again to Figure 3, in step S308, based on the feature points in the original images of two consecutive frames in the original image sequence, a feature point matching process is used to determine whether each non-target region of the original images of the two consecutive frames is a moving region (i.e., a region corresponding to a moving object).
[0052] In one embodiment, the feature point matching process includes comparing feature descriptors in two images to find similar feature points in the two images. Two feature points whose feature descriptors match form a matching pair. As described above, a feature descriptor can be represented in the form of a feature vector that symbolizes the gradient, direction, and feature intensity around a feature point. Thus, the feature point matching process may further include, but is not limited to, the calculation of distance or similarity between two feature vectors, such as Euclidean distance, Manhattan distance, cosine similarity, or magnitude used to represent other distances or similarities.
[0053] Figure 7 is a schematic diagram showing a feature point matching pair of a first source image 700 and a second source image 710 according to one embodiment of the present disclosure. The first source image 700 and the second source image 710 correspond to the first feature point detection result 500 and the second feature point detection result 510 in Figure 5, respectively, but not all feature points depicted in Figure 5 appear in Figure 7. In the example in Figure 7, only the nine similar feature points found in the feature point matching process in step S308, i.e., the nine matching pairs, are shown, including the first matching pair consisting of feature point 701 in the first source image 700 and feature point 711 in the second source image 710, the second matching pair consisting of feature point 703 in the first source image 700 and feature point 713 in the second source image 710, the third matching pair consisting of feature point 705 in the first source image 700 and feature point 715 in the second source image 710, the fourth matching pair consisting of feature point 706 in the first source image 700 and feature point 716 in the second source image 710, the fifth matching pair consisting of feature point 708 in the first source image 700 and feature point 718 in the second source image 710, and four other unsigned matching pairs. In these matching pairs, the first, second, and third matching pairs are located on the moving train 702, while the fourth and fifth matching pairs are located on the stationary bridge. Therefore, the displacements between feature points 701 and 711, 703 and 713, and 705 and 715 are greater than those between feature points 706 and 716, or between feature points 708 and 718. Thus, it can be determined that the region corresponding to the train 702 is a moving region.
[0054] The feature point matching process in step S308 can be performed using nearest neighbor matching, Random Sample Consensus (RANSAC), Kanade-Lucas-Tomasi feature tracker, or other similar algorithms. In one embodiment, the feature point matching process is performed using a brute-force matcher (BFMatcher) that searches for the best match by calculating the distance or similarity between two sets of feature descriptors (or feature vectors).
[0055] Figure 8 is a flowchart showing a more detailed step of determining the movement area in step S308 according to one embodiment of the present disclosure. As shown in Figure 8, step S308 may further include steps S802 to S808.
[0056] In step S802, based on the feature points in the original images of two consecutive frames, a feature point matching process is performed to obtain multiple matching pairs of these feature points. Taking Figure 7 as an example, after step S802, nine matching pairs of the first original image 700 and the second original image 710 can be obtained.
[0057] In step S804, the transformation matrix is estimated based on matching pairs within the target region of two consecutive frames of the original images.
[0058] The transformation matrix can represent the transformation relationship of the positions of the feature points and other pixels of each matching pair between the two original images. Therefore, step S804 can be understood as finding the transformation relationship of pixel positions within the target region of two consecutive frames of the original images. In subsequent steps, the shifted regions can be found by checking whether the transformation relationship of each non-target region differs significantly from that of the target region.
[0059] In step S806, the average projection error is calculated for each non-target region of the original image of two consecutive frames, based on the matching pair and transformation matrix in the non-target region.
[0060] The average projection error can represent the degree of dissimilarity between the pixel position transformation relationships of the non-target region and the target region. In other words, the larger the average projection error, the greater the difference between the pixel position transformation relationships representing the non-target region and the target region.
[0061] In step S808, for each non-target region of the original image of two consecutive frames, it is determined whether the non-target region is a moving region by comparing the average projection error with a threshold.
[0062] Specifically, if the average projection error exceeds a threshold, it means that the difference between the pixel position transformation relationship of the non-target region and the pixel position transformation relationship of the target region is considerably large, and therefore the non-target region is determined to be a moving region. The threshold can be a predefined numerical value or a variable determined by a specific calculation.
[0063] Figures 9A and 9B show the transformation matrix H of the target region according to one embodiment of the present disclosure. 3×3 This is a conceptual diagram showing how to determine the moving region using M. As shown in Figures 9A and 9B, the first source image 900 includes regions 902, 904, and 906, which correspond to regions 912, 914, and 916 of the second source image 910, respectively. Regions 902 (and 912), 904 (and 914), and 906 (and 916) have corresponding index values 0, 1, and 2, respectively, denoted by k. In this example, regions 902 and 912 with index value k=0 are set as target regions, and the other regions are set as non-target regions. Also, M kLet k represent the number of matching pairs in the region with index value k, and the i-th feature point in the region with index value k in the first original image 900 and the second original image 910 is
number
number
number
number
[0064] Transformation matrix H 3x3 The estimation of this is conceptually based on one optimal transformation matrix H 3x3 Find each feature point within the target region (i.e., i=1 to M) k The goal is to satisfy the following <Equation 1> as much as possible.
number
[0065] Next, in step S806, feature points in each non-target region (i.e., region with index value k≠0) on the first original image 900
number
Number
[0066] As shown in FIG. 9B, the feature points in the non-target region 904 of the first original image 900 can be projected onto the
Number
[0067] Subsequently, the Euclidean distance between the projection points
Number
Number
[0068] Taking FIG. 9B as an example, the average projection error value E of the non-target regions 904 and 914 with index value k = 1 1 is calculated as follows:
Number
[0069] In one embodiment, the estimation of the transformation matrix H 3×3 can include using an algorithm associated with the least squares method, least absolute deviations (LAD), least squares support vector machine (LS-SVM), polynomial fitting, or any other function fitting, but the present disclosure is not limited thereto.
[0070] In one embodiment, before step S808, the mean projection error and corresponding standard deviation of the target region, i.e., the mean and standard deviation of the projection error of each feature point within the target region, can be calculated based on the matching pair and transformation matrix in the target region of the original images of two consecutive frames. More specifically, the standard deviation σ is calculated as shown in <Equation IV> below:
number
[0071] However, based on the calculated mean projection error and standard deviation, a threshold can be determined to be used in step S808 for comparison with the mean projection error. In a preferred embodiment, the threshold is the mean projection error E 0 It is set as the sum of twice the standard deviation. For example, the mean projection error E of the target region. 0 If the value is 75 and the standard deviation is 50.4, then the threshold is 75 + 2 × 50.4 = 175.8. Taking Figures 9A and 9B as examples, the mean projection error E of target regions 902 and 912 is 0 And the corresponding standard deviation is calculated, and the corresponding threshold T is calculated accordingly. p This can be determined. The mean projection error value E of the non-target regions 904 and 914 can be determined. 1 is T p If larger, non-target regions 904 and 914 are determined to be moving regions.
[0072] In one embodiment, before step S804, the homography relationship of the target region in the original images of two consecutive frames can be checked based on the number of matching pairs within the target region in the original images of the two consecutive frames. If it is determined that the target region has a homography relationship in the original images of the two consecutive frames, the condition for estimating the transformation matrix in step S804 is met. Otherwise, steps S804 to S808 are skipped. That is, it is assumed that no moving region was detected from the original images of the two consecutive frames.
[0073] The aforementioned homography relationship refers to a reversible transformation from a real projective plane to a projective plane. When photographing a stationary object at a long distance, the original images of two consecutive frames can be considered to lie on approximately the same plane. Therefore, theoretically, the stationary region (represented by the target region) of the original image of the first frame should be transformable to the corresponding region of the original image of the second frame based on the projection relationship.
[0074] Referring again to Figure 3, in step S310, a series of still images are obtained by replacing the areas determined to be moving regions within the non-target region of each original image with featureless regions.
[0075] The aforementioned featureless region means that the pixels within this region do not have feature values. Generally, regions of pure color (such as white) can be used as featureless regions.
[0076] Figure 10 shows examples of still images 1000 and 1010 of two frames obtained after step S310 according to one embodiment of the present disclosure. As shown in Figure 10, in still images 1000 and 1010, the areas presented in the outline of the train (e.g., the front of the train and the passenger cars) are replaced with featureless areas 1002 and 1012 of pure white color.
[0077] Figure 11 is a system block diagram showing a computer system 1100 for 3D modeling according to one embodiment of the present disclosure. As shown in Figure 11, the computer system 1100 may include a processing unit 1102 and a storage device 1104.
[0078] The computer system 1100 may be, and is not limited to, a personal computer (e.g., a desktop computer or laptop computer) running an operating system (e.g., Windows, Mac OS, Linux, UNIX, etc.), a server computer, or a mobile device such as a tablet computer or smartphone.
[0079] The processing unit 1102 may include any one or more general-purpose processors or dedicated processors, or combinations thereof, for executing instructions. In various embodiments of the present disclosure, the processing unit 1102 is configured to perform the aforementioned moving object removal method, such as method 300. In one embodiment, the processing unit 1102 may further include a central processing unit (CPU) and a graphics processing unit (GPU), although not shown in Figure 11. A GPU is an electronic circuit specifically designed for computer graphics calculations and image processing, and is therefore more efficient than a general-purpose CPU in computer graphics calculations and image processing. Accordingly, in various embodiments of the present disclosure, appropriate tasks can be assigned depending on the characteristics of the CPU and GPU. For example, tasks for acquiring data or communicating with other devices can be assigned to the CPU, and tasks related to computer graphics and image processing can be assigned to the GPU. In further embodiments, the processing unit 1102 may further include a neural network processing unit (NPU) specifically optimized for deep learning, although not shown in Figure 11. Compared to a GPU, an NPU may have greater computing performance advantages in the operation of the aforementioned region partitioning models and / or target region identification models. Therefore, in this embodiment, tasks including these machine learning models can be assigned to the NPU.
[0080] The storage device 1104 may include, but is not limited to, devices (such as hard disk drives (HDDs), solid-state drives (SSDs), or optical discs) and / or combinations thereof of volatile memory (such as random-access memory (RAM), dynamic random-access memory (DRAM), or static random-access memory (SRAM)) and / or any one or more types of non-volatile memory (such as read-only memory, electrically erasable programmable read-only memory (EEPROM), flash memory, or non-volatile random-access memory (NVRAM)). In various embodiments of the disclosure, the storage device 1104 is used to store a program corresponding to the aforementioned moving object removal method. Once the processing unit 1102 loads this program from the storage device 1104, it can execute the aforementioned moving object removal method.
[0081] In the embodiment shown in Figure 11, the storage device 1104 stores a computer program product for 3D modeling, which is hereby referred to as the “3D modeling program” 1110. As shown in Figure 11, the 3D modeling program 1110 may include a user interface module 1112, a 3D modeling module 1114, and a moving object removal module 1116. The user interface module 1112 is used to provide a user interface, the 3D modeling module 114 is used to create a 3D model, and the moving object removal module is used to perform the aforementioned moving object removal method, for example, method 300. The processing unit 1102 can load the 3D modeling program 1110 from the storage device 1104 and execute it to perform the user interface module 1112, the 3D modeling module 1114, and the moving object removal module 1116, as well as other basic operations of the 3D modeling program 1110.
[0082] In one embodiment, the processing unit 1102 can be coupled to a display device and display a user interface provided by a user interface module 1112 thereon. The display device can be any device used to display visible information, such as an LCD display, LED display, OLED display, or plasma display, but the present invention is not limited to these.
[0083] The user interface provided by the user interface module 1112 may be a graphical user interface (GUI), a command-line interface (CLI), a touch interface, or a voice interface, but is not limited to these.
[0084] The 3D modeling module 1114 may include, but is not limited to, any known 3D modeling techniques, such as stereo-enabled, point cloud reconstruction, point cloud processing, and 3D rendering.
[0085] Figure 12 is a flowchart of a basic operation 1200 of a 3D modeling program 1110 according to one embodiment of the present disclosure. As shown in Figure 12, the basic operation 1200 of the 3D modeling program 1110 may include steps S1202 to S1210.
[0086] In step S1202, the original image sequence is obtained.
[0087] As described above, the original image sequence is a series of images taken from different viewpoints by an imaging device mounted on an unmanned aerial vehicle, and moving objects have not yet been removed. The original image sequence can be obtained via a network, storage medium, or various other communication interfaces, but is not limited to these.
[0088] In step S1204, the user interface module 1112 receives commands from the user of the computer system 1100 via the user interface. In response to the receipt of a moving object removal command, the process proceeds to step S1206. In response to the receipt of a direct modeling command, the process proceeds to step S1210.
[0089] In step S1206, the moving object removal module 1116 is called to execute the moving object removal method (method 300) described above.
[0090] In step S1208, the 3D modeling module 1114 is driven using the still image sequence obtained by executing the moving object removal method to create a 3D model.
[0091] In step S1210, the 3D modeling module 1114 is activated to create a 3D model using the source images from this sequence.
[0092] In one embodiment, the user interface provided by the user interface module 1112 is a graphical user interface (GUI) that presents the region division results of the region division process (for example, the first region division result 600 and the second region division result 610 shown in Figure 6), allowing the user of the computer system 1100 to select a target region from the region division results. This allows the moving object removal module 1116 to know the target region in step S306 and determine other regions other than the target region as non-target regions. Furthermore, checkboxes, toggle switches, drop-down lists, or other similar GUI elements or widgets are provided on this graphical user interface to allow the user of the computer system 1100 to select a moving object removal command or a direct modeling command, but this disclosure is not limited to the specific graphic design arrangement of the graphical user interface.
[0093] If an object's movement speed is too slow, it will frequently cover the target area in the source image sequence, potentially affecting the accuracy of 3D modeling. In view of this, in one embodiment, the moving object removal module 1116 further detects the area percentage of the moving area in the source image sequence and checks the number of objects determined not to contain the moving area in the source image sequence. In response to the area percentage of the moving area in the source image sequence exceeding a first specified threshold and the number of objects determined not to contain the moving area in the source image sequence not reaching a second specified threshold, the moving object removal module 1116 notifies the user interface module 1112 to present an exception message to the user interface. The content of the exception message may be set to guide the user to manually remove the moving object or to add more additional source images (e.g., images of slow-moving objects that do not exist, taken at other times) to facilitate 3D modeling.
[0094] Figure 13 is an interface block diagram showing a user interface 1300 that can be provided by a 3D modeling program after being loaded by a computer, according to one embodiment of the present disclosure. Similarly, Figure 14 is a schematic example of a user interface 1400 provided by the 3D modeling program when the above-described exception event is triggered. User interface 1300 includes at least an image import unit 1301, a moving object removal unit 1302, and a modeling unit 1305, as shown in Figure 13. Similarly, user interface 1400 includes at least an image import unit 1401, a moving object removal unit 1402, and a modeling unit 1405, as shown in Figure 14.
[0095] The image import units 1301 and 1401 allow the user to input a specified path, for example, "C: / Users / 3DModeling / RawData", or another similar path, and import the original image sequence from this path. The image import units 1301 and 1401 can be implemented, for example, through a file selection dialog, a text box, file drag and drop, a tree view, or other similar GUI elements or widgets, but are not limited to these.
[0096] The moving object removal units 1302 and 1402 are for the user to input a removal command, which is used to indicate whether a moving object removal method (e.g., moving object removal method 300 in Figure 3) is valid and to trigger the execution of the moving object removal method. In response to receiving a removal command entered by the user via the moving object removal units 1302 and 1402, the 3D modeling program causes the computer to execute the moving object removal method on the source image sequence imported from a specified path (entered by the user via the image import units 1302 and 1402). The moving object removal units 1302 and 1402 can be implemented, for example, by checkboxes, toggle switches, drop-down menus, or other similar GUI elements or widgets combined with confirmation or start buttons, but are not limited to these.
[0097] Modeling units 1305 and 1405 are for the user to input modeling commands, which are used to trigger the 3D modeling process. More specifically, if a remove command indicates that a moving object removal method is enabled, a 3D model is created using a still image sequence. Conversely, if a remove command indicates that a moving object removal method is not enabled, a 3D model is created using the original image sequence. Modeling units 1305 and 1405 can be implemented, for example, by a confirmation button, a start button, a toolbar button, a menu option, or other similar GUI elements or widgets, but are not limited to these.
[0098] In one embodiment, the user interfaces 1300 and 1400 may further include removal progress display units 1303 and 1403. The removal progress display units 1303 and 1403 are for displaying the processing progress of the moving object removal method, which may be, for example, a completion rate such as 20%, 50%, or 90%, or a text representation of a status such as "processing" or "completed." The removal progress display units 1303 and 1403 may be implemented as, for example, a progress bar, a circular progress bar, a digital display, a status label, a timeline, or other similar GUI elements or widgets, but are not limited to these.
[0099] In one embodiment, the user interface 1300 may further include a target area selection unit 1306. The target area selection unit 1306 is used to present area division results (such as the first area division result 600 and the second area division result 610 shown in Figure 6) and to allow the user to select a target area on these area division results. As an exemplary example, the selection / deletion unit 1407 of the user interface 1400 is used to present the above-mentioned area division results and to allow the user to select a target area on these area division results.
[0100] In one embodiment, the user interfaces 1300 and 1400 may further include image display units 1308 and 1408. The image display units 1308 and 1408 are used to present a sequence of source images and / or a sequence of still images.
[0101] In one embodiment, user interfaces 1300 and 1400 may further include removal result display units 1304 and 1404. If the execution of the moving object removal method is successful and in response to the acquired still image sequence, the removal result display units 1304 and 1404 display a success message. The success message is used to inform the user that moving objects (or moving regions) in the original image sequence have been removed, and its content may be, for example, "Removal of 13 moving regions out of 105 images was successful." In response to an exception event, i.e., if the area ratio of the moving region in the original image sequence exceeds a first specified threshold and the number of original images determined not to contain a moving region does not reach a second specified threshold, the removal result display units 1304 and 1404 display an exception message. The exception message is used to inform the user that the moving object removal method encountered an exception event (e.g., a moving object covers a target region) and was unsuccessful. In a further embodiment, the content of the exception message may be configured to guide the user to add more additional source images (e.g., images taken at other times that do not contain slowly moving objects) to the source image sequence. In another embodiment, the content of the exception message may be configured to guide the user to manually remove the moving object. After this exception message is presented, the user interface 1300 may further include a manual removal unit 1307 for the user to manually remove the moving object. As an exemplary example, a selective removal unit 1407 of the user interface 1400 may be used to allow the user to manually remove the moving object after this exception message is presented.
[0102] The solutions used to remove moving objects from 3D models provided by various embodiments of this disclosure simultaneously possess spatial (object) and temporal (motion) recognition capabilities, ensuring the integrity and accuracy of the 3D model by removing only moving objects in the image while retaining information on stationary objects. Furthermore, the user interface of this disclosure provides flexibility in enabling or disabling moving object removal, and offers user-interactive functions different from conventional 3D modeling, such as allowing the user to select a target area on the region segmentation results and view the removal results. Compared to some denoising processes performed after modeling, removing objects before modeling can further prevent noise from being introduced into the 3D model. The removed images are compatible with various commercial modeling software and have market applicability.
[0103] The above paragraphs are described in several aspects. Clearly, the teachings herein can be implemented in a variety of ways. Any particular structure or function disclosed in the examples is merely representative. Those skilled in the art should note that any of the disclosed embodiments may be implemented individually or in combination with two or more embodiments.
[0104] While embodiments are described above, this disclosure is not limited thereto. This disclosure can be modified and adjusted by those skilled in the art without departing from the spirit and scope of this disclosure. The scope of protection of the present invention shall be subject to the claims. [Explanation of symbols]
[0105] 100 First position 102 Flight Path 104 Second position 106 Flight Paths 108 Third Position 110 Target bodies 112 Feature Point A 114 Feature Point B 116 Ray of light 120 Image 1 122 a1 point 124 b1 point 130 Image 2 132 a2 points 134 b2 points 140 Third image 142 a3 points 144 b3 points 200 positions Flight path 202 204 position 210 Point A 212 Ray of light 214 Image 1 216 point a 220 A' point 222 Ray of light 224 Image 2 226 point a' 230 A* point 300 Moving object removal method S302~S310 Step 400 Original image Train 402 404 Bridges 406 Ground 410 Second original image Train 412 414 Bridges 416 Ground 500 First feature point detection result 502, 504 Feature points 510 Second Feature Point Detection Results 512 Feature Points 600 First region partitioning result 602 Train area 604 Bridge area 606 Ground area 610 Second region partitioning result 612 Train area 614 Bridge area 616 Ground area 700 Original image (first version) Train 702 701, 703, 705, 706, 708, 711, 713, 715, 716, 718 Feature points 710 Second original image S802~S808 Step 900 Original image (first version) 902 Target area 904, 906 Non-target area 910 Second original image 912 Target area 914, 916 Non-target areas 1000, 1010 still images 1002, 1012 Unremarkable regions 1100 Computer System 1102 Processing Unit 1104 Storage device 1110 3D Modeling Program 1112 User Interface Module 1114 3D Modeling Module 1116 Mobile Object Removal Module 1200 Basic Operations of 3D Modeling Programs S1202~S1210 Step 1300 User Interface 1301 Image Import Section 1302 Mobile object removal unit 1303 Removal progress display section 1304 Removal result display area 1305 Modeling Department 1306 Target area selection unit 1307 Manual removal section 1308 Image display section 1400 User Interfaces 1401 Image Import Section 1402 Mobile object removal unit 1403 Removal progress display section 1404 Removal result display area 1405 Modeling Department 1407 Manual removal section 1408 Image display section
Claims
1. A computer program product used for 3D modeling, The aforementioned computer program product, after being loaded into a computer, provides a graphical user interface (GUI). The aforementioned user interface is An image import section for importing the original image sequence from a path specified by the user, A mobile object removal unit for the user to input a removal command, The aforementioned includes a modeling section for the user to input modeling commands, The computer program product, in response to the received removal command, causes the computer to execute a moving object removal method on the original image sequence so that it can acquire a still image sequence. The computer program product, in response to the received modeling command, causes the computer to select one of the original image sequence and the still image sequence based on the removal command, and to create a 3D model. Computer program products.
2. The graphical user interface further includes a removal progress display unit that shows the processing progress of the moving object removal method. The computer program product according to claim 1.
3. The graphical user interface further includes a target region selection unit that presents region division results so that the user can select a target region. The computer program product according to claim 1.
4. The computer program product according to claim 1, wherein the graphical user interface further includes an image display unit for presenting at least one of the source image sequence and the still image sequence.
5. The graphical user interface further includes a removal result display unit, If the area ratio of the moving region in the original image sequence exceeds the first specified threshold, and the number of original images determined not to contain the moving region in the original image sequence does not reach the second specified threshold, the removal result display unit presents an exception message. The removal result display unit displays a success message in response to the acquisition of the still image sequence. The computer program product according to claim 1.
6. After the exception message is presented, the graphical user interface further includes a manual removal section for the user to manually remove the moving object. The computer program product according to claim 5.
7. The aforementioned exception message is configured to guide the user to add more original images to the original image sequence. The computer program product according to claim 5.
Citation Information
Patent Citations
Road / feature measuring device, feature identifying device, road / feature measuring method, road / feature measuring program, measuring device, measuring method, measuring program, measured position data, measuring terminal, measuring server device, draw
WO2008099915A1