Computer program product for use in 3D modeling and mobile body removing method therefor
Patent Information
- Application Number
- JP2024140646
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-11-22
- Filing Date
- 2024-08-22
- Publication Date
- 2025-06-03
- Estimated Expiration
- 2044-08-22
Smart Images

Figure 2025084676000001_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to 3D modeling technology, and particularly to a computer program product used for 3D modeling and a method for removing moving objects therefrom.
Background Art
[0002] Unmanned aerial vehicles (UAVs) are capable of overcoming topographical constraints to perform tasks, and can assist professional personnel in performing tasks and effectively improve efficiency. Therefore, the entire market has been booming in recent years. It is known that it is rapidly developing in fields such as construction / engineering / mining, energy / transportation / public facilities, and precision agriculture. In particular, using UAVs to perform 3D modeling of terrain, landscapes, and buildings has become one of the mainstream applications in recent years.
[0003] Conventional 3D modeling of UAVs uses a UAV equipped with a high-resolution camera to continuously capture images of a stationary object or a scene to be measured at different positions and different angles, estimate the 3D coordinates of the UAV using the feature points on each image, and then inversely estimate the measured 3D coordinates to create a 3D model.
[0004] FIG. 1 is a schematic diagram of 3D modeling photography using a UAV. As shown in FIG. 1, the UAV flies from a first position 100 in the air along a flight path 102 to a second position 104, and then flies along a flight path 106 to a third position 108. Photography is performed on a target object 110 at these three positions, and a first image 120, a second image 130, and a third image 140 used for modeling are obtained respectively. Feature point A 112 and feature point B 114 on the target object 110 are imaged as point a 1 point 122 and point b 1 point 124 in the first image 120, point a 2 point 132 and point b 2 point 134 in the second image 130, point a 3 point 142 and point b 3They are respectively projected onto point 144. Next, the characteristic similarity of each image is calculated, and by back-projecting the rays of multiple viewing angles into the 3D space, the 3D coordinates of feature point A112 and feature point B114 can be estimated and represented in the form of a 3D point cloud.
[0005] The above calculation of 3D coordinates is based on the principle that the projections of corresponding points of images taken at different angles intersect at the same point within the 3D space. However, since the above-mentioned shooting process needs to last for a certain period of time (taking a bridge with a length of 30 meters as an example, about 10 minutes of shooting is required), if a moving object (such as a train or a high-speed railway) passes during the shooting, the ray projections of these moving objects in different frame images cannot intersect at the exact position, and noise will be included in the constructed 3D model.
[0006] Figure 2 shows the causes of noise in 3D modeling. As shown in Figure 2, the unmanned aerial vehicle flies from a certain position 200 in the air to another position 204 along the flight path 202, takes pictures of the same object at these two positions, and respectively obtains the first image 214 and the second image 224 for modeling. During this period, the object moves from point A 210 to point A' 220. Before the unmanned aerial vehicle moves, point A 210 is projected onto point a 216 on the first image 214 through ray 212. Before the unmanned aerial vehicle moves, point A' 220 is projected onto point a' 226 on the second image 224 through ray 222. Based on the above 3D modeling principle, since the image characteristics of point a 216 and point a' 226 are similar, when these are back-projected into the 3D space, they will intersect at point A* 230. However, since the position of point A* 230 is neither point A 210 nor point A' 220, noise occurs.
[0007] In reality, in the shooting by an unmanned aircraft, it is impossible to restrict the entry and exit of objects into the shooting scene. Therefore, in order to ensure the accuracy of the final 3D model, it is necessary to manually remove noise one by one after 3D modeling, but this removal process is cumbersome and time-consuming. Currently, there is a method of performing noise removal processing on a 3D model after modeling, but since this depends on the projection relationship obtained in the initial 3D modeling, in certain situations, such a method fails. For example, when the input image has a moving object with a special appearance (such as being too large in size), an initial modeling error occurs, and the subsequent noise removal processing for the model may become ineffective.
[0008] In the future, as the demand for various 3D modelings using unmanned aircraft is expected to increase, the design of a moving object removal solution that can be used for 3D modeling has become an increasingly important issue.
Summary of the Invention
Problems to be Solved by the Invention
[0009] An object of the present disclosure is to provide a computer program product used for 3D modeling and a method for removing a moving object thereof.
Means for Solving the Problems
[0010] The moving object removal method according to the present disclosure is a moving object removal method used for 3D modeling, and includes a step of detecting a plurality of feature points for each of the original images in the original image sequence by a feature point detection process, a step of dividing each of the original images into a plurality of regions by a region division process, a step of specifying a target region and one or more non-target regions from the plurality of regions in each of the original images, a step of determining, by a feature point matching process, whether each of the non-target regions in the original images of two consecutive frames of the original image sequence is a moving region based on the plurality of feature points in the original images of the two consecutive frames, and a step of obtaining a still image sequence by replacing the moving regions and the regions determined to be such in the non-target regions of each of the original images with regions without features.
[0011] In one embodiment of the moving object removal method according to the present disclosure, the step of determining, by a feature point matching process, whether each of the non-target regions in the original images of two consecutive frames of the original image sequence is a moving region based on the plurality of feature points in the original images of the two consecutive frames includes a step of obtaining a plurality of matching pairs regarding the feature points by a feature point matching process based on the plurality of feature points in the original images of the two consecutive frames, a step of estimating a transformation matrix based on the plurality of matching pairs within the target region in the original images of the two consecutive frames, a step of calculating an average projection error for each of the non-target regions in the original images of the two consecutive frames based on the plurality of matching pairs within the non-target region and the transformation matrix, and a step of determining whether each of the non-target regions in the original images of the two consecutive frames is a moving region by comparing the average projection error with a threshold value.
[0012] One embodiment of the moving object removal method according to the present disclosure further includes, before the step of estimating the transformation matrix, a step of examining the homography relationship of the target region in the original images of two consecutive frames based on the number of a plurality of matching pairs within the target region in the original images of two consecutive frames.
[0013] One embodiment of the moving object removal method according to the present disclosure includes, for each of the non-target regions of the original images of two consecutive frames, a step of calculating the average projection error based on a plurality of matching pairs and the transformation matrix within the non-target region. For each of the matching pairs within the non-target region, the step of projecting the first feature point in the matching pair to a projection point using the transformation matrix, calculating the Euclidean distance between the projection point and the second feature point in the matching pair, and taking the average of the Euclidean distances of the matching pairs as the average projection error.
[0014] One embodiment of the moving object removal method according to the present disclosure further includes, before the step of comparing the average projection error with a threshold value, a step of calculating the average projection error of the target region and the standard deviation corresponding to the target region based on the matching pairs and the transformation matrix of the target region in the original images of two consecutive frames, and a step of determining the threshold value based on the average projection error and the standard deviation of the target region in the original images of two consecutive frames.
[0015] One embodiment of the moving object removal method according to the present disclosure is that each of the feature points has a feature descriptor.
[0016] One embodiment of the moving object removal method according to the present disclosure is that the feature point matching procedure includes a step of comparing these feature descriptors within the original images of two consecutive frames to obtain a plurality of matching pairs regarding the feature points.
[0017] One embodiment of the moving object removal method according to the present disclosure is that the feature point detection process includes the step of using the Scale-Invariant Feature Transform (SIFT) algorithm.
[0018] One embodiment of the moving object removal method according to the present disclosure is that the region division process includes the step of using a Vision Transformer Adapter (ViT-Adapter).
[0019] One embodiment of the moving object removal method according to the present disclosure is that the feature point matching process includes the step of using a Brute-Force Matcher (BFMatcher).
[0020] One embodiment of the present disclosure further provides a computer program product for 3D modeling including a user interface module, a 3D modeling module, and a moving object removal module. The user interface module provides a user interface. The 3D modeling module creates a 3D model. The moving object removal module executes a moving object removal method. After being loaded into a computer, the computer program product can execute the following steps: obtaining an original image sequence; when receiving a moving object removal instruction from the user interface, calling the moving object removal module to execute the moving object removal method, driving the 3D modeling module, and creating a 3D model using a series of still images obtained by executing the moving object removal method; and when receiving a direct modeling instruction from the user interface, driving the 3D modeling module and creating a 3D model using the original image sequence.
[0021] One embodiment of the computer program product according to the present disclosure is that the user interface is a graphical user interface (GUI) that presents the region division result of the region division process and allows the user of the computer to select a target region on the region division result.
[0022] In one embodiment of the computer program product according to the present disclosure, the moving object removal module further detects the area ratio of the moving region in the original image sequence, examines the number determined not to include the moving region in the original image sequence, and in response to the area ratio of the moving region in the original image sequence exceeding a first specified threshold and the number determined not to include the moving region in the original image sequence not reaching a second specified threshold, the moving object removal module notifies the user interface module to present an exception message on the user interface.
[0023] Another embodiment of the present disclosure further provides a computer program product for 3D modeling. After being loaded by a computer, the computer program product can provide a graphical user interface (GUI). The graphical user interface includes an image import section, a moving object removal section, and a modeling section. The image import section is for the user to input a specified path to import the original image sequence. The moving object removal section is for the user to input a removal command. The modeling section is for the user to input a modeling command. In response to the received removal command, the computer program product causes the computer to execute a moving object removal method on the original image sequence and obtain a still image sequence. In response to the received modeling command, the computer program product causes the computer to select and use one of the original image sequence and the still image sequence based on the removal command to create a 3D model.
[0024] In one embodiment of the computer program product according to the present disclosure, the graphical user interface further includes a removal progress display section for presenting the processing progress of the moving object removal method.
[0025] One embodiment of a computer program product according to the present disclosure is that the graphical user interface further includes a target area selection unit that presents a region division result for the user to select a target area on the region division result.
[0026] One embodiment of a computer program product according to the present disclosure is that the graphical user interface further includes an image display unit for presenting at least one of the original image sequence and the still image sequence.
[0027] In one embodiment, the graphical user interface further includes a removal result display unit. In response to the area ratio of the moving area in the original image sequence exceeding a first specified threshold and the number of frames determined not to include the moving area in the original image sequence not reaching a second specified threshold, the removal result display unit presents an exception message. In response to the acquired still image sequence, the removal result display unit presents a success message.
[0028] In one embodiment, after the exception message is presented, the graphical user interface further includes a manual removal unit for the user to manually remove the moving object.
[0029] In one embodiment, the exception message is configured to guide the user to further add an original image to the original image sequence.
[0030] The solutions provided by various embodiments of the present disclosure for removing moving objects in 3D modeling have the recognition functions of space (objects) and time (movement) simultaneously, and only remove the moving objects in the image while retaining the information of static objects, ensuring the integrity and accuracy of the 3D model. In addition, the user interface of the present disclosure further provides the flexibility to enable or disable the removal of moving objects, and provides user-interactive functions different from conventional 3D modeling, such as allowing the user to select a target area on the region division result and display the removal result. Compared with the noise removal process performed after some modeling, removing objects before modeling can further prevent noise from being mixed into the 3D model. The removed image is compatible with various commercial modeling software and has applicability in the market.
[0031] The present disclosure can be better understood from the following description of exemplary embodiments and the accompanying drawings. Further, it should be understood that in the flowcharts of the present disclosure, the execution order of each block may be changed, and / or some blocks may be changed, deleted, or combined.
Brief Description of Drawings
[0032]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9A
Figure 9B
Figure 10
Figure 11
Figure 12
Figure 13
Figure 14
Mode for Carrying Out the Invention
[0033] The following description lists various embodiments of the present invention, but is not intended to limit the content of the present invention. The actual scope of the invention is defined by the scope of the patent application.
[0034] In each of the embodiments listed below, the same or similar elements or components are represented by the same reference numerals.
[0035] The ordinal terms such as "first", "second", "third", etc. used in the claims are used only for convenience of explanation and do not mean a priority relationship between each other.
[0036] The following description regarding embodiments of the apparatus or system is also applicable to embodiments of the method, and vice versa.
[0037] FIG. 3 is a flowchart showing a moving object removal method 300 for 3D modeling according to an embodiment of the present disclosure. As shown in FIG. 3, the method 300 can include steps S302 to S310.
[0038] In step S302, a plurality of feature points in each original image in the original image sequence are detected by a feature point detection process.
[0039] The original image sequence can be a series of images in which a moving object has not been removed, taken from different angles by a photographing device mounted on an unmanned aerial vehicle. The present disclosure does not limit the type of the unmanned aerial vehicle or the photographing device and the photographing scene.
[0040] The feature point can be understood as the most distinguishable position or region within an image, and is usually a point where characteristics such as brightness change, texture, edge, color, and shape in the image are prominent (for example, having local extrema). It should be understood that due to the change in the viewing angle (or the position of the unmanned aerial vehicle), the position distribution of the feature points in different original images within the original image sequence is also different regardless of whether the photographed scene is completely stationary. In one embodiment, each feature point has a feature descriptor, and the feature descriptor can be represented in the form of a feature vector that symbolizes the gradient, direction, and feature intensity around the feature point.
[0041] FIG. 4 is a schematic diagram showing original images of two frames taken by an unmanned aircraft flying from right to left according to an embodiment of the present disclosure, that is, a first original image 400 and a second original image 410. As shown in FIG. 4, the first original image 400 includes a train 402, a bridge 404, and the ground 406, and the second original image 410 includes a train 412, a bridge 414, and the ground 416. In this example, it is assumed that the trains 402 and 412 are the same moving train photographed at different time points, and the bridge and the ground are stationary. From the first original image 400 and the second original image 410, it can be observed that the appearance of the train 412 has changed significantly compared to the train 402.
[0042] FIG. 5 is a schematic diagram showing a first feature point detection result 500 and a second feature point detection result 510 obtained after passing through step S302 of the first original image 400 and the second original image 410 according to an embodiment of the present disclosure. As shown in FIG. 5, the upper left part of the first feature point detection result 500 has feature points (such as feature point 502) on the stationary bridge and feature points (such as feature point 504) located on the moving train. However, in the second feature point detection result 510, there are no feature points on the stationary bridge in the upper left part, and it has feature points on the moving train such as feature point 512. That is, the feature points on the stationary bridge are partially blocked by the moving train.
[0043] The feature point detection process in step S302 can be implemented using various feature detection algorithms or corner detectors, such as the speeded up robust features (SURF) algorithm, the Accelerated-KAZE algorithm, the Harris corner detector, Features from Accelarated Segment Test (FAST), Binary Robust Invariant Scalable Keypoints (BRISK), and so on. In one embodiment, the feature point detection process is implemented using the scale-invariant feature transform (SIFT) algorithm. This concept involves performing a convolution operation on the image using Gaussian filters of different scales, obtaining the Gaussian variance based on the continuous Gaussian blurred images, and thereby obtaining feature points that are invariant to scaling and rotation. Each feature point has a feature descriptor, which can be represented in the form of a feature vector, representing the gradient, direction, and feature intensity around the feature point in the form of a feature vector. In actual operation, approximately 50,000 SIFT feature points can be detected from one 3840x2160 image, and a single SIFT feature point contains 128-dimensional information representing the magnitude and direction of the gradient within the region around the feature point.
[0044] Referring back to FIG. 3, in step S304, each original image is divided into a plurality of regions by the region division process.
[0045] More specifically, in the region division process, an index value is assigned to each pixel in the original image, and pixels having the same index value form a single region. For example, pixels having the index value "1" constitute the first region, and pixels having the index value "2" constitute the second region, and so on. In other words, the pixels in the first region have the index value "1", and the pixels in the second region have the index value "2", and so on. Thereby, in subsequent steps, each region can be identified based on the index value and corresponding processing can be performed.
[0046] FIG. 6 is a schematic diagram showing first region division results 600 and second region division results 610 of a first original image 400 and a second original image 410 respectively obtained after passing through step S304 according to an embodiment of the present disclosure. As shown in FIG. 6, the first region division result 600 includes a train region 602, a bridge region 604, and a ground region 606, and the second region division result 610 includes a train region 612, a bridge region 614, and a ground region 616. It should be noted that the terms "train region", "bridge region", and "ground region" described above are named to make it easier for the reader to understand the correspondence between the region division result and the original image, but in reality, these regions are identified by index values. The region division process does not specify and identify to which types of objects such as trains, bridges, and the ground these regions belong.
[0047] The region division process in step S304 is implemented by algorithms used for region division, such as Simple Linear Iterative Clustering (SLIC), Graph-Based Segmentation (Felzenszwalb’s Graph-Based Segmentation), Quick Shift, Region Growing, etc. Alternatively, a region division process constructed using a convolutional neural network (CNN)-based machine learning model such as U-Net, SegNet, or a fully convolutional network (FCN) can also be used. The training process can include obtaining labeled data, selecting a loss function, and configuring an optimization algorithm, etc., and various conventional methods can be adopted, but the present disclosure is not limited thereto. Further, the region division model can be trained locally, or after being trained in advance on another computer device (e.g., a server), a trained region division model can be obtained via a network (e.g., downloaded from the cloud), a storage medium (e.g., an external hardware), or another communication interface (e.g., USB), and the present disclosure is not limited thereto. In one embodiment, the region division process is implemented using a Visual Transformer Adapter (ViT-Adapter). The Visual Transformer Adapter is a deep learning-based classifier, and its concept is to improve the performance and versatility of the model by adding a small neural network module called an "adapter" into different layers of a pre-trained visual transformer model.
[0048] Referring to FIG. 3 again, in step S306, from within these regions of each original image, a target region and one or more non-target regions (i.e., regions other than the target region) are determined.
[0049] In one embodiment, the target area can be determined by a user's selection from the area division results presented in a graphical user interface (GUI). In another embodiment, a machine learning model for identifying the target area, i.e., a target area identification model, can be pre-trained, and then, in step S306, the trained target area identification model can be used to identify the target area from within these areas of each original image. The target area identification model can be a convolutional neural network (CNN)-based model, and the training process can include obtaining labeled data, selecting a loss function, and configuring an optimization algorithm, etc., and various conventional methods can be used, but the present disclosure is not limited thereto. Further, the target area identification model can also be trained locally, or after being trained in another computer device (e.g., a server) first, the trained target area identification model can be obtained via a network (e.g., downloaded from the cloud), a storage medium (e.g., an external hard drive), or another communication interface (e.g., USB), and the present disclosure is not limited thereto.
[0050] In a single modeling task, the target object of the modeling is known in advance. For example, for the modeling of a bridge, the targets are the bridge body and the surrounding terrain. Taking FIG. 6 as an example, the bridge area 604 can be selected as the target area. Since the original images in the original image sequence have a high degree of overlap with each other (generally 50% or more), the index value of the original bridge area 604 is further extended to the bridge area 614 of the next original image based on the overlap ratio (e.g., only the areas with an overlap of 50% or more are considered), and the bridge area 614 is regarded as the target area of the second original image 410. In other words, by executing step S306 once, the remaining original images in the original image sequence can know their target areas by inference.
[0051] Referring to FIG. 3 again, in step S308, based on the feature points in the original images of two consecutive frames in the original image sequence, it is determined by a feature point matching process whether each non-target region of the original images of the two consecutive frames is a moving region (i.e., the region corresponding to the moving object).
[0052] In one embodiment, the feature point matching process includes comparing feature descriptors in two images and finding similar feature points in the two images. Two feature points with matching feature descriptors form a matching pair. As described above, the feature descriptor can be represented in the form of a feature vector that symbolizes the gradient, direction, and feature intensity around the feature point. Therefore, the feature point matching process can further include calculating the distance or similarity between two feature vectors, such as the Euclidean distance, Manhattan distance, cosine similarity, or magnitude used to represent other distances or similarities, but the present disclosure is not limited thereto.
[0053] FIG. 7 is a schematic diagram showing feature point matching pairs of a first original image 700 and a second original image 710 according to an embodiment of the present disclosure. The first original image 700 and the second original image 710 respectively correspond to the first feature point detection result 500 and the second feature point detection result 510 in FIG. 5, but not all the feature points depicted in FIG. 5 appear in FIG. 7. In the example of FIG. 7, only nine similar feature points found in the feature point matching process of step S308, that is, nine matching pairs are shown, including a first matching pair composed of a feature point 701 in the first original image 700 and a feature point 711 in the second original image 710, a second matching pair composed of a feature point 703 in the first original image 700 and a feature point 713 in the second original image 710, a third matching pair composed of a feature point 705 in the first original image 700 and a feature point 715 in the second original image 710, a fourth matching pair composed of a feature point 706 in the first original image 700 and a feature point 716 in the second original image 710, a fifth matching pair composed of a feature point 708 in the first original image 700 and a feature point 718 in the second original image 710, and four other unlabeled matching pairs. In these matching pairs, the first matching pair, the second matching pair, and the third matching pair are located on the moving train 702, and the fourth matching pair and the fifth matching pair are located on the stationary bridge. Therefore, there is a greater displacement between the feature point 701 and the feature point 711, between the feature point 703 and the feature point 713, and between the feature point 705 and the feature point 715 than between the feature point 706 and the feature point 716 or between the feature point 708 and the feature point 718. Accordingly, it can be determined that the area corresponding to the train 702 is a moving area.
[0054] The feature point matching process in step S308 can be implemented using nearest neighbor matching, Random Sample Consensus (RANSAC), the Kanade-Lucas-Tomasi feature tracker, or other similar algorithms. In one embodiment, the feature point matching process is implemented using a brute force matcher (BFMatcher) that searches for an optimal match by calculating the distance or similarity between two sets of feature descriptors (or feature vectors).
[0055] FIG. 8 is a flowchart showing more detailed steps of the moving area determination in step S308 according to an embodiment of the present disclosure. As shown in FIG. 8, step S308 can further include steps S802 to S808.
[0056] In step S802, based on the feature points in two consecutive frames of the original image, a plurality of matching pairs of these feature points are obtained by a feature point matching process. Taking FIG. 7 as an example, after step S802, nine matching pairs of the first original image 700 and the second original image 710 can be obtained.
[0057] In step S804, a transformation matrix is estimated based on the matching pairs within the target area of two consecutive frames of the original image.
[0058] The transformation matrix can represent the transformation relationship between the feature points of each matching pair and the positions of other pixels between the two original images. Therefore, step S804 can be understood as obtaining the transformation relationship of the pixel positions within the target area of two consecutive frames of the original image. In subsequent steps, the moving area can be found by examining whether the transformation relationship of each non-target area is significantly different from that of the target area.
[0059] In step S806, for each non-target region of the original images of two consecutive frames, an average projection error is calculated based on the matching pairs and transformation matrix in the non-target region.
[0060] The average projection error can represent the dissimilarity between the pixel position transformation relationship of the non-target region and the pixel position transformation relationship of the target region. In other words, the larger the average projection error, the greater the difference between the pixel position transformation relationship representing the non-target region and the pixel position transformation relationship of the target region.
[0061] In step S808, for each non-target region in the original images of two consecutive frames, the average projection error is compared with a threshold to determine whether the non-target region is a moving region.
[0062] Specifically, when the average projection error exceeds a threshold, it means that the difference between the pixel position transformation relationship of the non-target area and the pixel position transformation relationship of the target area is quite large, so the non-target area is determined to be a moving area. The threshold can be a predefined numerical value, or a variable determined by a specific calculation.
[0063] 9A and 9B show a transformation matrix H 3×3 9A and 9B, a first original image 900 includes an area 902, an area 904, and an area 906, which correspond to an area 912, a second area 914, and a third area 916, respectively, of a second original image 910. The areas 902 (and 912), 904 (and 914), and 906 (and 916) have corresponding index values 0, 1, and 2, respectively, indicated by k. In this example, the areas 902 and 912 with index value k=0 are set as target areas, and the other areas are set as non-target areas. In addition, M kLet it represent the number of matching pairs in the region where the index value is k, and the i-th feature point in the region where the index value is k in the first original image 900 and the second original image 910 is
Number
Number
Number
Number
[0064] The transformation matrix H 3x3 is estimated to conceptually find one optimal transformation matrix H 3x3 and make the following <Equation 1> satisfied as much as possible for each feature point (i.e., i = 1~M k ) within the target region.
Number
[0065] Next, in step S806, for each feature point within each non-target region (i.e., the region where the index value k≠0) on the first original image 900
Number
Number
[0066] As shown in FIG. 9B, the feature points in the non-target region 904 of the first original image 900 are projected onto the respective ones by the transformation matrix H
Number
[0067] Subsequently, the Euclidean distance between the projection points in the second original image 910
Number
Number
[0068] Taking FIG. 9B as an example, the average projection error value E of the non-target regions 904 and 914 with index value k = 1 1 is calculated as follows:
Number
[0069] In one embodiment, the estimation of the transformation matrix H 3×3 is the least squares method, Least Absolute Deviations (LAD), least squares support vector machine (LS-SVM), It can include using algorithms related to polynomial fitting or any other function fitting, but the present disclosure is not limited thereto.
[0070] In one embodiment, before step S808, based on the matching pairs and transformation matrices in the target regions of the original images of two consecutive frames, the average projection error and the corresponding standard deviation of the target region, that is, the average and standard deviation of the projection errors of each feature point within the target region can be calculated. More specifically, the standard deviation σ is calculated as in the following <Equation Four>:
Equation
[0071] However, based on the calculated average projection error and standard deviation, a threshold value used for comparison with the average projection error in step S808 can be determined. In a preferred embodiment, the threshold value is set as the sum of the average projection error E 0 and twice the standard deviation. For example, if the average projection error E 0 of the target region is 75 and the standard deviation is 50.4, the threshold value is 75 + 2×50.4 = 175.8. Taking FIGS. 9A and 9B as examples, the average projection error E 0 of the target regions 902 and 912 and the corresponding standard deviation are calculated, and accordingly, the corresponding threshold value T p can be determined. When the average projection error value E 1 of the non-target regions 904 and 914 is greater than T p , the non-target regions 904 and 914 are determined to be moving regions.
[0072] In one embodiment, before step S804, based on the number of matching pairs within the target region of the original images of two consecutive frames, the homography relationship of the target region in the original images of the two consecutive frames can be examined. If it is determined that the target region has a homography relationship in the original images of the two consecutive frames, the conditions for estimating the transformation matrix in step S804 are satisfied. Otherwise, steps S804 to S808 are skipped. That is, it is considered that no moving region is detected from the original images of the two consecutive frames.
[0073] The aforementioned homography relationship refers to a reversible transformation from the real projective plane to the projective plane. When photographing a stationary object at a long distance, since the original images of two consecutive frames can be regarded as being on substantially the same plane, theoretically, the stationary region (represented by the target region) of the original image of the first frame must be transformed into the corresponding region of the original image of the second frame based on the projective relationship.
[0074] Referring to FIG. 3 again, in step S310, a series of still images are obtained by replacing the regions determined to be moving regions within the non-target regions of each original image with regions without features.
[0075] The above-mentioned featureless region means that the pixels within this region do not have feature values. Generally, a region of a pure color (such as white) can be used as a region without features.
[0076] FIG. 10 shows examples of still images 1000 and 1010 of two frames obtained after step S310 according to an embodiment of the present disclosure. As shown in FIG. 10, in still images 1000 and 1010, the regions presented in the outline of the train (for example, the head of the train and the passenger cars) are replaced with featureless regions 1002 and 1012 of pure white.
[0077] FIG. 11 is a system block diagram showing a computer system 1100 for 3D modeling according to an embodiment of the present disclosure. As shown in FIG. 11, the computer system 1100 can include a processing device 1102 and a storage device 1104.
[0078] The computer system 1100 can be a personal computer (e.g., a desktop computer or a laptop computer) that executes an operating system (e.g., Windows, Mac OS, Linux, UNIX, etc.), or a server computer, or a mobile device such as a tablet computer or a smartphone, and the present disclosure is not limited thereto.
[0079] The processing device 1102 can include any one or more general-purpose processors or dedicated processors for executing instructions, and combinations thereof. In various embodiments of the present disclosure, the processing device 1102 is installed to execute the aforementioned moving object removal method such as method 300. In one embodiment, although not shown in FIG. 11, the processing device 1102 can further include a central processing unit (CPU) and a graphics processing unit (GPU). Since the GPU is an electronic circuit specifically designed to perform computer graphics operations and image processing, it is more efficient than a general-purpose CPU in computer graphics operations and image processing. Therefore, in various embodiments of the present disclosure, appropriate tasks can be assigned according to the characteristics of the CPU and the GPU. For example, tasks such as acquiring data or communicating with other devices can be assigned to the CPU, and tasks related to computer graphics and image processing can be assigned to the GPU. In a further embodiment, although not shown in FIG. 11, the processing device 1102 can further include a neural network processing unit (NPU) specifically optimized for deep learning. Compared with the GPU, the NPU may have more advantages in computing performance in the operation of the aforementioned region division model and / or target region identification model. Therefore, in this embodiment, tasks including these machine learning models can be assigned to the NPU.
[0080] The memory device 1104 can include a volatile memory (such as a random access memory (RAM), a dynamic random access memory (DRAM), a static random access memory (SRAM), etc.) and / or any one or more types of non-volatile memory (read-only memory, electrically erasable programmable read-only memory (EEPROM), flash memory, non-volatile random access memory (NVRAM)), devices (such as a hard disk drive (HDD), a solid state drive (SSD), or an optical disk, etc.), and combinations thereof, but the present disclosure is not limited thereto. In various embodiments of the present disclosure, the memory device 1104 is used to store a program corresponding to the aforementioned moving object removal method. When the processing device 1102 loads this program from the memory device 1104, the aforementioned moving object removal method can be executed.
[0081] In the embodiment shown in FIG. 11, the memory device 1104 stores a computer program product for 3D modeling, herein referred to as the "3D modeling program" 1110. As shown in FIG. 11, the 3D modeling program 1110 can include a user interface module 1112, a 3D modeling module 1114, and a moving object removal module 1116. The user interface module 1112 is used to provide a user interface, the 3D modeling module 1114 is used to create a 3D model, and the moving object removal module is used to execute the aforementioned moving object removal method, such as method 300. The processing device 1102 loads the 3D modeling program 1110 from the memory device 1104 and executes it to execute the user interface module 1112, the 3D modeling module 1114, and the moving object removal module 1116, as well as other basic operations of the 3D modeling program 1110.
[0082] In one embodiment, the processing device 1102 is coupled to a display device and can display a user interface provided by the user interface module 1112 thereon. The display device can be any device used to display visual information such as an LCD display, an LED display, an OLED display, or a plasma display, but the present invention is not limited thereto.
[0083] The user interface provided by the user interface module 1112 can be a graphical user interface (GUI), a command line interface (CLI), a touch interface, or a voice interface, but the present disclosure is not limited thereto.
[0084] The 3D modeling module 1114 can include any known 3D modeling techniques such as stereo correspondence, point cloud reconstruction, point cloud processing, 3D rendering, etc., but the present disclosure is not limited thereto.
[0085] FIG. 12 is a flowchart of the basic operation 1200 of the 3D modeling program 1110 according to an embodiment of the present disclosure. As shown in FIG. 12, the basic operation 1200 of the 3D modeling program 1110 can include steps S1202 to S1210.
[0086] In step S1202, an original image sequence is obtained.
[0087] As described above, the original image sequence is a series of images taken by a photographing device mounted on an unmanned aerial vehicle from different viewpoints, and the moving object has not been removed yet. The original image sequence can be obtained via a network, a storage medium, or various other communication interfaces, but the present disclosure is not limited thereto.
[0088] In step S1204, the user interface module 1112 receives instructions from the user of the computer system 1100 via the user interface. In response to receiving the moving object removal instruction, the process proceeds to step S1206. In response to receiving the direct modeling instruction, the process proceeds to step S1210.
[0089] In step S1206, the moving object removal module 1116 is called to execute the aforementioned moving object removal method (Method 300).
[0090] In step S1208, the 3D modeling module 1114 is driven using the still image sequence obtained by executing the moving object removal method to create a 3D model.
[0091] In step S1210, the 3D modeling module 1114 is driven to create a 3D model using the original images of this sequence.
[0092] In one embodiment, the user interface provided by the user interface module 1112 is a graphical user interface (GUI), which presents the region division results of the region division process (for example, the first region division result 600 and the second region division result 610 shown in FIG. 6), enabling the user of the computer system 1100 to select a target region from the region division results. Thereby, the moving object removal module 1116 can know the target region in step S306 and determine other regions outside the target region as non-target regions. Also, on this graphical user interface, checkboxes, toggle switches, dropdown lists, or other similar GUI elements, or widgets are provided, enabling the user of the computer system 1100 to select a moving object removal instruction or a direct modeling instruction. However, the present disclosure is not limited to the specific graphic design layout of the graphical user interface.
[0093] If the moving speed of the object is too slow, the target area will be frequently covered in the original image sequence, which may affect the accuracy of 3D modeling. In view of this, in one embodiment, the moving object removal module 1116 further detects the area ratio of the moving area in the original image sequence and examines the number determined not to include the moving area in the original image sequence. In response to the area ratio of the moving area in the original image sequence exceeding the first specified threshold and the number determined not to include the moving area in the original image sequence not reaching the second specified threshold, the moving object removal module 1116 notifies the user interface module 1112 to present an exception message to the user interface. The content of the exception message can be guided and set so that the user can manually remove the moving object or add more additional original images (for example, images of non-existent and slowly moving moving objects taken at other times) to facilitate 3D modeling.
[0094] FIG. 13 is an interface block diagram showing a user interface 1300 that can be provided by a 3D modeling program after being loaded by a computer according to an embodiment of the present disclosure. Similarly, FIG. 14 is a schematic example of a user interface 1400 provided by the 3D modeling program when the above-described exception event is triggered. The user interface 1300 includes at least the image import section 1301, the moving object removal section 1302, and the modeling section 1305 shown in FIG. 13. Similarly, the user interface 1400 includes at least the image import section 1401, the moving object removal section 1402, and the modeling section 1405 shown in FIG. 14.
[0095] The image import units 1301 and 1401 are for the user to input a specified path, such as "C: / Users / 3DModeling / RawData", or a path in other similar forms, and import the original image sequence from this path. The image import units 1301 and 1401 can be implemented, for example, with a file selection dialog, a text box, file drag and drop, a tree view, or other similar GUI elements or widgets, but the present disclosure is not limited thereto.
[0096] The moving object removal units 1302 and 1402 are for the user to input a removal command. The removal command indicates whether the moving object removal method (for example, the moving object removal method 300 in FIG. 3) is effective and is used to trigger the execution of the moving object removal method. In response to receiving the removal command input by the user via the moving object removal units 1302 and 1402, the 3D modeling program causes the computer to execute the moving object removal method on the original image sequence imported from the specified path (input by the user via the image import units 1302 and 1402). The moving object removal units 1302 and 1402 can be implemented, for example, with a checkbox, a toggle switch, a drop-down menu, or other similar GUI elements or widgets combined with a confirmation button or a start button, but the present disclosure is not limited to these.
[0097] The modeling units 1305 and 1405 are for the user to input modeling commands, and the modeling commands are used to trigger the process of 3D modeling. More specifically, when the removal command indicates that the moving object removal method is effective, a 3D model is created using a still image sequence. Conversely, when the removal command indicates that the moving object removal method is not effective, a 3D model is created using the original image sequence. The modeling units 1305 and 1405 can be implemented, for example, with a confirmation button, a start button, toolbar buttons, menu options, or other similar GUI elements or widgets, but the present disclosure is not limited thereto.
[0098] In one embodiment, the user interfaces 1300 and 1400 can further include removal progress display units 1303 and 1403. The removal progress display units 1303 and 1403 are for presenting the processing progress of the moving object removal method, and the processing progress is, for example, a completion rate such as 20%, 50%, 90%, or a text representation of a state such as "processing" or "completed". The removal progress display units 1303 and 1403 can be implemented, for example, with a progress bar, a circular progress bar, a digital display, a status label, a timeline, or other similar GUI elements or widgets, but the present disclosure is not limited thereto.
[0099] In one embodiment, the user interface 1300 can further include a target area selection unit 1306. The target area selection unit 1306 is used to present area division results (such as the first area division result 600 and the second area division result 610 shown in FIG. 6), enabling the user to select a target area on this area division result. As an exemplary example, the selection and deletion unit 1407 of the user interface 1400 is used to present the above-mentioned area division results, enabling the user to select a target area on this area division result.
[0100] In one embodiment, user interfaces 1300 and 1400 can further include image display units 1308 and 1408. Image display units 1308 and 1408 are used to present the original image sequence and / or the still image sequence.
[0101] In one embodiment, user interfaces 1300 and 1400 can further include removal result display units 1304 and 1404. In response to the successful execution of the moving object removal method and the acquired still image sequence, removal result display units 1304 and 1404 present a success message. The success message is used to inform the user that the moving object (or moving area) in the original image sequence has been removed, and its content can be, for example, "The removal of 13 moving areas out of 105 images was successful." In response to an exception event, that is, the area ratio of the moving area in the original image sequence exceeds the first specified threshold and the number of images determined not to contain the moving area in the original image does not reach the second specified threshold, removal result display units 1304 and 1404 present an exception message. The exception message is used to inform the user that the moving object removal method has encountered an exception event (e.g., the moving object covers the target area) and has not been successful. In a further embodiment, the content of the exception message can be set to guide the user to add more additional original images (e.g., images taken at other times without slowly moving objects) to the original image sequence. In another embodiment, the content of the exception message can be set to guide and enable the user to manually remove the moving object. After this exception message is presented, user interface 1300 can further include a manual removal unit 1307 for the user to manually remove the moving object. As an illustrative example, the selective removal unit 1407 of user interface 1400 can be used to enable the user to manually remove the moving object after this exception message is presented.
[0102] The solutions provided by various embodiments of the present disclosure for removing moving objects in 3D modeling have a recognition function for both space (objects) and time (movement), and only remove the moving objects in the image while retaining the information of static objects, ensuring the integrity and accuracy of the 3D model. In addition, the user interface of the present disclosure further provides flexibility to enable or disable the removal of moving objects, and provides user-interactive functions different from conventional 3D modeling, such as allowing the user to select a target area on the region division result and display the removal result. Compared with the noise removal process performed after some modeling, removing objects before modeling can further prevent noise from being mixed into the 3D model. The removed image is compatible with various commercial modeling software and has applicability in the market.
[0103] The above paragraph is described in multiple aspects. Obviously, the teachings of this specification can be implemented in various ways. Any specific structure or function disclosed in the embodiments is merely a representative example. According to the teachings of this specification, those skilled in the art should note that any disclosed aspect can be implemented individually or two or more aspects can be combined and implemented.
[0104] Although the present disclosure is described above as an embodiment, the present disclosure is not limited thereto. The present disclosure can be modified and adjusted by those skilled in the art without departing from the spirit and scope of the present disclosure. The protection scope of the present invention shall be in accordance with the scope of the claims.
Description of Reference Numerals
[0105] 100 First position 102 Flight path 104 Second position 106 Flight path 108 Third position 110 Target body 112 Feature point A 114 Feature point B 116 Light ray 120 First image 122 a 1 Point 124 b 1 Point 130 Second image 132 a 2 Point 134 b 2 Point 140 Third image 142 a 3 Point 144 b 3 Point 200 Position 202 Flight path 204 Position 210 Point A 212 Light ray 214 First image 216 Point a 220 Point A’ 222 Light ray 224 Second image 226 Point a’ 230 Point A* 300 Moving object removal method Steps S302~S310 400 First original image 402 Train 404 Bridge 406 Ground 410 Second original image 412 Train 414 Bridge 416 Ground 500 First feature point detection result Feature points 502, 504 510 Second feature point detection result Feature point 512 600 First region segmentation result Train region 602 Bridge region 604 Ground region 606 610 Second region segmentation result Train region 612 Bridge region 614 Ground region 616 700 First original image Train 702 Feature points of 701, 703, 705, 706, 708, 711, 713, 715, 716, 718 Second original image 710 Steps S802~S808 First original image 900 Target area 902 Non-target areas 904, 906 Second original image 910 Target area 912 Non-target areas 914, 916 Still images 1000, 1010 Regions without features 1002, 1012 Computer system 1100 Processing device 1102 Storage device 1104 3D modeling program 1110 User interface module 1112 3D modeling module 1114 Moving object removal module 1116 Basic operations of 3D modeling program 1200 Steps S1202~S1210 User interface 1300 Image import section 1301 Moving object removal section 1302 Removal progress display section 1303 Removal result display section 1304 Modeling section 1305 Target area selection section 1306 Manual removal section 1307 Image display section 1308 User interface 1400 Image import section 1401 Moving object removal section 1402 Removal progress display section 1403 Removal result display section 1404 Modeling section 1405 Manual removal section 1407 1408 Image display unit
Claims
1. A moving object removal method for use in 3D modeling, comprising the steps of: detecting a plurality of feature points for each of the source images in the sequence of source images through a feature point detection process; Segmenting each of the original images into a plurality of regions through a region segmentation process; identifying a target region and one or more non-target regions from the plurality of regions in each of the source images; Based on the plurality of feature points in the original images of the two consecutive frames of the original image sequence, determining whether each of the non-target regions of the original images of the two consecutive frames is a moving region through a feature point matching process; and replacing the determined moving regions in the non-target regions of each of the source images with featureless regions to obtain a still image sequence. How to remove moving objects.
2. The step of determining whether each of the non-target regions of the original images of the two consecutive frames of the original image sequence is the moving region through the feature point matching process based on the plurality of feature points in the original images of the two consecutive frames of the original image sequence, includes: obtaining a plurality of match pairs of the feature points through the feature point matching process based on the plurality of feature points in the original images of the two consecutive frames; estimating a transformation matrix based on the plurality of matching pairs within the target region in the original images of the two consecutive frames; calculating an average projection error for each of the non-target regions of the original images of the two consecutive frames based on the plurality of matching pairs in the non-target region and the transformation matrix; determining whether each of the non-target regions of the original images of the two consecutive frames is a moving region by comparing the average projection error with a threshold; Including, The moving body removal method according to claim 1 .
3. Before estimating the transformation matrix, a homography relationship of the target area in the original images of the two consecutive frames is checked based on the number of the plurality of matching pairs in the target area in the original images of the two consecutive frames; Further comprising: The moving body removal method according to claim 2 .
4. The step of calculating the average projection error for each of the non-target regions of the original images of the two consecutive frames based on the plurality of matching pairs in the non-target region and the transformation matrix includes: For each matching pair in the non-target region, projecting a first feature point in the matching pair to a projection point using the transformation matrix, and calculating a Euclidean distance between the projection point and a second feature point in the matching pair; the average of the Euclidean distances of the matching pairs being the average projection error; Including, The moving body removal method according to claim 2 .
5. Prior to the step of comparing the average projection error to the threshold, Calculating the average projection error of the target area and a standard deviation corresponding to the target area based on the matching pairs of the target areas of the original images of the two consecutive frames and the transformation matrix; determining the threshold value based on the average projection error and the standard deviation of the target area of the original images of the two consecutive frames; Further comprising: The moving body removal method according to claim 2 .
6. Each of the feature points has a feature descriptor. The moving body removal method according to claim 1 .
7. The feature point matching process includes a step of comparing the feature descriptors in the original images of the two consecutive frames to obtain a plurality of matching pairs for the feature points. The moving body removal method according to claim 6.
8. The feature point detection process includes using a scale invariant feature transform (SIFT) algorithm; The moving body removal method according to claim 6.
9. The segmentation process includes using a Vision Transformer Adapter (ViT-Adapter), The moving body removal method according to claim 1 .
10. The feature point matching process includes using a brute force matcher (BF Matcher); The moving body removal method according to claim 1 .
11. a user interface module for providing a user interface; a 3D modeling module for creating a 3D model; A computer program product for use in 3D modeling, comprising: a moving object removal module for executing the moving object removal method according to any one of claims 1 to 10, The computer program product, after being loaded into a computer, obtaining an original image sequence; receiving a moving object removal command from the user interface, calling the moving object removal module to execute the moving object removal method and driving the 3D modeling module to create the 3D model using the still image sequence obtained by executing the moving object removal method; upon receiving a direct modeling command from the user interface, driving the 3D modeling module to create the 3D model using the source image sequence; causing the computer to execute Computer program products.
12. the user interface being a graphical user interface (GUI) that presents the segmentation results of the segmentation process so that a user of the computer can select the target region for the segmentation results.
12. A computer program product as claimed in claim 11.
13. The moving object removal module further comprises: Detecting an area ratio of the moving region in the original image sequence, and checking the number of original images in the original image sequence that are determined not to include the moving region; When the area percentage of the moving region in the source image sequence exceeds a first specified threshold and the number of source images in the source image sequence that are determined not to include the moving region does not reach a second specified threshold, the moving object removal module notifies the user interface module to present an exception message in the user interface.
12. A computer program product as claimed in claim 11.
14. 1. A computer program product for use in 3D modeling, comprising: the computer program product provides a graphical user interface (GUI) after being loaded into a computer; The user interface includes: an image import unit for importing an original image sequence from a path input by a user; a moving object removal unit for the user to input a removal command; a modeling section for the user to input modeling commands; the computer program product causes the computer to perform a moving object removal method on the source image sequence to obtain a still image sequence in response to receiving the removal command; the computer program product causes the computer to, in response to receiving the modeling command, select one of the source image sequence and the still image sequence based on the removal command, and create a 3D model. Computer program products.
15. The graphical user interface further includes a removal progress display unit that displays a processing progress status of the moving object removal method.
15. A computer program product according to claim 14.
16. the graphical user interface further includes a target region selection unit that presents the region segmentation result so that the user can select a target region.
15. A computer program product according to claim 14.
17. the graphical user interface further comprising an image display for presenting at least one of the original image sequence and the still image sequence.
15. A computer program product according to claim 14.
18. The graphical user interface further includes a removal result display section; When an area ratio of the moving region in the original image sequence exceeds a first specified threshold, and a number of original images determined not to include the moving region in the original image sequence does not reach a second specified threshold, the removal result display unit presents an exception message; the removal result display unit presents a success message in response to the still image sequence being acquired.
15. A computer program product according to claim 14.
19. and after the exception message is presented, the graphical user interface further includes a manual removal unit for the user to manually remove the moving object.
20. The computer program product of claim 18.
20. the exception message is configured to guide the user to add further source images to the source image sequence; 20. The computer program product of claim 18.
Citation Information
Patent Citations
Generation device, generation method and program of three-dimensional model
JP2019106145A
Image processing apparatus, image processing system, image processing method, and program
JP2022096217A
Information processing device, information processing method, and program
WO2021210492A1