Terminal device, information processing system, information processing method and program
The terminal device and method align 3D models with real structures by generating and matching 3D point cloud data using feature points and relative vector information, addressing inaccuracies caused by deformations and measurement errors.
Patent Information
- Application Number
- JP2025076148
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2025-05-01
- Publication Date
- 2025-10-30
- Estimated Expiration
- 2045-05-01
AI Technical Summary
Existing 3D models of structural information often mismatch with actual structures due to deformations and measurement errors, leading to inaccurate superimposition of information on real structures.
A terminal device and method that generates and matches 3D point cloud data from both the structure and its 3D model using feature points, relative vector information, and incorporates techniques like SfM and SLAM to align camera positions and angles accurately.
Enables precise superimposition of structural information on actual structures by reducing processing load and ensuring accurate alignment of 3D models with real-world structures.
Smart Images

Figure 0007762394000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to a terminal device, an information processing system, an information processing method, and a program. [Background technology]
[0002] BIM / CIM (Building Information Modeling / Construction Information Modeling / Management) models are known as information models (hereafter referred to as structural information) that link various information required for a series of processes such as planning, design, construction, and maintenance to a 3D model of a structure. In recent years, several technologies have been proposed to link real structures with structural information.
[0003] Patent Document 1 proposes recording the location of damage discovered during inspection of a structure in a 3D model of the structural information, and superimposing information associated with the 3D model on the actual structure.
[0004] Patent Document 2 proposes superimposing a three-dimensional model of structural information on an actual structure, and then utilizing the three-dimensional model as a map for controlling moving objects inside the structure. [Prior art documents] [Patent documents]
[0005] [Patent Document 1] Patent application 2024-111907 [Patent Document 2] International Publication No. 2024 / 166659 Summary of the Invention [Problem to be solved by the invention]
[0006] 3D models of structural information are digital twins of actual structures, and ideally, the two should overlap with each other without any dimensional difference. However, actual structures can be deformed by factors such as temperature, and position information measured in the real world can contain measurement errors. As a result, discrepancies can occur between the position and shape of a structure in the real world and the position and shape of the 3D model of the structural information.
[0007] For example, if information associated with a specific component of a 3D model of structural information is to be superimposed on the actual structure using the component's position information (expressed in the coordinate system of the 3D model) as is, the information may be displayed at a position that is different from the correct position on the actual structure.Similarly, if the position of damage discovered in an actual structure is measured using GNSS (Global Navigation Satellite System) or SLAM (Simultaneous Localization and Mapping), and the position information (expressed in the coordinate system of the real world) is used as is to record the damage on a 3D model, the damage may be recorded at a position that is different from the correct position on the 3D model.
[0008] On the other hand, technologies such as SLAM are capable of estimating self-location by dynamically generating maps, while technologies such as VPS (Visual Positioning System) can be distinguished as technologies that estimate self-location based on visual information from cameras, etc. The challenge with technologies such as SLAM is how to estimate self-location by creating maps beforehand or afterward, while the challenge with technologies such as VPS is how to estimate self-location based only on visual information from cameras, etc. Both technologies have their own challenges.
[0009] To avoid such problems and manage the coordinate system of the 3D model of the structural information as a reference, and to use it as an alternative to maps in technologies such as SLAM, it is necessary to use feature points set on the corners, ridges, faces, etc. of the structure as clues, as shown in Figure 11. Although the structure as a whole undergoes elastic deformation, it undergoes rigid deformation locally, so an operation (matching) is required to rigidly move and rotate the shape of the structure measured in the real world and match it exactly to the shape of the 3D model of the structural information. Note that a similar operation can also be achieved by matching the position and direction of the camera in the real world with the position and direction of the camera in the 3D model.
[0010] The present invention relates to such a technique and aims to provide a terminal device, an information processing system, an information processing method, and a program that are capable of accurately superimposing a structure on a three-dimensional model of structural information. [Means for solving the problem]
[0011] According to one embodiment, the terminal device includes an image acquisition unit that controls a camera to acquire an image of a structure, an image processing unit that generates 3D point cloud data including feature points of the structure based on the image, a structural information acquisition unit that acquires structural information including a 3D model of the structure, a map generation unit that generates 3D point cloud data including feature points of the 3D model based on the structural information, and a matching unit that matches the 3D point cloud data indicating the shape of the structure with the 3D point cloud data indicating the shape of the 3D model, and the matching unit generates relative vector information for all the feature points including information on the distance or vector from the feature point to other feature points in the same point cloud data, and performs the matching based on the relative vector information. According to one embodiment, the terminal device further includes a display unit that uses three-dimensional point cloud data indicating the shape of the structure as a map and displays predetermined information superimposed on the structure. According to one embodiment, the matching unit of the terminal device histograms the relative vector information and performs the matching based on the histogrammed relative vector information. According to one embodiment, the image processing unit of the terminal device removes edges indicating parallel or vertical straight lines that constitute the structure from the image, and extracts the feature points from the remaining edges. According to one embodiment, the terminal device has a matching unit that identifies the position and angle of a camera in the three-dimensional model using line information and surface information of the image and the three-dimensional model, and performs the matching. According to one embodiment, a terminal device includes: a SLAM processing unit that generates first three-dimensional point cloud data including feature points of a structure; an image acquisition unit that controls a camera to acquire an image of the structure; an image processing unit that generates second three-dimensional point cloud data including feature points of the structure based on the image; a correction unit that outputs the first three-dimensional point cloud data if the accuracy of the SLAM processing unit is within an acceptable range, and outputs the second three-dimensional point cloud data if the accuracy is outside the acceptable range; a structural information acquisition unit that acquires structural information including a three-dimensional model of the structure; a map generation unit that generates three-dimensional point cloud data including feature points of the three-dimensional model based on the structural information; and a matching unit that matches the three-dimensional point cloud data output by the correction unit with three-dimensional point cloud data that indicates the shape of the three-dimensional model, and the matching unit generates relative vector information for all of the feature points, including information on the distance or vector from the feature point to other feature points in the same point cloud data, and performs the matching based on the relative vector information. According to one embodiment, the information processing system further includes an image processing server, wherein the image acquisition unit acquires successive still images of the structure and transmits them to the image processing server, the image processing unit receives the second three-dimensional point cloud data from the image processing server, and the image processing server generates the second three-dimensional point cloud data including feature points of the structure based on the successive still images and transmits it to the image processing unit. According to one embodiment, a terminal device places a three-dimensional virtual grid in a three-dimensional model space of a structure, performs 360-degree spherical rendering or rendering in all directions using virtual cameras placed on each grid, extracts feature points from the rendered image, and performs machine learning using the feature points extracted from the rendered image as learning data and the coordinates of the virtual camera as training data to generate a trained model. According to one embodiment, the terminal device is an information processing device that performs estimation processing using the trained model, extracts feature points from a photographed image of the structure, inputs the feature points extracted from the photographed image of the structure into the trained model, and estimates the camera position of the photographed image of the structure. According to one embodiment, the terminal device includes an image acquisition unit that controls a camera to acquire an image of a structure, an image processing unit that generates 3D point cloud data including feature points of the structure based on the image, a structural information acquisition unit that acquires structural information including a 3D model of the structure, a map generation unit that generates 3D point cloud data including feature points of the 3D model based on the structural information, and a matching unit that matches the 3D point cloud data indicating the shape of the structure with the 3D point cloud data indicating the shape of the 3D model, and the matching unit generates relative vector information for all the feature points including information on the distance or vector from the feature point to other feature points in the same point cloud data, and performs the matching based on the relative vector information. According to one embodiment, the information processing method is a method executed by an information processing device, and includes an image acquisition step of controlling a camera to acquire an image of a structure, an image processing step of generating 3D point cloud data including feature points of the structure based on the image, a structural information acquisition step of acquiring structural information including a 3D model of the structure, a map generation step of generating 3D point cloud data including feature points of the 3D model based on the structural information, and a matching step of matching the 3D point cloud data indicating the shape of the structure with the 3D point cloud data indicating the shape of the 3D model, wherein in the matching step, relative vector information including information on the distance or vector from the feature point to other feature points in the same point cloud data is generated for all the feature points, and the matching is performed based on the relative vector information. According to one embodiment, an information processing method is a method executed by an information processing device, and includes: a SLAM processing step of generating first three-dimensional point cloud data including feature points of a structure; an image acquisition step of controlling a camera to acquire an image of the structure; an image processing step of generating second three-dimensional point cloud data including feature points of the structure based on the image; a correction step of outputting the first three-dimensional point cloud data if the accuracy of the SLAM processing unit is within an acceptable range, and outputting the second three-dimensional point cloud data if the accuracy is outside the acceptable range; a structural information acquisition step of acquiring structural information including a three-dimensional model of the structure; a map generation step of generating three-dimensional point cloud data including feature points of the three-dimensional model based on the structural information; and a matching step of matching the three-dimensional point cloud data output by the correction unit with three-dimensional point cloud data indicating the shape of the three-dimensional model, wherein in the matching step, relative vector information including information on the distance or vector from the feature point to other feature points in the same point cloud data is generated for all the feature points, and the matching is performed based on the relative vector information. According to one embodiment, an information processing method is a method executed by an information processing device, and includes the steps of: arranging three-dimensional virtual grids in a three-dimensional model space of a structure; performing 360-degree spherical rendering or rendering in all directions using virtual cameras arranged on each grid; extracting feature points from the rendered image; and performing machine learning using the feature points extracted from the rendered image as training data and the coordinates of the virtual camera as training data to generate a trained model. According to one embodiment, the information processing method is a method executed by an information processing device, and includes an image acquisition step of controlling a camera to acquire an image of a structure, an image processing step of generating 3D point cloud data including feature points of the structure based on the image, a structural information acquisition step of acquiring structural information including a 3D model of the structure, a map generation step of generating 3D point cloud data including feature points of the 3D model based on the structural information, and a matching step of matching the 3D point cloud data indicating the shape of the structure with the 3D point cloud data indicating the shape of the 3D model, wherein in the matching step, relative vector information including information on the distance or vector from the feature point to other feature points in the same point cloud data is generated for all the feature points, and the matching is performed based on the relative vector information. According to one embodiment, a program causes a computer to perform any of the above methods. [Effects of the Invention]
[0012] The present invention can provide a terminal device, an information processing system, an information processing method, and a program that can accurately superimpose a structure and a three-dimensional model of structural information. [Brief explanation of the drawings]
[0013] [Figure 1] FIG. 2 is a diagram illustrating a hardware configuration of the terminal device 10 (20). [Figure 2] 2 is a block diagram showing the functional configuration of the terminal device 10. FIG. [Figure 3] 4 is a flowchart showing the operation of the terminal device 10. [Figure 4] FIG. 10 is a diagram illustrating the concept of matching processing. [Figure 5] 10A and 10B are diagrams illustrating a process of removing straight lines that have a low contribution to spatial recognition. [Figure 6] FIG. 2 is a block diagram showing the functional configuration of a terminal device 20. [Figure 7] 10 is a flowchart showing the operation of the terminal device 20. [Figure 8] 1 is a diagram illustrating a system configuration and a functional configuration of an information processing system 1. FIG. [Figure 9] FIG. 2 is a block diagram showing the functional configuration of a terminal device 20. [Figure 10] FIG. 10 is a diagram illustrating the concept of matching processing. [Figure 11] FIG. 10 is a diagram illustrating the concept of matching processing. DETAILED DESCRIPTION OF THE INVENTION
[0014] <<Embodiment 1>> FIG. 1 is a diagram illustrating a hardware configuration of a terminal device 10 according to a first embodiment of the present invention.
[0015] The terminal device 10 is an information processing device having a processor, a memory, a communication device, an input / output device, etc. The processor reads and executes a program stored in the memory, thereby logically realizing each functional unit described below.
[0016] The terminal device 10 is an information processing device used by a worker, and is typically a tablet computer, a smartphone, an HMD (e.g., Hololens, etc.), etc. In addition to the above-mentioned processor, memory, etc., the terminal device 10 may include hardware such as a display device that outputs information to the worker, an input device (touch panel, eye-gaze input device, pointing device, etc.) that accepts information input by the worker, a camera for capturing images, and various sensors for SLAM (gyro, acceleration sensor, depth sensor, ToF sensor, LiDAR, etc.).
[0017] 2 is a block diagram showing the functional configuration of the terminal device 10 according to the first embodiment. The terminal device 10 includes a display unit 101, an image acquisition unit 102, an image processing unit 103, a structural information acquisition unit 104, a map generation unit 105, a matching unit 106, and a superimposed display processing unit 107.
[0018] 3 is a flowchart showing the operation of the terminal device 10. The terminal device 10 operates in accordance with steps S101 to S106 shown below.
[0019] The display unit 101 performs processing to display an image on a display device (display). For example, the display unit 101 can display information to be superimposed on a structure, typically various drawings of the structure, a three-dimensional model, still images showing inspection results of the structure, moving images, character data, etc., on the display device so that the worker can visually confirm the information.
[0020] If the display device is a transmissive type, the worker can point the display device toward the structure and directly see the structure that is visible behind the display device (optical see-through). In this state, the display unit 101 can project desired information onto the display device, superimposing the information on the structure.
[0021] If the display device is a non-transmissive type, the display device can display in real time the image of the structure being captured by the imaging unit (described later) (video see-through). In this state, the display unit 101 can superimpose desired information onto the video see-through image and display it on the display device, thereby superimposing the information on the structure.
[0022] The image acquisition unit 102 controls the camera to capture an image of the structure and outputs an image (video) of the structure (step S101).
[0023] The image processing unit 103 generates three-dimensional point cloud data indicating the shape of the structure based on the image acquired by the image acquisition unit 102 (step S102). First, two-dimensional point cloud data is generated from the image, for example, by the following procedure. (1) Detect edges contained in the image. (2) Edge intersections are extracted as feature points. (3) The extracted feature points are collected and converted into 2D point cloud data.
[0024] Edge detection and feature point extraction from images can be performed using image recognition libraries such as OpenCV. Although OpenCV is a technology based on still images, it is also possible to generate 2D point cloud data from videos. For example, by extracting feature points from each frame of a video using OpenCV and then using optical flow technology to identify identical feature points across frames, it is possible to generate efficient 2D point cloud data from the video.
[0025] Next, the two-dimensional point cloud data is converted into three-dimensional point cloud data, for example, by the following procedure. (4) Applying SfM (Structure from Motion) to multiple video frames and adding 3D information to feature points generates 3D point cloud data. (5) Using SfM, the current position and posture at which the image is being captured are identified.
[0026] SfM is a technology that reconstructs the 3D shape of a structure from multiple still images taken of the structure from different viewpoints (positions or orientations). SfM not only makes it possible to determine the 3D coordinates of all 2D feature points contained in 2D point cloud data, but also the 3D coordinates and orientation of the position where the current image is being taken. The current position and orientation determined by SfM can be used as is as an estimation result of the self-position. Once matching (described later) is performed, it is also possible to convert the self-position and orientation into the coordinate system of the 3D model of the structural information.
[0027] The structural information acquisition unit 104 acquires structural information of the structure from a predetermined storage area (step S103). The structural information is, for example, a BIM / CIM model. The structural information includes information indicating the shape of the structure, specifically, a three-dimensional model. The predetermined storage area may be provided within the terminal device 10, or may be provided in another information processing device (such as a server) that can communicate with the terminal device 10.
[0028] The map generation unit 105 generates three-dimensional point cloud data based on the structural information acquired by the structural information acquisition unit 104 (step S104). For example, a large amount of three-dimensional point data indicating the vertices of various objects constituting the three-dimensional model is generated and collected to form the three-dimensional point cloud data. In this embodiment, the three-dimensional point cloud data generated by the map generation unit 105 is used as a map when superimposing and displaying information.
[0029] The matching unit 106 matches the 3D point cloud data generated by the image processing unit 103 with the 3D point cloud data generated by the map generation unit 105 (step S105). In other words, a process is performed to precisely overlay the 3D model of the structural information onto the real structure in the camera image. This process makes it possible to use the 3D point cloud data generated based on the structural information as a map for overlaying information in real space. In other words, if information is displayed at a predetermined position on the 3D model, the information will be accurately overlaid at the corresponding position on the real structure.
[0030] The matching unit 106 performs matching, for example, according to the following procedure. (1) For each point included in the 3D point cloud data generated by the image processing unit 103, the distance or vector to other points is calculated. This calculation is performed for all other points. Note that, to reduce the processing load, calculation may be performed only for other points that exist within a predetermined threshold (for example, within a distance of X meters). A data set that compiles all the calculated distance or vector information is called relative vector information. The relative vector information is linked to each point and saved, for example, as attribute information of the point data. (2) The same process as (1) is performed for each point included in the three-dimensional point cloud data generated by the map generating unit 105. (3) The relative vector information of each point included in the 3D point cloud data generated by the image processing unit 103 is compared with the relative vector information of each point included in the 3D point cloud data generated by the map generating unit 105, and points with similar relative vector information are matched. Matching of the two 3D point cloud data is completed by matching at least three pairs of points.
[0031] If at least three pairs of feature points can be matched, a transformation matrix (a rigid transformation matrix that converts between the feature points of the actual structure and the feature points of the 3D model of the structural information) can be obtained. Then, coordinate transformation can be performed for the other feature points using the transformation matrix. Similarly, coordinate transformation can be performed for the camera position. That is, the matching unit 106 can use the transformation matrix to convert the camera position and angle obtained by the image processing unit 103 using SfM into the coordinate system of the 3D model of the structural information.
[0032] The superimposed display processing unit 107 reads out information from a predetermined storage area and uses the 3D point cloud data generated by the map generation unit 105 as a map to superimpose the information on the structure (step S106). The information to be superimposed is assumed to be stored in advance in a predetermined storage area. The predetermined storage area may be provided within the terminal device 20, or may be provided in another information processing device (such as a server) that can communicate with the terminal device 20. The information may be, for example, an object that can be displayed in two or three dimensions, and specifically may be text, figures, photographs, videos, etc. The display position, size, angle, etc. of the information are typically assumed to be defined in advance in the coordinate system of the three-dimensional model of the structural information.
[0033] When displaying information, the superimposed display processing unit 107 determines the display position of the information by taking into account the offset amount according to the matching result by the matching unit 106. For example, assume that the display position A of certain information is defined as (XA, YA, ZA) in the coordinate system of the 3D model of the structural information. Now, the matching unit 106 matches point B (XB, YB, ZB) extracted from the image to display position A. Then, the superimposed display processing unit 107 converts the display position of the information from A to B. In other words, the offset amount (XB-XA, YB-YA, ZB-ZA) between point A and point B is added to the original display position A, and point B is determined to be the display position of the information. As a result, the information is superimposed and displayed in the correct position of the actual structure.
[0034] (Technical Effects) According to the first embodiment, if an image of a structure photographed on the spot is available, it can be matched with pre-given structural information. In the matching process, 3D point cloud data extracted from the image is compared with 3D point cloud data extracted from the structural information using relative vector information. This enables highly accurate matching.
[0035] <Variation 1> As a modification of the first embodiment, the matching unit 106 can perform matching (step S105) by, for example, the following procedure. (1) For each point included in the 3D point cloud data generated by the image processing unit 103, the distance or vector to other points is calculated. This calculation is performed for all other points. Note that, to reduce the processing load, calculation may be performed only for other points that exist within a predetermined threshold (for example, within a distance of X meters). A data set that compiles all the calculated distance or vector information is called relative vector information. The relative vector information is linked to each point and saved, for example, as attribute information of the point data. (2) The same process as (1) is performed for each point included in the three-dimensional point cloud data generated by the map generating unit 105. (3) The relative vector information calculated in (1) and (2) is converted into a histogram using statistical methods. (4) Histogrammed relative vector information is compared, and similar points are matched. Matching of two 3D point cloud data is completed by matching at least three pairs of points.
[0036] An example of a matching method that incorporates the statistical approaches (3) and (4) above is shown below. 1. To preprocess known points, the known points are converted from the local coordinate system defined within the model to a coordinate system suitable for processing. This allows position information to be handled using a unified standard. 2. To obtain the observation point (restoration point), the observation point is obtained as the world coordinate of the object placed in the scene. These are recorded consistently on a global basis. 3. To extract features and represent them using histograms, a histogram is generated for each point from two perspectives: "direction" and "distance." Angle histogram: Calculates the angular distribution based on the direction from a point to surrounding points and displays it as a histogram. Distance histogram: The distance from one point to another is normalized and the distribution is summarized as a histogram. These two histograms are concatenated to form a vector that describes the geometric features of each point. This feature vector is used to calculate the similarity between the known point and the observed point. 4. To determine the correspondence, we calculate the cost from the difference between feature vectors and create a correspondence cost matrix between points. Using the Hungarian method, we determine the optimal one-to-one correspondence that minimizes the overall cost. 5. To remove outliers and align the points, we randomly select a small number of corresponding points and perform multiple trials of tentative transformations. After each trial, we evaluate the consistency of the transformation, remove outliers, and adopt the transformation with the highest support. 6. To re-estimate the transformation with high accuracy, three points are selected from the reliable points after outlier removal and the transformation is re-estimated. The scale is estimated from the distance ratio and directional relationship, and a rotation matrix is derived based on the relationship between the points. This allows us to obtain highly accurate transformation parameters including position, rotation, and scale. 7. Apply the transformation to the model by applying the calculated transformation to the target model and adjusting the position, orientation, and size of the model. This automatically aligns the model to match the arrangement of the known points.
[0037] (Technical effect of Modification 1) As shown in Figure 4, it is expected that the position, angle, and scale of the 3D model of the structure information and the shape of the actual structure will differ. When matching feature points with different positions, angles, and scales, if relative vector information is used as is, a huge amount of preprocessing will be required. Therefore, by using relative vector information that has been histogrammed using a statistical method, as in Variation 1, it is possible to significantly reduce the processing load of matching.
[0038] <Variation 2> As a modification of the first embodiment, the image processing unit 103 can also create three-dimensional point cloud data (step S102) using the following method.
[0039] The image processing unit 103 generates three-dimensional point cloud data, for example, in the following procedure. (1) Detect edges contained in the image. (2) Of the straight lines detected as edges, extract the straight lines that intersect at a single point (vanishing point) when extended. (3) The extracted straight lines are excluded from the processing target, and the remaining edges are targeted, and the intersections of the edges are extracted as feature points. (4) The extracted feature points are collected and converted into 2D point cloud data. (5) Applying SfM (Structure from Motion) to multiple video frames and adding 3D information to feature points generates 3D point cloud data. (6) Using SfM, the position and orientation at which the image was taken are identified.
[0040] (Technical effect of Modification 2) As shown in Figure 5, artificial structures such as bridges contain many parallel lines. However, rather than using many feature points extracted from these parallel lines, it is believed that using only a few feature points extracted from elements with unique characteristics, such as diagonal truss members, can both reduce processing load and achieve accurate spatial recognition. Therefore, in Variation 2, parallel lines are excluded from the processing target as a preprocessing step. For example, among the edges detected from the image, lines converging toward a vanishing point in the bridge axis direction, lines converging toward a vanishing point perpendicular to the bridge axis direction, or lines converging toward a vertical direction in the bridge axis direction are considered to be parallel and vertical lines that make up the structure. Meanwhile, the remaining edges are considered to be highly likely to contribute (be meaningful) to spatial recognition, and feature points are extracted from these edges. This enables accurate matching while reducing processing load.
[0041] <Variation 3> In the first embodiment, feature points are extracted from the photographed image and the 3D model of the structure and then matched. In the third modification, feature points are extracted by incorporating line information and surface information from the photographed image and the 3D model of the structure and then matched.
[0042] Conventional SfM (Structure from Motion) estimates the camera position and angle using feature points in the real world. The method proposed in this requirement identifies the camera position and angle in the BIM model by performing SfM that incorporates BIM surface and line information, and then superimposes the real image and the BIM model image.
[0043] Ordinary SfM is applied to feature points with 3D information in world coordinates projected onto camera coordinates (plane: 2D) using a perspective projection matrix, but this requirement further introduces a conversion matrix from BIM plane coordinates or line coordinates to world coordinates, and applies it to feature points on a surface (2 degrees of freedom) or feature points on a line (1 degree of freedom) projected onto camera coordinates. In principle, this concept can also be applied to feature points on curved surfaces (2D) (Figure 10).
[0044] <<Embodiment 2>> In the first embodiment, the environmental map of the structure and the self-position estimation are performed by image processing. In the second embodiment, the environmental map of the structure and the self-position estimation are performed by SLAM.
[0045] However, with the SLAM function (ARKit (registered trademark), etc.) installed in smart devices such as tablet terminals, the effective range of the spatial recognition function is limited to a few meters around, and large structures such as bridges may not be recognized accurately. In the second embodiment, when such limitations prevent sufficient accuracy from being obtained in creating an environmental map and estimating the user's own position using SLAM, a technique is disclosed that uses image processing to complement the accuracy.
[0046] The hardware configuration of the terminal device 20 is the same as that of the terminal device 10 (see FIG. 1).
[0047] 6 is a block diagram showing a functional configuration of the terminal device 20 according to the second embodiment. The terminal device 20 includes a display unit 201, an image acquisition unit 202, an image processing unit 203, a structural information acquisition unit 204, a map generation unit 205, a matching unit 206, a superimposed display processing unit 207, a SLAM processing unit 208, and a correction unit 209.
[0048] 7 is a flowchart showing the operation of the terminal device 20. The terminal device 20 operates in accordance with steps S201 to S208 shown below.
[0049] Similar to the display unit 101, the display unit 201 performs processing to display an image on a display device (display).
[0050] The SLAM processing unit 208 executes the SLAM function to create an environmental map (map, 3D point cloud data) of the real space and estimate the vehicle's own position (step S201). SLAM is a well-known technology, so a detailed explanation of the SLAM algorithm will be omitted in this paper. The SLAM processing unit 208 can execute SLAM using information obtained from various sensors, such as a camera, a gyro sensor, an acceleration sensor, a depth sensor, a ToF (Time of Flight) sensor, and a LiDAR (Light Detection and Ranging).
[0051] The image acquisition unit 202, like the image acquisition unit 102, controls the camera to capture an image of the structure and outputs an image (video) of the structure (step S202).
[0052] The image processing unit 203 generates three-dimensional point cloud data indicating the shape of the structure based on the images acquired by the image acquisition unit 202, for example, by the following procedure (step S203). (1) Detect edges contained in the captured image. (2) Edge intersections are extracted as feature points. (3) The extracted feature points are collected and converted into 2D point cloud data. (4) Applying SfM (Structure from Motion) to multiple video frames and adding 3D information to feature points generates 3D point cloud data.
[0053] The structure information acquisition unit 204, like the structure information acquisition unit 104, acquires the structure information of the structure from a predetermined storage area (step S204).
[0054] Similar to the map generating unit 105, the map generating unit 205 generates three-dimensional point cloud data based on the structural information acquired by the structural information acquiring unit 104 (step S205).
[0055] When the correction unit 209 detects a decrease in accuracy of the environmental map creation and self-position estimation by the SLAM processing unit 208, it complements the 3D point cloud data generated by the SLAM processing unit 208 with the 3D point cloud data generated by the image processing unit 203 (step S206). For example, the correction unit 209 receives input of both the 3D point cloud data generated by the SLAM processing unit 208 and the 3D point cloud data generated by the image processing unit 203, and outputs the 3D point cloud data generated by the SLAM processing unit 208 to the matching unit 206 if the accuracy of the SLAM processing unit 208 is within an allowable range, and outputs the 3D point cloud data generated by the image processing unit 203 to the matching unit 206 if the accuracy is outside the allowable range.
[0056] For example, the accuracy of the SLAM processing unit 208 can be considered to have decreased in the following cases. When the estimated self-location result changes beyond a predetermined threshold within a predetermined time. When the estimated self-position falls within a predefined prohibited area. For example, a position that interferes with an object, or an impossible position such as in the air or underground can be defined as a prohibited area.
[0057] The matching unit 206 performs matching between the three-dimensional point cloud data output by the correction unit 209 and the three-dimensional point cloud data generated by the map generation unit 205 (step S207). Matching is performed using the same method as that used by the matching unit 106.
[0058] The superimposed display processing unit 207, like the superimposed display processing unit 107, reads information from a predetermined storage area and uses the three-dimensional point cloud data generated by the map generation unit 205 as a map to superimpose the information on the structure (step S208).
[0059] When displaying information, the superimposed display processing unit 207 takes into consideration the amount of offset according to the matching result by the matching unit 106 and determines the display position of the information.
[0060] (Technical Effects) According to the second embodiment, even if the accuracy of the SLAM of the terminal device 20 cannot be guaranteed, accurate information superimposition on the structure can be performed by performing complementation through image processing.
[0061] <Variation 4> As a modification of the second embodiment, the image processing unit 203 can also create the three-dimensional point cloud data (step S203) by the following method. Fig. 8 is a diagram showing the system configuration and functional configuration of the information processing system 1 according to the fourth modification. The information processing system 1 includes a terminal device 20 and an image processing server 30. (1) The image processing unit 203 of the terminal device 20 extracts still images from the captured images (moving images) at predetermined intervals (for example, 20 times per second) and transmits them to the image processing server 30 one by one. (2) The image processing server 30 uses known techniques such as ORB-SLAM (Oriented FAST and Rotated BRIEF SLAM) to sequentially extract feature points from successive still images, create an environmental map, and estimate the user's own position. (3) The image processing server 30 sequentially transmits the created environmental map (three-dimensional point cloud data) and its own position (camera coordinates) to the terminal device 20.
[0062] The image processing server 30 is typically a PC, and is an information processing device having a processor, memory, communication devices, input / output devices, etc. The processor reads and executes programs stored in the memory, thereby logically realizing each functional unit described below. The image processing server 30 may be realized in a cloud computing environment, or may be realized as an IoT device that can be directly connected to the terminal device 20.
[0063] ORB-SLAM is a well-known technology that creates an environmental map and estimates the user's position using only camera images. Typically, ORB-SLAM accumulates a series of still images extracted from a video or other source in advance and performs analysis using batch processing. In contrast, in Modification 1, the terminal device 20 sequentially generates a series of still images from the camera image and transmits them to the image processing server 30 running ORB-SLAM. The image processing server 30 then performs SLAM processing in real time using the still images it receives, and sequentially returns the processing results to the terminal device 20.
[0064] Note that the present invention is not limited to the use of ORB-SLAM, and any other technology that creates an environmental map and estimates the user's position from video can be adopted instead of ORB-SLAM.
[0065] (Technical effect of Modification 4) The processing performed by the terminal device 20 is limited to generating continuous still images, and real-time processing with a high processing load can be executed by the image processing server 30. This reduces the processing load on the terminal device 20. This technology makes it possible to use not only tablet devices and HMDs as the terminal device 20, but also devices with low processing capabilities, such as high-resolution single-lens reflex cameras and 360-degree panoramic cameras. For example, inspection robots and drones equipped with such devices can also be used as the terminal device 20.
[0066] <Variation 5> As a modification of the second embodiment, a technique is disclosed in which, when the accuracy of self-localization by SLAM is no longer sufficient, the accuracy is complemented by a method that combines simulation, image processing, and machine learning.
[0067] 9 is a block diagram showing the functional configuration of the terminal device 20 according to Modification 5. The terminal device 20 of Modification 5 includes a self-position estimation unit 210 instead of the image processing unit 203.
[0068] The self-position estimation unit 210 estimates the self-position, for example, in the following procedure. [Learning Phase] (1) The three-dimensional model space of the structural information is divided into virtual grids of a predetermined size (for example, 50 cm square). (2) A virtual camera is placed on each grid, and 360-degree spherical rendering or rendering in all directions is performed. (3) Extract feature points from a 360-degree spherical rendering image or a rendering image facing in all directions. (4) Machine learning is performed using the extracted feature points as training data and the camera coordinates as teacher data to generate a trained model.
[0069] [Estimation Phase] (5) Feature points are extracted from the photographed image (still image cut out from the video) of the structure acquired by the image acquisition unit 202. (6) The extracted feature points are input into the trained model, and the camera coordinates at which the image was taken are estimated. The self-location estimation results are output in the coordinate system of the 3D model.
[0070] The learning phase is executed in advance to generate a trained model. In the actual self-location estimation process, only the estimation phase is executed. The learning phase does not necessarily have to be executed in the terminal device 20; it may be executed in another information processing device, and only the trained model may be stored in a predetermined storage area of the terminal device 20 so that it can be used by the self-location estimation unit 210.
[0071] For 360-degree spherical rendering and rendering in all directions, it is also possible to perform nighttime rendering, rendering that takes into account weather, rendering after defining the position of sunlight according to the date, etc. This makes it possible to generate more realistic learning data that includes information on ambient light, weather, and shadows, for example.
[0072] It is also possible to reduce the processing load of learning and estimation by narrowing the learning and estimation range to, for example, only the area where an inspector is scheduled to visit, rather than the entire structure. For example, a three-dimensional area indicating the inspection range can be stored in advance in a predetermined memory area, and in the learning phase and estimation phase, the above processes (1) to (6) can be executed within the inspection range.
[0073] When the correction unit 209 detects a decrease in accuracy of the environmental map creation and self-position estimation by the SLAM processing unit 208, it complements the self-position estimated by the SLAM processing unit 208 with the self-position estimated by the self-position estimation unit 210. For example, the correction unit 209 receives input of both the self-position estimated by the SLAM processing unit 208 and the self-position estimated by the self-position estimation unit 210, and outputs the self-position generated by the SLAM processing unit 208 to the matching unit 206 if the accuracy of the SLAM processing unit 208 is within an allowable range, and outputs the self-position generated by the self-position estimation unit 210 to the matching unit 206 if the accuracy is outside the allowable range.
[0074] For example, the accuracy of the SLAM processing unit 208 can be considered to have decreased in the following cases. When the estimated self-location result changes beyond a predetermined threshold within a predetermined time. When the estimated self-position falls within a predefined prohibited area. For example, a position that interferes with an object, or an impossible position such as in the air or underground can be defined as a prohibited area.
[0075] (Technical effect of Modification 5) According to the fifth modification, even if the accuracy of the SLAM of the terminal device 20 cannot be guaranteed, it is possible to perform complementation by combining simulation, image processing, and machine learning, and to perform accurate information superimposition on structures.
[0076] The present invention is not limited to the above-described embodiments. Various modifications that can be understood by those skilled in the art can be made to the configurations and specific processing procedures of the embodiments without departing from the spirit of the present invention. Furthermore, each embodiment can be combined with other embodiments as appropriate.
[0077] Furthermore, the order of information processing shown in the above embodiments is not limited to the order described, as long as it does not deviate from the spirit of the present invention. The order of processing shown in each embodiment can be changed as appropriate within the scope that does not affect the processing results.
[0078] Each processing means constituting the present invention may be configured by hardware, or any process may be realized by having a CPU execute a computer program. Furthermore, the computer program may be stored and supplied to a computer using various types of temporary or non-temporary computer-readable media. Temporary computer-readable media include, for example, electromagnetic signals supplied to a computer via wire or wirelessly. [Explanation of symbols]
[0079] 1. Information Processing Systems 10 Terminal Equipment 101 Display section 102 Image acquisition unit 103 Image processing section 104 Structure information acquisition part 105 Map Generation Unit 106 Matching Section 107 Overlay display processing unit 20 Terminal equipment 201 Display section 202 Image acquisition unit 203 Image Processing Unit 204 Structure information acquisition unit 205 Map Generation Unit 206 Matching Department 207 Overlay display processing unit 208 SLAM processing section 209 Correction Unit 210 Self-position estimation section 30 Image Processing Server
Claims
1. an image acquisition unit that controls a camera to acquire an image of the structure; an image processing unit that generates three-dimensional point cloud data including feature points of the structure based on the image; a structural information acquisition unit that acquires structural information including a three-dimensional model of the structure; a map generation unit that generates three-dimensional point cloud data including feature points of the three-dimensional model based on the structural information; a matching unit that performs matching between three-dimensional point cloud data indicating the shape of the structure and three-dimensional point cloud data indicating the shape of the three-dimensional model, The matching unit generates relative vector information including information on a distance or vector from each feature point to another feature point in the same point cloud data for each feature point, and performs the matching based on the relative vector information. Terminal device.
2. The display unit further includes a display unit that uses three-dimensional point cloud data indicating the shape of the structure as a map and displays predetermined information superimposed on the structure. The terminal device according to claim 1.
3. The matching unit converts the relative vector information into a histogram and performs the matching based on the histogrammed relative vector information. The terminal device according to claim 1.
4. The image processing unit removes edges representing parallel or vertical straight lines that constitute the structure from the image, and extracts the feature points from the remaining edges. The terminal device according to claim 1.
5. The matching unit identifies the position and angle of a camera in the three-dimensional model using line information and surface information of the image and the three-dimensional model, and performs the matching. The terminal device according to claim 1.
6. a SLAM processing unit that generates first three-dimensional point cloud data including feature points of the structure; an image acquisition unit that controls a camera to acquire an image of the structure; an image processing unit that generates second three-dimensional point cloud data including feature points of the structure based on the image; a correction unit that outputs the first three-dimensional point cloud data when the accuracy of the SLAM processing unit is within an allowable range, and outputs the second three-dimensional point cloud data when the accuracy is outside the allowable range; a structural information acquisition unit that acquires structural information including a three-dimensional model of the structure; a map generation unit that generates three-dimensional point cloud data including feature points of the three-dimensional model based on the structural information; a matching unit that matches the three-dimensional point cloud data output by the correction unit with three-dimensional point cloud data that indicates the shape of the three-dimensional model, The matching unit generates relative vector information including information on a distance or vector from each feature point to another feature point in the same point cloud data for each feature point, and performs the matching based on the relative vector information. Terminal device.
7. A terminal device according to claim 6, an image processing server; the image acquisition unit acquires continuous still images of the structure and transmits the images to the image processing server; the image processing unit receives the second three-dimensional point cloud data from the image processing server; The image processing server The second three-dimensional point cloud data including the feature points of the structure is generated based on the continuous still images, and is transmitted to the image processing unit. Information processing system.
8. A self-position estimation unit is included instead of the image acquisition unit, The self-location estimation unit Placing a three-dimensional virtual grid in the three-dimensional model space of the structure; 360-degree spherical rendering or omnidirectional rendering is performed using virtual cameras placed on each grid. extracting feature points from the rendering image; a learning phase in which machine learning is performed using the feature points extracted from the rendering image as learning data and the coordinates of the virtual camera as training data to generate a learned model; extracting feature points from the captured image of the structure; an estimation phase in which the feature points extracted from the photographed image of the structure are input to the trained model, and a camera position of the photographed image of the structure is estimated as the self-position; When the correction unit detects a decrease in accuracy of the SLAM processing unit, the correction unit complements the self-position estimated by the SLAM processing unit with the self-position estimated by the self-position estimation unit in the estimation phase.
7. The terminal device according to claim 6.
9. A method executed by an information processing device, comprising: an image acquisition step of controlling the camera to acquire an image of the structure; an image processing step of generating three-dimensional point cloud data including feature points of the structure based on the image; a structural information acquisition step of acquiring structural information including a three-dimensional model of the structure; a map generation step of generating three-dimensional point cloud data including feature points of the three-dimensional model based on the structural information; a matching step of matching three-dimensional point cloud data indicating the shape of the structure with three-dimensional point cloud data indicating the shape of the three-dimensional model, In the matching step, relative vector information including information on the distance or vector from each feature point to another feature point in the same point cloud data is generated for each feature point, and the matching is performed based on the relative vector information. Information processing methods.
10. A method executed by an information processing device, comprising: a SLAM processing step of generating first three-dimensional point cloud data including feature points of the structure; an image acquisition step of controlling a camera to acquire an image of the structure; an image processing step of generating second three-dimensional point cloud data including feature points of the structure based on the image; a correction step of outputting the first three-dimensional point cloud data when the accuracy of the SLAM processing unit is within an allowable range, and outputting the second three-dimensional point cloud data when the accuracy is outside the allowable range; a structural information acquisition step of acquiring structural information including a three-dimensional model of the structure; a map generation step of generating three-dimensional point cloud data including feature points of the three-dimensional model based on the structural information; a matching step of matching the three-dimensional point cloud data output by the correction unit with three-dimensional point cloud data indicating the shape of the three-dimensional model, In the matching step, relative vector information including information on the distance or vector from each feature point to another feature point in the same point cloud data is generated for each feature point, and the matching is performed based on the relative vector information. Information processing methods.
11. A program for causing a computer to execute the method described in claim 9 or 10.
Citation Information
Patent Citations
Measurement system, measuring device, and measurement method
JP2019082400A
Information processing system, information processing method and program
JP7745042B1
Information processing method, program, and movable body control system
WO2024166659A1