Information processing method, information processing device, and program
By calculating the area and normal direction of 3D model faces and using a trained model to determine attribute information from optimized viewing angles, the method addresses the inefficiencies and inaccuracies of existing 3D model data input methods, providing cost-effective and accurate attribute assignment.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- CALTA INC
- Filing Date
- 2024-10-17
- Publication Date
- 2026-04-23
AI Technical Summary
Existing methods for inputting attribute information into 3D models, such as BIM, are time-consuming, costly, and prone to errors, and existing AI-based classification methods struggle to accurately extract detailed attribute information for a wide variety of 3D models.
A computer calculates the area and normal direction of each face in a 3D model, determines a viewing direction based on these calculations, acquires an image from that direction, inputs it into a trained model to obtain attribute information, and assigns it to the model.
This method reduces the effort and cost of manually inputting attribute information and minimizes errors, enabling accurate and efficient assignment of attribute information to a wide range of 3D models using AI.
Smart Images

Figure JP2024036981_23042026_PF_FP_ABST
Abstract
Description
Information Processing Method, Information Processing Apparatus, and Program
[0001] The present invention relates to an information processing method, an information processing apparatus, and a program for attaching attribute information to a 3D model.
[0002] Conventionally, software for inputting attribute information such as windows and doors into a 3D model such as BIM (Building Information Modeling) by a user's operation has been used.
[0003] Patent Document 1 discloses generating a 3D point cloud using a plurality of images taken from above, and generating a texture map representing the 3D shape of a roof based on the generated 3D point cloud. Further, Patent Document 1 discloses performing a coloring process in which roof features existing in an image are classified using artificial intelligence so that individual different roof structure features are uniquely colored, and applying a coloring classification label corresponding to the roof structure feature to the point cloud to indicate a specific roof structure feature in a specific color.
[0004] Japanese Patent Translation Publication No. 2023-505212
[0005] Inputting attribute information into a large number of 3D models by a user's operation in software is time-consuming and costly, and input errors may occur. Also, although there is a method of classifying roofs using artificial intelligence as described in Patent Document 1, it is difficult to extract detailed attribute information other than roofs with higher accuracy for a wide variety of 3D models, and there is a possibility of incorrect output.
[0006] Therefore, an object of the present invention is to provide an advantageous technique for attaching attribute information to a 3D model.
[0007] An information processing method, as one aspect of the present invention for solving the above problems, is an information processing method for assigning attribute information to a three-dimensional model of an object, characterized in that a computer calculates the area and normal direction of each face of the three-dimensional model, determines the viewing direction based on the calculated area and normal direction, acquires an image of the three-dimensional model viewed from the viewing direction, inputs the image viewed from the viewing direction into a trained model that outputs attribute information from an input image, obtains attribute information of the three-dimensional model, and assigns the attribute information to the three-dimensional model.
[0008] According to the present invention, it is possible to provide a technique that is advantageous for assigning attribute information to a three-dimensional model.
[0009] This is a diagram showing the configuration of the information processing device. This is a diagram showing a flowchart of the information processing method. This is a diagram representing the TIN model. This is a diagram representing the normal vector. This is a diagram showing each angular range of the normal vector. This is a diagram showing the concept of weight peaks. This is a diagram representing the 3D model of the building. This is a diagram representing the bounding box, line of sight, and viewpoint. This is a diagram showing a 2D image of the 3D model as seen from the line of sight direction. This is a diagram showing a 2D image of the 3D model as seen from a second line of sight direction. This is a diagram showing an image of the 3D model as seen from the top of the bounding box in the comparative example. This is a diagram showing an image of the 3D model as seen from the bottom of the bounding box in the comparative example.
[0010] Embodiments of the present invention include, for example, the following configurations.
[0011] [Item 1] An information processing method for assigning attribute information to a three-dimensional model of an object, characterized in that a computer calculates the area and normal direction of each face of the three-dimensional model, determines the viewing direction based on the calculated area and normal direction, acquires an image of the three-dimensional model viewed from the viewing direction, inputs the image viewed from the viewing direction into a trained model that outputs attribute information from an input image to obtain attribute information of the three-dimensional model, and assigns the attribute information to the three-dimensional model. [Item 2] The information processing method according to Item 1, characterized in that for each face of the three-dimensional model, the viewing direction is determined based on the direction of a first component vector, which is a normal vector corresponding to the face with the largest total area of faces facing the same direction. [Item 3] The information processing method according to Item 1, characterized in that the direction along the first component vector starting from the center of the three-dimensional model is determined as the viewing direction. [Item 4] The information processing method according to Item 1, characterized in that a plurality of viewing directions are determined, a plurality of images viewed from the plurality of viewing directions are input into the trained model to obtain attribute information corresponding to the plurality of images. [Item 5] The information processing method according to item 1, characterized in that a three-dimensional model of the object is superimposed on geospatial information, and attribute information is obtained from an image of the three-dimensional model including the geospatial information as viewed from the line of sight. [Item 6] A program for causing a computer to execute the information processing method according to any one of items 1 to 5. [Item 7] An information processing device for assigning attribute information to a three-dimensional model of an object, comprising a processing unit, wherein the processing unit calculates the area and normal direction of each face of the three-dimensional model, determines the line of sight based on the calculated area and normal direction, obtains an image of the three-dimensional model as viewed from the line of sight, inputs the image as viewed from the line of sight to a trained model that outputs attribute information from an input image, obtains attribute information of the three-dimensional model, and assigns the attribute information to the three-dimensional model.
[0012] Preferred embodiments of this disclosure will be described in detail below with reference to the attached drawings. In this specification and the drawings, components having substantially the same functional configuration are denoted by the same reference numerals, and redundant descriptions will be omitted.
[0013] <Details of the Embodiment> This embodiment describes a technique for assigning attribute information to objects represented by three-dimensional data such as a three-dimensional (3D) model or a three-dimensional point cloud. Examples of objects include, but are not limited to, buildings such as detached houses, apartments, condominiums, skyscrapers, shops, commercial facilities, and factories; structures such as roads, transportation facilities, bridges, dams, wind power generation facilities, power transmission towers, communication towers, and agricultural facilities; and natural objects such as topography, forests, and farmland (fields, rice paddies, and orchards). Attribute information differs from the shape and location of the object itself and provides descriptive data and characteristics about the object. It is used to understand, identify, classify, and analyze the object in more detail and can include various properties and states that are directly or indirectly related to the object. Examples of assigning attribute information to three-dimensional data include, but are not limited to, the following: For building data, attribute information is assigned to each component such as walls, windows, doors, entrances, balconies, roofs, chimneys, antennas, air conditioning equipment, decks, columns, beams, and solar panels. This information includes the type of material (reinforced concrete, wood, etc.), strength, height, year of construction, and number of floors. It also includes gardens, swimming pools, barns, and patios attached to the building. This information is used for structural analysis, maintenance planning, and urban planning. Geological attributes (type of soil, soil strength, groundwater location, etc.) are added to the topographic data. This information is used for civil engineering works and disaster risk assessment. Information such as plant and tree species, height, and health status is added to the forest data. This makes forest management and monitoring of vegetation health easier. As these examples show, by adding attribute information, 3D data gains more meaning and can be used as more useful information.
[0014] Figure 1 shows a configuration diagram of the information processing device 1 in this embodiment. The information processing device 1 has a computer that can be installed in a general-purpose computer such as a personal computer, or in a device such as a smartphone or tablet. The information processing device 1 executes an information processing method that assigns attribute information to an object represented as three-dimensional data using installed software or applications (programs).
[0015] The information processing device 1 includes a processing unit 11, a memory 12, a storage unit 13, a communication unit 14, an input unit 15, and a display unit 16. These are electrically connected to each other via a bus 17.
[0016] The processing unit 11 is a computing device that controls the operation of the entire information processing device 1, controls the transmission and reception of data between each part, and performs information processing necessary for program execution and authentication processing. The processing unit 11 includes a computing device such as a CPU, GPU, or FPGA, and executes programs stored in the storage 13 and loaded into the memory 12 to perform various information processing described later.
[0017] The memory 12 (storage unit) includes a main memory composed of a volatile storage device such as DRAM, and an auxiliary storage device composed of a non-volatile storage device such as flash memory or HDD. The memory 12 is used as a work area for the processing unit 11, and also stores the BIOS and various setting information that are executed when the information processing device 1 is started up.
[0018] The storage 13 (storage medium) includes storage devices such as HDDs and SSDs, and stores various programs such as applications and programs. In addition, a database containing data used for each process is built on the storage 13.
[0019] The communication unit 14 connects the information processing device 1 to a network. The communication unit 14 communicates with external devices directly or via a network access point using methods such as wired LAN, wireless LAN, Wi-Fi (Wireless Fidelity, registered trademark), infrared communication, Bluetooth (registered trademark), short-range or contactless communication. The input unit 15 is an information input device such as a keyboard, mouse, or touch panel.
[0020] The display unit (display device) 16 includes a display that shows information calculated by the processing unit 11 or information received from an external source by the communication unit 14. The display unit 16 provides a graphical user interface (GUI) that displays various types of information on its screen. The display unit 16 is not limited to being integrated with the information processing device 1, but may also be a display (display device) provided separately from the information processing device 1 and connected to the information processing device 1.
[0021] Bus 17 is connected in common to all of the above components and transmits, for example, address signals, data signals, and various control signals.
[0022] Next, we will explain an information processing method for assigning attribute information to an object represented by three-dimensional data. Figure 2 shows a flowchart of this method.
[0023] First, the information processing device 1 acquires 3D data for assigning attribute information and creates a 3D model (S1). The 3D data includes 3D point cloud data where each point cloud represents a 3D coordinate. Methods for creating this 3D data include laser scanning methods such as LiDAR, methods for calculating 3D data from multiple images or videos taken from multiple different angles such as Structure from Motion (SfM), methods for projecting patterned light (such as stripes) onto an object and capturing the distortion when the light reflects off the object with a camera to determine the 3D shape, and methods for calculating distance by shining light onto an object and measuring the time it takes for the light to reflect back.
[0024] From here on, we will explain the SfM method as an example. First, the target object is photographed using a camera (imaging device) on a mobile device such as a drone, and image data is acquired. The camera on the mobile device photographs the target object from various angles. The photography can be done by shooting video using the camera on the mobile device, or by taking multiple still images. Also, if a certain distance is maintained between the target object and the camera when taking photographs, the three-dimensional (3D) information will be reconstructed more accurately.
[0025] Furthermore, filming may be performed by the mobile body moving autonomously, by the mobile body being moved automatically by a control device, or by a person operating a control device to control the position and orientation of the mobile body. When the mobile body is moved automatically, the control device sets waypoints that include points along the path the mobile body is moving and controls the position of the mobile body so that it moves according to the waypoints. The control device transmits the waypoints of the mobile body as target positions to the mobile body and performs position control on the mobile body, for example, according to the difference between the target position and the measured current position. The control device can also control the shooting conditions of the mobile body's camera during video recording, such as the orientation (angle, direction), focal position, brightness, and sensitivity.
[0026] Image data captured by a mobile camera is transmitted to an information processing device (computer) that performs SfM. If necessary, pre-processing can be performed on the captured images before inputting the data into the SfM algorithm. For example, unwanted images may be removed, or the brightness and contrast of the images may be adjusted.
[0027] Next, the information processing device that performs SfM uses multiple captured or pre-processed images to perform SfM. First, it uses multiple images to detect feature points in the images. Feature points are unique points in an image that are easy to match from different viewpoints, such as the edges and corners of an object. Common algorithms include SIFT (Scale-Invariant Feature Transform) and SURF (Speeded Up Robust Features).
[0028] Next, the feature points detected from each image are mapped across multiple images. Common points in images taken from different viewpoints are matched to identify which parts of which images represent the same point. Feature-based methods (e.g., feature descriptors like SIFT or SURF) or nearest neighbor search algorithms are used for matching.
[0029] Then, the camera's position and orientation (angle, direction) at the time each image was taken is estimated from the positional information (2D coordinates) of the matching feature points in the two images. For example, a mathematical model called epipolar geometry is used to calculate the relative position and orientation of the cameras by integrating information from multiple viewpoints based on parallax. Basic matrices and essential matrices are calculated to estimate the relative position and orientation between cameras. Alternatively, the camera's position and orientation can be calculated using AI such as Gaussian Splatting. In this way, the camera's position and orientation in 3D space are determined, and the image and the camera's position and orientation at the time the image was taken can be associated and stored in a database or other storage unit.
[0030] Furthermore, a 3D point cloud is constructed using feature points matched from each viewpoint. Specifically, based on the estimated camera position and orientation, the coordinates of corresponding 3D points are calculated from the coordinates of matching 2D feature points across multiple images. This process is based on triangulation. This method allows us to determine where the features of an object seen in a 2D image are located in 3D space. The 3D point cloud is represented by 3D coordinates, with each point representing a position corresponding to the surface of the object. This allows us to grasp the approximate shape of the object. If necessary, by simultaneously adjusting the positions of all cameras and the 3D point cloud and performing bundle adjustment to minimize errors, more accurate camera positions and 3D point clouds can be obtained.
[0031] The 3D point cloud obtained at this stage is sparsely datated, so to obtain a denser point cloud, it is necessary to find and add more corresponding points. Algorithms such as stereo matching and optical flow can be used for this purpose. Stereo matching uses two or more image pairs to find corresponding points for each pixel and generate a dense point cloud. Optical flow calculates the movement of pixels between temporally consecutive images to generate a more detailed point cloud. It is also possible to obtain a dense 3D point cloud using methods such as Patch-based Multi-view Stereo (PMVS), which estimates the position of each point using the pixel information of the image. Furthermore, if the point cloud contains noise such as outliers, noise removal and optimization may be performed.
[0032] Furthermore, since the camera position and 3D point cloud calculated by SfM are estimated relatively, to obtain the actual size and distance, one method is to use external scale information (e.g., dimensions of known objects, GPS data). For example, when photographing the object, known survey points or targets can be set up in advance at the shooting site, and after SfM is completed, their coordinate values can be entered and fitted. Alternatively, targets such as AR markers that indicate arbitrary coordinates can be placed in the image data during SfM processing, and these coordinate values can be provided during SfM processing for fitting, thereby representing the output of the SfM processing in a real coordinate system.
[0033] Next, the information processing device 1 generates a three-dimensional model (3D model) by connecting points based on the constructed point cloud data. Modeling is a step to represent the shape smoothly by polygonizing the 3D point cloud, so that the detailed shape of the object is more clearly represented. The 3D model is composed of, for example, a TIN (Triangular Irregular Network) model. A TIN model is a model composed of multiple triangular planar figures that connect three points that are close together from the obtained three-dimensional point cloud, and these are connected to each other without intersecting, forming a digital data structure that represents three-dimensional data as a collection of triangles. The surface shape of the TIN model represents the surface shape of the measured object. Alternatively, the 3D model may be a model represented by two-dimensional polygonal figures such as quadrilaterals, or a model represented by a mesh of vertices, edges, and faces (polygons).
[0034] The 3D point cloud and 3D model obtained as described above can be output in various formats.
[0035] Next, the processing unit of the information processing device 1 overlays the acquired 3D model or 3D point cloud data with an image representing 2D geographic information (S2). Geographic information can be obtained from a Geographic Information System (GIS), which handles data containing various location-related information, including map information and various statistical information, on an electronic map. The overlay in S2 is not necessarily required if geographic information is not available or if it is not necessary for attribute information assignment.
[0036] Next, the processing unit of the information processing device 1 performs surface analysis on the acquired 3D model (S3). The 3D model is represented, for example, as a TIN model. A TIN model is composed of a set of faces where many triangular planes are adjacent, and each triangle has a unique area and face orientation. Figure 3 shows a diagram representing a part of a TIN model. A TIN model is represented as a set of triangles t formed by connecting each point p in the 3D point cloud with a line (edge) e. In S3, first, the area of each triangle constituting the model and the normal vector v perpendicular to the plane of each triangle are calculated using the model data. The area and normal vector can be calculated using the coordinates of the three vertices of the triangle. Here, the normal vector is represented as a vector with three-dimensional components and is a unit vector with magnitude r of 1. Figure 4 shows a diagram of the coordinate representation of the normal vector. The normal vector can be represented in coordinates (X1, Y1, Z1) of a 3D Cartesian coordinate system, or in spherical coordinates with azimuth angle φ and elevation angle Θ. The direction of the normal vector can be set to point outwards from the 3D model.
[0037] The processing unit of the information processing device 1 assigns the magnitude of the area as weight information to the calculated normal vector of each triangle in the plane. For example, the normal vector of a certain triangle t1 and the area of that triangle t1 are associated as a set of data, and the normal vector of another triangle t2 and the area of that triangle t2 are associated as a set of data, and the data is saved.
[0038] Next, the weights of multiple normal vectors are statistically processed. In the statistical processing, the sum of the magnitudes of the weights of the normal vectors contained in each range, which is obtained by dividing the three-dimensional space into ranges at regular intervals, is calculated. Each range can be set as a range by dividing the azimuth angle and elevation angle into ranges d at regular intervals, similar to dividing the surface of a sphere of radius 1 into ranges at regular angular intervals, as shown in Figure 5. For example, the azimuth angle can be divided into 360 equal parts to set ranges of 0 degrees to less than 1 degree, 1 degree to less than 2 degrees, ..., 359 degrees to less than 360 degrees, and the elevation angle can be similarly divided into 360 equal parts to set ranges where the azimuth angle and elevation angle are divided into 1-degree intervals. Alternatively, the angle can be rounded to the first decimal place to obtain an integer angle. That is, the range from 0 to less than 0.5 degrees can be set to 0, and the range from 0.5 to less than 1.5 degrees can be set to 1. Although the angle is expressed in degrees, it can also be expressed in radians. Furthermore, each range can be defined as a fixed interval in the X, Y, and Z coordinates of a three-dimensional Cartesian coordinate system, covering the surface of a sphere with radius 1. The sum of the weights in each range represents the sum of the area of the triangular faces with normal vectors included in that range.
[0039] The processing unit of the information processing device 1 can calculate statistical values such as the range in which the sum of the weights has a maximum value, the range in which the sum of the weights shows a peak, the mean, the variance, and the deviation by performing statistical processing of the weights. A peak is a location with a local maximum value in a local range, and may be a point where a value actually exists, a point interpolated from discrete points where a value actually exists, or a peak on a curve approximating discrete points. Figure 6 shows a diagram representing the concept of a peak. In Figure 6, the horizontal axis represents the azimuth angle and elevation angle in three-dimensional space, and the vertical axis represents the sum of the weights. Although Figure 6 is two-dimensional, the horizontal axis conceptually represents an angle in three-dimensional space. The component with the largest peak value (maximum value) can be named the first peak, the component with the second largest peak value can be named the second peak, and so on, from largest to smallest peak value. The normal vector of the angle corresponding to the first peak can be called the first component vector, the normal vector of the angle corresponding to the second peak can be called the second component vector, and the normal vector of the angle corresponding to the third peak can be called the third component vector. Furthermore, if two peaks are close together and have a range, the mean or median of that range can be used as the peak position.
[0040] For example, if the first peak has the maximum weighted value in the range of azimuth angle Θ 12 to 13 degrees and elevation angle φ 0 to 1 degree, representative values of the angle, such as the minimum, maximum, average, and median, are determined within that range. Then, the first component vector having the determined representative values for azimuth and elevation is determined. If the angle is expressed as an integer rounded to the first decimal place, the integer angle represents the representative value.
[0041] The first peak represents the maximum total area of triangles on the surface of the TIN model whose normal direction is oriented in the angular range corresponding to the first peak. Therefore, by extracting the first peak, it is possible to identify which direction of the faces constituting the 3D model has the largest total area. The first component vector is the normal vector corresponding to the face in the 3D model whose total area is largest when all faces are oriented in the same direction or within the same range.
[0042] Further, the second peak represents that, outside the first peak region, it is the maximum value where the total area of the triangles whose normal directions face in the angular range corresponding to the second peak is the second largest. Therefore, by extracting the second peak, it is possible to identify which direction the surface with the second largest area among the surfaces constituting the three-dimensional model faces. Thus, by performing statistical processing using the area of each surface of the three-dimensional model as a weight, it is possible to analyze the direction and size of the surfaces of the three-dimensional model.
[0043] The processing unit of the information processing apparatus 1 determines the line-of-sight direction and the viewpoint for viewing the three-dimensional model based on the statistical value of the result of statistically processing the weights (S4). The line-of-sight direction can be determined so that the attribute information can be correctly obtained when obtaining the attribute information in S6. The line-of-sight direction is set to the direction of viewing the three-dimensional data from the outside, that is, the direction from the outside of the object toward the center. The line-of-sight direction can be determined based on, for example, the direction of the first component vector. As an example, the direction opposite to the first component vector with the starting point being the center position of the three-dimensional model can be set as the line-of-sight direction. The center position of the three-dimensional model can be obtained as the center position of a rectangular parallelepiped (bounding box) surrounding the three-dimensional model. Thus, the direction along the first component vector with the center of the three-dimensional model as the starting point can be determined as the line-of-sight direction.
[0044] The position of the viewpoint that is the starting point of the line of sight may be an arbitrary position outside the bounding box surrounding the three-dimensional model. For example, at the time of image acquisition in S5, a position close to the three-dimensional model at a distance where the entire three-dimensional model is included within the viewing angle is preferable.
[0045] Note that the line-of-sight direction is not limited to the direction opposite to the first component vector described above as long as the attribute information can be correctly obtained when obtaining the attribute information in S6, and a direction shifted from that direction may also be used. The allowable range of the shift can be determined within the limit that the surface corresponding to the first component vector among the surfaces constituting the three-dimensional model is included in the two-dimensional image of the three-dimensional model when obtaining the two-dimensional image of the three-dimensional model. This is because the surface corresponding to the first component vector occupies the largest surface area in the three-dimensional model and is likely to represent the features and attributes of the three-dimensional model.
[0046] The line-of-sight direction may also be determined based on, among other things, the second component vector or the third component vector. For example, the direction opposite to the second component vector or the third component vector with the starting point being the center position of the three-dimensional model can be set as the line-of-sight direction. Also, the line-of-sight direction may be set as a plurality of directions.
[0047] The line-of-sight direction may be set such that the viewpoint is above the center of the three-dimensional model (object). Many three-dimensional data are calculated from images viewed from above in the sky or vertically upward. If the viewpoint is below the center, the image viewed from that viewpoint has less information, and an image lacking information is likely to occur.
[0048] Next, a two-dimensional image of the three-dimensional model viewed from the determined viewpoint and line-of-sight direction is obtained (S5). A two-dimensional image can be obtained by placing a plane perpendicular to the line-of-sight direction between the viewpoint and the three-dimensional model and projecting the image of the three-dimensional model onto that plane.
[0049] Next, using a learned model that outputs attribute information from an input image, the attribute information of the three-dimensional model (object) is obtained from the two-dimensional image of the three-dimensional model viewed from the line-of-sight direction (S6). The learned model that outputs attribute information from an input image is a model that has been mechanically learned using a large number of two-dimensional images and a large number of attribute information as teacher data, and is created by artificial intelligence (AI) such as a large-scale neural network. The learning model may be configured within the information processing device 1 or may be provided by an external computer separate from the information processing device 1. Examples of external computers include servers and public clouds. In S6, when the two-dimensional image obtained in S5 is input to the learned model that outputs attribute information from an input image, the learned model outputs the attribute information of things included in the image, etc. The output may be a description text of the image or may be in the form of keywords.
[0050] If multiple viewing directions are set, multiple images are acquired from each viewing direction and these multiple images are input into the trained model. The trained model outputs attribute information corresponding to each image, but all of the output attribute information may be used as the object's attribute information, only the attribute information common to all images may be selected as the object's attribute information, or the attribute information may be determined by majority vote.
[0051] Next, attribute information output from the trained model is attached to the 3D model (object) (S7). Methods for attaching attribute information include, for example, tagging the 3D model or linking with a database. Attribute information can be tagged for each part of the 3D model corresponding to each attribute information and saved in a database or file format. For example, attribute data can be directly linked to each part of the 3D model. File formats include the IFC format for BIM data and Shapefile for GIS. With database linkage, attribute data related to 3D data can be saved in a separate database, and that attribute data can be retrieved when a specific part of the 3D data is selected.
[0052] (Example 1) As an example of a 3D model, Figure 7 shows a diagram of a 3D model of a building. In S3, when surface analysis of the 3D model is performed, it is calculated that the area of the surface having a normal direction within the same angular range as the normal direction of the entrance surface shown in Figure 7 is the largest. Therefore, the first component vector with the same angular range as the normal direction of the entrance surface, starting from the center position of the 3D model, is determined, and the opposite direction of this first component vector is determined as the line of sight direction. In Figure 8, a bounding box BB that is tangent to and surrounds the 3D model is shown by a solid line. The line of sight direction VI is the opposite direction of the first component vector starting from the center position CP of the bounding box BB. The viewpoint VP can be determined so that the 3D model of the building is included within the field of view.
[0053] Figure 9 shows a 2D image of the 3D model as seen from the determined line of sight direction VI and viewpoint VP. The image in Figure 9 includes the front of the building, including the entrance. The image in Figure 9 is then input into the pre-trained model to obtain attribute information for the 3D model. The pre-trained model, with the image in Figure 9 as input, outputs attribute information such as a white entrance, windows, chimney, blue exterior walls, a brown roof, and stairs.
[0054] Figure 10 shows a 2D image of the 3D model viewed from the line of sight, where the line of sight is defined as the opposite direction of the second component vector starting from the center position CP of the rectangular prism BB. When the image in Figure 10 is input to the trained model, it outputs attribute information such as windows, chimney, blue exterior wall, and brown roof.
[0055] As shown in the example above, when comparing attribute information obtained from an image of a 3D model viewed from a viewing direction determined based on the first component vector with attribute information obtained from an image of a 3D model viewed from a viewing direction determined based on the second component vector, the first component vector yields more attribute information and more accurately represents the 3D model. Therefore, obtaining attribute information from only the image of the 3D model viewed from a viewing direction determined based on the first component vector requires less computation and less time to obtain accurate attribute information, compared to obtaining attribute information from 2D images of the 3D model viewed from each of multiple viewing directions.
[0056] As described above, according to the above embodiment, by using a trained model that outputs attribute information from input images, the effort and cost of inputting attribute information into a large number of 3D models through user operation in the software can be reduced, and input errors can be suppressed. Furthermore, in the above embodiment, by analyzing the surfaces of the 3D model, determining the viewing direction of the 3D model based on the analysis results, and acquiring a 2D image of the 3D model viewed from the determined viewing direction, an image can be obtained from which attribute information can be more accurately determined. In this way, even when using artificial intelligence and a trained model that outputs attribute information from input images for a wide variety of 3D models, the possibility of outputting incorrect attribute information is reduced, and attribute information can be determined more accurately. Therefore, correct attribute information can be added to a wide variety of 3D data more easily, creating new methods of data utilization and improving the efficiency of data-driven work.
[0057] For example, by adding attributes such as windows and doors, as well as information about the location of windows and doors, to a 3D model of a building, in addition to information such as the building's height and roof, this information can be used for building maintenance, drone delivery services, autonomous driving services, and architectural planning.
[0058] (Comparative Example) Next, a comparative example will be explained. In the comparative example, we will explain an example in which the viewing direction of the 3D model of the building in Figure 7 is determined based on the normal direction of the faces of the rectangular prism BB that surrounds and touches the 3D model. The rectangular prism BB has six faces, and the normal directions of the six faces, that is, the up and down, left and right, and front and back directions, can be set as the viewing direction. If the viewing direction is the opposite direction of the outward normal vector of the top surface of the rectangular prism BB, starting from the center position CP of the rectangular prism BB, the image of the 3D model viewed from a viewpoint at a certain distance will be as shown in Figure 11. When the image in Figure 11 is input into the trained model described above to obtain attribute information of the 3D model, the brown roof and chimney are output, but the windows, entrance, exterior walls and stairs are not output because they do not exist in the image.
[0059] If the line of sight is defined as the opposite direction of the outward normal vector of the bottom surface of the rectangular prism BB, starting from the center position CP of the rectangular prism BB, then the image of the 3D model viewed from a viewpoint at a certain distance will be as shown in Figure 12. In the image in Figure 12, the shape of the building's base can be recognized, but because actual measurement data is not available, the image is lacking in information. When the image in Figure 12 is input into the trained model described above, the information that it is a building is output, but detailed attribute information about the building's components is not output.
[0060] As another example, if a 3D object has an inclined surface such as a slope, setting a bounding box for the object will result in the object being viewed from the side of the slope when viewed from the side of the bounding box. In this way, 3D objects can be diverse, so if a bounding box is set for an object and the viewing direction is set based on the surface of the bounding box, the image viewed from that viewing direction may not have accurate attribute information. If the viewing direction is not set appropriately, even if an image of the 3D model is input into the trained model described above, it will not be possible to obtain appropriate attribute information.
[0061] Preferred embodiments of this disclosure have been described in detail above with reference to the attached drawings, but the technical scope of this disclosure is not limited to these examples. Furthermore, not all components shown in the embodiments are essential components of this disclosure. Also, features shown in each embodiment are applicable to other embodiments insofar as they do not contradict each other.
[0062] The information processing device 1 is not limited to a single device; it may consist of multiple distributed information processing devices, or it may be configured as a virtual server or container in the cloud.
[0063] When a 3D model is superimposed on an image representing 2D geographic information, the image viewed from the line of sight determined in S4 contains 2D geographic information. Therefore, when this image is input to the trained model described above, attribute information related to the 2D geographic information is output. For example, in the case of a 3D model of a building that is known to be close to the coastline from the geographic information, attribute information of "a building on the coast" is obtained. In this way, by superimposing a 3D model of an object onto geospatial information, attribute information can be obtained from an image of the 3D model containing geospatial information, viewed from the determined line of sight.
[0064] This disclosure relates to an information processing method, an information processing device, and a program for assigning attribute information to a three-dimensional model, and has industrial applicability.
[0065] 1. Information Processing Device
Claims
1. An information processing method for assigning attribute information to a three-dimensional model of an object, characterized in that a computer calculates the area and normal direction of each face of the three-dimensional model, determines the viewing direction based on the calculated area and normal direction, acquires an image of the three-dimensional model viewed from the viewing direction, inputs the image viewed from the viewing direction into a trained model that outputs attribute information from an input image to obtain attribute information of the three-dimensional model, and assigns the attribute information to the three-dimensional model.
2. The information processing method according to claim 1, characterized in that, for each face of the three-dimensional model, the line of sight direction is determined based on the direction of the first component vector of the normal vector corresponding to the face with the largest total area obtained by combining faces facing the same direction.
3. The information processing method according to claim 1, characterized in that the direction along the first component vector, which starts from the center of the three-dimensional model, is determined as the line of sight direction.
4. The information processing method according to claim 1, characterized in that a plurality of viewing directions are determined, a plurality of images viewed from the plurality of viewing directions are input to the trained model, and attribute information corresponding to the plurality of images is obtained.
5. The information processing method according to claim 1, characterized in that a three-dimensional model of the object is superimposed on geospatial information, and attribute information is obtained from an image of the three-dimensional model including the geospatial information as viewed from the line of sight.
6. A program for causing a computer to execute the information processing method described in any one of claims 1 to 5.
7. An information processing device for assigning attribute information to a three-dimensional model of an object, comprising a processing unit, the processing unit calculating the area and normal direction of each face of the three-dimensional model, determining the viewing direction based on the calculated area and normal direction, acquiring an image of the three-dimensional model viewed from the viewing direction, inputting the image viewed from the viewing direction into a trained model that outputs attribute information from an input image to obtain attribute information of the three-dimensional model, and assigning the attribute information to the three-dimensional model.
Citation Information
Patent Citations
System and method for modeling structures using point clouds derived from stereoscopic image pairs - Patent Application 20070122997
JP2023505212A
Building model reconstruction method and system
CN115661378A
Image output device and image output program
JP2017167933A
Semantic Fusion
JP2022531536A