Information processing method, information processing device, and program

KR103003654B1Active Publication Date: 2026-08-11CALTA INC
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
KR1020257009159
Authority / Receiving Office
KR · KR
Patent Type
Patents
Current Assignee / Owner
Filing Date
2024-10-17
Publication Date
2026-08-11
Estimated Expiration
2044-10-17

Smart Images

  • Figure 112025031541454-PCT00002_ABST
    Figure 112025031541454-PCT00002_ABST
Patent Text Reader

Abstract

A technology advantageous for assigning attribute information to a 3D model is provided. In an information processing method for assigning attribute information to a 3D model of an object, the area and normal direction of each face of the 3D model are calculated, the line of sight is determined based on the size of the calculated area and the normal direction, an image of the 3D model viewed from the line of sight is acquired, and the image viewed from the line of sight is input to a trained model that outputs attribute information from the input image to obtain attribute information of the 3D model and assign attribute information to the 3D model.
Need to check novelty before this filing date? Find Prior Art

Description

Technology Field

[0001] The present invention relates to an information processing method, an information processing device, and a program for assigning attribute information to a three-dimensional model. Background Technology

[0002] Conventionally, software that inputs attribute information such as windows and doors into 3D models, such as BIM (Building Information Modeling), through user operation is used.

[0003] Patent Document 1 discloses generating a three-dimensional point cloud using a plurality of images taken from above, and generating a texture map representing the three-dimensional shape of a roof based on the generated three-dimensional point cloud. Additionally, Patent Document 1 discloses classifying roof features present in the images using artificial intelligence, performing a coloring process such as uniquely coloring individual different roof structure features, and applying a coloring classification label corresponding to the roof structure features to the point cloud to represent specific roof structure features in a specific color. Prior art literature

[0004] Patent Document 1: Japanese Patent Publication No. 2023-505212 The problem to be solved

[0005] Inputting attribute information into multiple 3D models through user operation in software is laborious and costly, and input errors may occur.

[0006] In addition, as described in Patent Document 1, there is a method for classifying roofs using artificial intelligence, but for various 3D models, it is difficult to extract detailed attribute information other than the roof with high precision, and there is a possibility of producing incorrect output.

[0007] Therefore, the present invention aims to provide a technology advantageous for assigning attribute information to a three-dimensional model. means of solving the problem

[0008] An information processing method as one aspect of the present invention for solving the above problem is an information processing method for assigning attribute information to a three-dimensional model of an object, wherein a computer calculates the area and normal direction of each face of the three-dimensional model, determines a viewing direction based on the size of the calculated area and the normal direction, acquires an image of the three-dimensional model viewed from the viewing direction, inputs the image viewed from the viewing direction to a learning completed model that outputs attribute information from an input image, obtains attribute information of the three-dimensional model, and assigns attribute information to the three-dimensional model. Effects of the invention

[0009] According to the present invention, a technique advantageous for assigning attribute information to a three-dimensional model can be provided. Brief explanation of the drawing

[0010] Figure 1 is a diagram illustrating the configuration of an information processing device. Figure 2 is a diagram illustrating the flowchart of an information processing method. Figure 3 is a diagram showing the TIN model. Figure 4 is a diagram showing the normal vector. Figure 5 is a diagram showing the angle range of the normal vector. Figure 6 is a diagram showing the concept of the peak of the weight. Figure 7 is a drawing showing a three-dimensional model of a building. Figure 8 is a diagram showing the bounding box, line of sight, and viewpoint. Figure 9 is a drawing showing a 2D image of a 3D model viewed from the line of sight. Figure 10 is a drawing showing a 2D image of a 3D model viewed from a second line of sight. Figure 11 is a drawing showing an image of a three-dimensional model viewed from the upper surface direction of a bounding box in a comparative example. Figure 12 is a drawing showing an image of a three-dimensional model viewed from the lower direction of the bounding box in the comparative example. Specific details for implementing the invention

[0011] An embodiment of the present invention comprises, for example, the following configuration.

[0012] [Item 1]

[0013] In an information processing method for assigning attribute information to a three-dimensional model of an object,

[0014] The computer,

[0015] Calculate the area and normal direction of each face of the above 3D model, and determine the line of sight based on the size of the calculated area and the normal direction,

[0016] Acquire an image of the above 3D model viewed from the above line of sight, and

[0017] The image viewed from the above-mentioned line of sight is input into a trained model that outputs attribute information from an input image to obtain attribute information of the above-mentioned 3D model, and

[0018] Information processing method characterized by assigning the above attribute information to the above three-dimensional model.

[0019] [Item 2]

[0020] An information processing method described in Item 1, characterized by determining the line of sight direction based on the direction of a first component vector, which is a normal vector corresponding to the face with the largest sum of areas from faces facing the same range of directions, for each face of the above-described three-dimensional model.

[0021] [Item 3]

[0022] An information processing method described in Item 1, characterized by determining the direction along the first component vector, starting from the center of the above three-dimensional model, as the line of sight direction.

[0023] [Item 4]

[0024] An information processing method described in Item 1, characterized by determining a plurality of the above-mentioned gaze directions, inputting a plurality of images viewed from the plurality of the above-mentioned gaze directions into a learning completed model, and obtaining attribute information corresponding to the plurality of images.

[0025] [Item 5]

[0026] An information processing method described in Item 1, characterized by superimposing a three-dimensional model of the object onto geospatial information and obtaining attribute information from an image of the three-dimensional model containing the geospatial information viewed from the line of sight.

[0027] [Item 6]

[0028] A program for executing on a computer the information processing method described in any one of items 1 through 5.

[0029] [Item 7]

[0030] In an information processing device that assigns attribute information to a three-dimensional model of an object,

[0031] It has a processing unit, and the processing unit,

[0032] Calculate the area and normal direction of each face of the above 3D model, and determine the line of sight based on the size of the calculated area and the normal direction,

[0033] Acquire an image of the above 3D model viewed from the above line of sight, and

[0034] The image viewed from the above-mentioned line of sight is input into a trained model that outputs attribute information from an input image to obtain attribute information of the above-mentioned 3D model, and

[0035] An information processing device characterized by assigning the above attribute information to the above three-dimensional model.

[0036] Suitable embodiments of the present disclosure will be described in detail below with reference to the attached drawings. In addition, in this specification and drawings, components having substantially the same functional configuration are given the same reference numerals to avoid redundant descriptions.

[0037] <Details of Embodiment>

[0038] In this embodiment, a technique for assigning attribute information to an object represented by three-dimensional data, such as a three-dimensional (3D) model or a three-dimensional point cloud, is described. Examples of objects include buildings such as single-family homes, apartments, mansions, high-rise buildings, shops, commercial facilities, and factories; structures such as roads, transportation facilities, bridges, dams, wind power generation facilities, transmission towers, communication towers, and agricultural facilities; and natural objects such as terrain, forests, and farmland (fields, rice paddies, orchards), but are not limited to these. Attribute information provides descriptive data or features regarding an object, unlike the shape or location of the object itself, and is used to understand the object in more detail, and to identify, classify, and analyze it. It may include various properties or states directly or indirectly related to the object.

[0039] Examples of assigning attribute information to 3D data include, but are not limited to, the following. For building data, attribute information is assigned to each component, such as walls, windows, doors, entrances, balconies, roofs, chimneys, antennas, HVAC equipment, decks, columns, beams, and solar panels. This information includes the type of material (reinforced concrete, wood, etc.), strength, height, installation year, and number of floors. Additionally, gardens, pools, barns, and patios attached to the building are also included. This is utilized for structural analysis of buildings, maintenance planning, and urban planning. Furthermore, geological attributes (type of strata, soil strength, location of groundwater, etc.) are assigned to topographic data. This is utilized for civil engineering works and disaster risk assessment. For forest data, information such as the type, height, and health status of plants and trees is assigned. This facilitates forest management and monitoring of vegetation health. As shown in these examples, the assignment of attribute information gives 3D data a deeper meaning and allows it to be utilized as more useful information.

[0040] FIG. 1 illustrates a configuration diagram of an information processing device (1) in the present embodiment. The information processing device (1) has a computer equipped in a device such as a general-purpose computer, a personal computer, a smartphone, or a tablet. The information processing device (1) executes an information processing method that assigns attribute information to an object appearing as three-dimensional data by means of installed software or an application (program).

[0041] The information processing device (1) has a processing unit (11), a memory (12), a storage (13), a communication unit (14), an input unit (15), and a display unit (16). These are electrically connected to each other through a bus (17).

[0042] The processing unit (11) is an arithmetic unit that controls the operation of the entire information processing device (1), controls the transmission and reception of data between each unit, and performs information processing necessary for program execution and authentication processing. The processing unit (11) includes an arithmetic processing unit, such as a processor, CPU, GPU, or FPGA, and executes a program stored in storage (13) and deployed in memory (12) to perform various information processing described later.

[0043] The memory (12) (memory unit) includes a main memory composed of a volatile memory device such as DRAM, and an auxiliary memory composed of a non-volatile memory device such as flash memory or HDD. The memory (12) is used as a work area of ​​the processing unit (11), and also stores BIOS and various setting information that are executed when the information processing device (1) starts up.

[0044] Storage (13) (storage medium) includes storage devices such as HDDs and SSDs, and stores various programs such as applications and programs. In addition, a database storing data used for each process is built in the storage (13).

[0045] The communication unit (14) connects the information processing device (1) to a network. The communication unit (14) communicates directly with an external device or through a network access point, for example, via a wired LAN, wireless LAN, Wi-Fi (Wireless Fidelity, registered trademark), infrared communication, Bluetooth (registered trademark), short-range or contactless communication, etc. The input unit (15) is an information input device, for example, a keyboard, mouse, touch panel, etc.

[0046] The display unit (display device) (16) is equipped with a display that displays information obtained by processing unit (11) or information received from the outside by communication unit (14). The display unit (16) provides a graphical user interface (GUI) that displays various information on its screen. The display unit (16) is not limited to being installed integrally with the information processing device (1), but may be a display (display device) that is installed separately from the information processing device (1) and connected to the information processing device (1).

[0047] The bus (17) is connected in common to each of the above-mentioned components and transmits, for example, address signals, data signals and various control signals.

[0048] Next, an information processing method for assigning attribute information to an object represented as three-dimensional data is described. A flowchart of the above method is illustrated in FIG. 2.

[0049] First, the information processing device (1) acquires three-dimensional data to assign attribute information and creates a three-dimensional model (S1). As three-dimensional data, there is three-dimensional point cloud data in which each point cloud represents three-dimensional coordinates. To create these three-dimensional data, methods include a measurement method using laser scanning such as LiDAR, a method of calculating three-dimensional data from multiple images or moving images taken from multiple different angles such as Structure from Motion (SfM), a method of obtaining a three-dimensional shape by projecting patterned light (such as a striped shape) onto an object and capturing the distortion when the light is reflected from the object with a camera, and a method of calculating distance by irradiating light onto an object and measuring the time until it is reflected back from the object.

[0050] Next, as an example, a method using SfM will be explained. First, image data is acquired by taking a picture of an object using a camera (imaging device) of a moving object (moving device) such as a drone. The camera of the moving object takes a picture of the object from various angles. The picture may be taken by taking a moving image using the camera of the moving object, or by taking multiple still images. In addition, when taking the picture, if a certain distance is maintained between the object and the camera, the three-dimensional (3D) information is reconstructed more accurately.

[0051] In addition, filming may be performed by the moving body moving autonomously, by automatically moving the moving body using a control device, or by a person controlling the position and attitude of the moving body by operating the control device. When the moving body is moved automatically, the control device sets waypoints including points on the path of movement of the moving body and controls the position of the moving body so that the moving body moves according to the waypoints. The control device transmits the waypoints of the moving body to the moving body as target positions and performs position control for the moving body, for example, based on the difference between the target position and the measured current position. In addition, the control device may also control filming conditions such as the attitude (angle, direction), focus position, brightness, and sensitivity of the camera of the moving body during video recording.

[0052] Image data captured by the camera of the moving object is transmitted to an information processing device (computer) that executes SfM. Additionally, if necessary, preprocessing may be performed on the images obtained from the capture before inputting the data into the SfM algorithm. For example, unnecessary images may be removed, or the brightness or contrast of the images may be adjusted.

[0053] Next, an information processing device executing SfM performs SfM using multiple captured images or preprocessed images. First, feature points of the images are detected using multiple images. A feature point is a point that is unique within an image and easy to match even from different viewpoints, such as an edge or corner of an object. Common algorithms include SIFT (Scale-Invariant Feature Transform) and SURF (Speeded Up Robust Features).

[0054] Next, feature points detected from each image are matched across multiple images. Common points within images captured from different viewpoints are matched to identify which parts of which images represent the same point. For this matching, feature-based methods (e.g., feature descriptors such as SIFT or SURF) or nearest neighbor search algorithms are used.

[0055] Then, the position and orientation (angle, direction) of the camera when each image is captured are estimated from the position information (2D coordinates) of the feature points matched in the two images. For example, using a mathematical model called epipolar geometry, the relative position and orientation of the camera are calculated by integrating information from multiple viewpoints based on parallax. The relative position and orientation between cameras are estimated by calculating a basis matrix or an essential matrix. Alternatively, the position and orientation of the camera can be calculated using AI such as Gaussian Splatting. In this way, the position and orientation of the camera in 3D space are determined, and the image is correlated with the position and orientation of the camera when the image is captured, allowing them to be stored in a memory such as a database.

[0056] In addition, a 3D point cloud is constructed using the feature points matched at each viewpoint. Specifically, based on the estimated camera position and orientation, the coordinates of the corresponding 3D points are calculated from the coordinates of the 2D feature points that match between multiple images. This process is a method based on triangulation. Through this method, the location of the features of an object reflected in a 2D image within 3D space is determined. The 3D point cloud is represented by 3D coordinates, and each point indicates a position corresponding to the location on the surface of the object. This allows for the identification of the approximate shape of the object. If necessary, more accurate camera positions and 3D point clouds can be obtained by performing bundle adjustments that minimize errors by simultaneously adjusting the positions of all cameras and the 3D point cloud.

[0057] Since the 3D point cloud obtained at this stage is sparse data, it is necessary to find more corresponding points and add them to obtain a dense point cloud. Algorithms such as stereo matching and optical flow can be used for this purpose. In stereo matching, two or more image pairs are used to find corresponding points for each pixel, thereby generating a dense point cloud. In optical flow, pixel movement between temporally consecutive images is calculated to generate a more detailed point cloud. Additionally, a dense 3D point cloud can be obtained by using methods such as Patch-based Multi-view Stereo (PMVS), which estimates the location of each point using image pixel information. Furthermore, if the point cloud contains noise such as outliers, optimization may be performed by removing the noise.

[0058] In addition, since camera positions and 3D point clouds calculated by SfM are estimated relatively, there is a method to use external scale information (e.g., dimensions of known objects, GPS data) to obtain actual sizes or distances. For example, when photographing an object, known survey points or targets are prepared in advance at the shooting site, and after SfM is completed, the coordinate values ​​are input and fitted. Alternatively, targets such as AR markers, which have known coordinates, are placed in the image data during SfM processing, and the output of the SfM processing is represented in a real-world coordinate system by assigning those coordinate values ​​and fitting during SfM processing.

[0059] Next, the information processing device (1) creates a three-dimensional model (3D model) by connecting points based on the constructed point cloud data. Modeling is a step for smoothly expressing the shape by polygonizing the 3D point cloud, and the detailed shape of the object appears more clearly. The 3D model is composed, for example, of a TIN (Triangular Irregular Network) model. The TIN model is a model composed of a plane formed by connecting multiple triangular planar figures that connect three points close to each other among the obtained 3D point cloud, without intersecting, and is a digital data structure that represents 3D data as a set of triangles. The plane shape of the TIN model represents the surface shape of the measured object. In addition, the 3D model may be a model expressed as a two-dimensional polygonal figure such as a rectangle, or a model expressed as a mesh of vertices, edges, and faces (polygons).

[0060] The three-dimensional point cloud and 3D model obtained as described above are output in various formats.

[0061] Next, the processing unit of the information processing device (1) superimposes an image representing two-dimensional geographic information with the acquired 3D model or three-dimensional point cloud data (S2). The geographic information can be obtained from a Geographic Information System (GIS) that handles data containing various information regarding location, including map information and various statistical information, on an electronic map. The superimposition of S2 is not necessarily required, such as when geographic information is not obtained or when it is not necessary for assigning attribute information.

[0062] Next, the processing unit of the information processing device (1) performs surface analysis on the acquired 3D model (S3). The 3D model is represented, for example, as a TIN model. A TIN model is composed of a set of faces where a plurality of triangle planes are adjacent, and each triangle has a unique area and a face direction. Figure 3 shows a diagram showing a part of the TIN model. The TIN model is represented as a set of triangles t formed by connecting each point p of a 3D point group with a line (side) e. In S3, first, using the data of the model, the area of ​​each triangle constituting the model and the normal vector v perpendicular to the plane of each triangle are calculated. The area and the normal vector can be calculated using the coordinates of the three vertices of the triangle. Here, the normal vector is represented as a vector having 3D components and is a unit vector with a magnitude r of 1. Figure 4 shows a diagram of the coordinate representation of the normal vector. The normal vector can be represented as coordinates (X1, Y1, Z1) in a 3D Cartesian coordinate system, or as spherical coordinates with an azimuth angle φ and an elevation angle Θ. The direction of the normal vector can be set to face outward from the 3D model.

[0063] The processing unit of the information processing device (1) assigns the size of the area to the calculated normal vector for each triangle plane as information of weight. For example, the normal vector of a triangle t1 and the area of ​​triangle t1 are associated with data as a set of sets, and the normal vector of another triangle t2 and the area of ​​triangle t2 are associated with data as a set of sets, thereby preserving the data.

[0064] Then, the weights of multiple normal vectors are processed statistically. In the statistical processing, the sum (total sum) of the magnitudes of the weights of the normal vectors included in each range is calculated for each range in which the three-dimensional space is divided at regular intervals. Each range can be set as a range in which the azimuth and elevation angles are each divided at regular intervals d, such as the range in which the surface of a sphere with radius 1 is divided at regular angle intervals as shown in FIG. 5. For example, the azimuth angle can be divided into 360 equal parts to set a range of 0 degrees or more and less than 1 degree, a range of 1 degree or more and less than 2 degrees, ..., 359 degrees or more and less than 360 degrees, and the elevation angle can be similarly divided into 360 equal parts to set a range in which the azimuth and elevation angles are divided at 1-degree intervals. Alternatively, the first decimal place of the angle may be rounded to an integer angle. That is, the range of 0 to less than 0.5 degrees can be set to 0, and the range of 0.5 degrees or more and less than 1.5 degrees can be set to 1. Angles are expressed in degrees, but may also be in radians.

[0065] In addition, each range may be set at equal intervals of the X, Y, and Z coordinates of a 3D orthogonal coordinate system to cover the surface of a sphere with radius 1. The sum of the magnitudes of the weights in each range represents the magnitude of the sum of the areas of the faces of triangles having normal vectors included in that range.

[0066] The processing unit of the information processing device (1) can calculate statistical values ​​such as the range where the sum of the weights has a maximum value, the range where the sum of the weights shows a peak, the average value, the variance value, and the deviation by statistical processing of the weights. The peak is a location that has a maximum value in a local range, and it may be a point where the value actually exists, a point that interpolates a discrete point where the value actually exists, or a peak on a curve that approximates the discrete point. A diagram showing the concept of a peak is illustrated in FIG. 6. The horizontal axis of FIG. 6 represents the angles of azimuth and elevation in three-dimensional space, and the vertical axis represents the sum of the weights. FIG. 6 is two-dimensional, but the horizontal axis conceptually represents the angle in three-dimensional space. The component with the largest peak value (maximum value) can be named the first peak, the component with the second largest peak value can be named the second peak, and the peaks can be named the first peak, second peak, and third peak in order from the largest peak value. The normal vector at the angle corresponding to the first peak can be called the first component vector, the normal vector at the angle corresponding to the second peak can be called the second component vector, and the normal vector at the angle corresponding to the third peak can be called the third component vector. In addition, if two peaks are close together and have a width over a certain range, the average value or the median value of that range can be set as the peak position.

[0067] For example, if the weighted sum has a maximum first peak in the range of azimuth Θ 12 to 13 degrees and elevation φ 0 to 1 degree, representative values ​​of the angle, such as the minimum, maximum, average, and median values ​​of the angle in that range, are determined. Then, a first component vector having the azimuth and elevation angles of the determined representative values ​​is determined. If the angle is expressed as an integer by rounding to the first decimal place, the integer angle represents the representative value.

[0068] The first peak indicates that among the surfaces of the TIN model, the area of ​​the sum of triangles facing the direction of the angle range corresponding to the first peak is maximum. Therefore, by extracting the first peak, it is possible to determine which direction the area of ​​the face constituting the 3D model is facing is maximum. The first component vector is the normal vector corresponding to the face with the largest sum of areas from faces facing the same direction or the same range of directions for each face of the 3D model.

[0069] In addition, the second peak indicates that, outside the first peak region, the area of ​​the sum of triangles whose normal direction is oriented in the direction of the angle range corresponding to the second peak is the second largest maximum value. Therefore, by extracting the second peak, it is possible to determine which direction the area of ​​the face constituting the 3D model is oriented to have the second maximum value. In this way, by performing statistical processing using the area of ​​each face of the 3D model as a weight, it is possible to analyze the direction and size of the faces of the 3D model.

[0070] The processing unit of the information processing device (1) determines the viewing direction and viewpoint for the 3D model based on the statistical value of the result of statistical processing of the weights (S4). The viewing direction can be determined so that the attribute information can be accurately obtained when obtaining attribute information in S6. The viewing direction is set as the direction of viewing the 3D data from the outside, that is, the direction from the outside of the object toward the center. The viewing direction can be determined, for example, based on the direction of the first component vector. As an example, the direction in the reverse direction of the first component vector, with the starting point as the center position of the 3D model, can be set as the viewing direction. The center position of the 3D model can be obtained as the center position of the rectangular prism (bounding box) surrounding the 3D model. In this way, the direction following the first component vector with the center of the 3D model as the starting point can be determined as the viewing direction.

[0071] The position of the point of view that serves as the starting point of the line of sight may be any position outside the bounding box surrounding the 3D model, but, for example, when acquiring an image of S5, a distance at which the entire 3D model is included within the field of view and a position close to the 3D model is preferred.

[0072] In addition, when obtaining attribute information in S6, the line of sight direction is not limited to the inverse direction of the first component vector mentioned above, provided that the attribute information can be obtained accurately, and may be a direction deviated from that direction. The allowable range of deviation can be determined to the extent that, when obtaining a 2D image of the 3D model, the surface corresponding to the first component vector among the surfaces constituting the 3D model is included in the image. This is because the surface corresponding to the first component vector occupies the maximum surface area in the 3D model and is easy to represent the features and attributes of the 3D model.

[0073] In addition, the line of sight direction may be determined based on a second component vector or a third component vector. For example, the direction in the reverse direction of a second component vector or a third component vector with the starting point as the center position of a three-dimensional model may be set as the line of sight direction. Furthermore, the line of sight direction may be set as multiple directions.

[0074] It is advisable to set the viewing direction so that the viewpoint is above the center of the 3D model (object). This is because 3D data is often generated from images viewed from above in the vertical direction, and if the viewpoint is below the center, the image viewed from that point contains less information, making it easy for an image with insufficient information to occur.

[0075] Next, a 2D image of the 3D model is obtained from the determined viewpoint and the direction of view (S5). A 2D image can be obtained by placing a plane perpendicular to the direction of view between the viewpoint and the 3D model and projecting the image of the 3D model onto that plane.

[0076] Next, attribute information of a 3D model (object) is obtained from a 2D image viewed from the direction of the gaze of the 3D model using a trained model that outputs attribute information from an input image (S6). The trained model that outputs attribute information from an input image is a model that is mechanically trained using multiple 2D images and multiple attribute information as teaching data, and is created by artificial intelligence (AI), such as a large-scale neural network. The training model may be configured within the information processing device (1) or provided by an external computer separate from the information processing device (1). The external computer may be in the form of a server or a public cloud. In S6, when the 2D image obtained in S5 is input to the trained model that outputs attribute information from an input image, the trained model outputs attribute information such as an object included in the image. The output may be in the form of a description of the image or in the form of keywords.

[0077] If multiple viewing directions are set, multiple images viewed from each viewing direction are acquired and input into the trained model. The trained model outputs attribute information corresponding to each image; however, all output attribute information may be used as the attribute information of the target object, only the attribute information common to each image among the output attribute information may be determined as the attribute information of the target object, or the attribute information may be determined by majority vote.

[0078] Next, attribute information output from the trained model is applied to a 3D model (object) (S7). As a method for applying attribute information, there are methods such as tagging the 3D model or linking with a database. For each part of the 3D model corresponding to each attribute information, attribute information can be tagged and stored in a database or file format. For example, attribute data can be directly linked to each part of the 3D model. File formats include the IFC format of BIM data or the Shapefile of GIS. In the case of linking with a database, attribute data related to 3D data can be stored in a separate database, and the attribute data can be retrieved when a specific part of the 3D data is selected.

[0079] (Example 1)

[0080] As an example of a three-dimensional model, a drawing of a three-dimensional model of a building is shown in FIG. 7. In S3, when a face analysis of the three-dimensional model is performed, it is calculated that the area of ​​the face having a normal direction within an angle range identical to the normal direction of the face of the entrance shown in FIG. 7 is maximized. Therefore, a first component vector within an angle range identical to the normal direction of the face of the entrance, starting from the center position of the three-dimensional model, is obtained, and the inverse direction of the first component vector is determined as the line of sight direction. In FIG. 8, a rectangular prism (bounding box) (BB) that encloses and touches the three-dimensional model is shown by solid lines. The inverse direction of the first component vector starting from the center position (CP) of the rectangular prism (BB) is the line of sight direction (VI). The viewpoint (VP) can be determined so that the three-dimensional model of the building is included within the angle of view.

[0081] Figure 9 shows a 2D image of a 3D model viewed from a determined line of sight (VI) and viewpoint (VP). The image in Figure 9 includes the front of a building, including an entrance. Then, the image in Figure 9 is input into the above-mentioned trained model to obtain attribute information of the 3D model. The trained model, with the image in Figure 9 input, outputs white entrance, windows, chimneys, blue exterior walls, brown roofs, and stairs as attribute information.

[0082] FIG. 10 illustrates a 2D image of a 3D model viewed from the line of sight, in which the inverse direction of the second component vector, starting from the center position (CP) of the rectangular prism (BB), is the line of sight. The trained model, with the image of FIG. 10 input, outputs a window, a chimney, a blue exterior wall, and a brown roof as attribute information.

[0083] As shown in the example above, when comparing the attribute information obtained from the image of the 3D model viewed from the viewing direction determined based on the first component vector with the attribute information obtained from the image of the 3D model viewed from the viewing direction determined based on the second component vector, the first component vector provides more attribute information and accurately represents the 3D model. Therefore, compared to the method of obtaining attribute information from the 2D image of the 3D model viewed from each of the multiple viewing directions, obtaining attribute information only from the image of the 3D model viewed from the viewing direction determined based on the first component vector allows for obtaining accurate attribute information with less computation and in a shorter amount of time.

[0084] As described above, according to the above-described embodiment, by using a fully trained model that outputs attribute information from an input image, the effort and cost of inputting attribute information into multiple 3D models through user operation in software can be reduced, and input errors can be suppressed. Furthermore, in the above-described embodiment, an analysis of the faces of a 3D model is performed, and based on the analysis results, a viewing direction for the 3D model is determined. By acquiring a 2D image of the 3D model viewed from the determined viewing direction, an image from which attribute information can be obtained more accurately can be obtained. By doing so, for various types of 3D models, even when using artificial intelligence and a fully trained model that outputs attribute information from an input image, the possibility of outputting incorrect attribute information is reduced, and attribute information can be obtained more accurately. Therefore, accurate attribute information can be assigned to various types of 3D data more easily, new data utilization methods can be created, and the efficiency of work utilizing data can be improved.

[0085] For example, regarding a 3D model of a building, in addition to information such as the building's height or roof, information such as attributes like windows or doors, or the location of windows or doors, can be provided, and that information can be utilized for building maintenance, drone delivery services, autonomous driving services, and architectural planning review.

[0086] (Comparative Example)

[0087] Next, a comparative example is described. In the comparative example, an example is described in which the viewing direction of the 3D model of the building in FIG. 7 is determined based on the normal direction of the face of the rectangular prism (BB) that encloses the 3D model. The rectangular prism (BB) has six faces, and the normal directions of the six faces—namely, up and down, left and right, and front and back—can be set as the viewing direction. When the inverse direction of the outward normal vector of the top face of the rectangular prism (BB), starting from the center position (CP) of the rectangular prism (BB), is set as the viewing direction, the image of the 3D model viewed from a certain distance is as shown in FIG. 11. When the image of FIG. 11 is input into the above-mentioned completed model to obtain attribute information of the 3D model, a brown roof and a chimney are output, but windows, entrances, exterior walls, and stairs are not output because they do not exist in the image.

[0088] When the inverse direction of the outward normal vector of the lower surface of the rectangular prism (BB), starting from the center position (CP) of the rectangular prism (BB), is taken as the line of sight, the image of the 3D model viewed from a certain distance is as shown in FIG. 12. In the image of FIG. 12, the shape of the floor surface of the building can be recognized, but since actual measurement data does not exist, it is an image lacking information. When the image of FIG. 12 is input into the above-mentioned trained model, information that it is a building is output, but attribute information regarding the detailed components of the building is not output.

[0089] As another example, if an object in 3D data has an inclined surface such as a bevel, setting a bounding box for the object results in the object being viewed from the back of the bevel when viewed from the horizontal direction of the bounding box. As such, since there can be a wide variety of objects in 3D data, if a bounding box is set for an object and the viewing direction is set based on the surface of the bounding box, the image viewed from the viewing direction may become an image from which attribute information cannot be accurately obtained. If the viewing direction is not properly set, even if an image of the 3D model is input into the above-mentioned trained model, appropriate attribute information cannot be obtained.

[0090] Although suitable embodiments of the present disclosure have been described in detail above with reference to the accompanying drawings, the technical scope of the present disclosure is not limited to these examples. Furthermore, not all components appearing in the embodiments are essential components of the present disclosure. Additionally, features appearing in each embodiment may be applied to other embodiments as long as they do not contradict each other.

[0091] The information processing device (1) is not limited to a single device and may be composed of a distributed plurality of information processing devices, or may be composed of a virtual server or container in the cloud.

[0092] When a 3D model is superimposed with an image representing 2D geographic information, the image viewed from the line of sight determined in S4 contains 2D geographic information. Therefore, when that image is input into the model that has completed the learning process, attribute information related to the 2D geographic information is output. For example, in the case of a 3D model of a building that can be identified as being close to the coastline from the geographic information, attribute information such as "a building in the coastal area" is obtained. In this way, by superimposing a 3D model of an object onto geospatial information, attribute information can be obtained from an image of a 3D model containing geospatial information viewed from a determined line of sight. Industrial applicability

[0093] The present disclosure relates to an information processing method, an information processing device, and a program for assigning attribute information to a three-dimensional model, and has industrial applicability. Explanation of the symbols

[0094] 1: Information processing device

Claims

Claim 1 An information processing method for assigning attribute information to a three-dimensional model of an object, wherein a computer calculates the surface area of ​​each surface of the three-dimensional model and the normal direction of each surface, determines a viewing direction based on the normal direction to which a weight is assigned based on the size of the calculated surface area for each surface, acquires an image of the three-dimensional model viewed from the viewing direction, inputs the image viewed from the viewing direction to a learned model that outputs attribute information from an input image to obtain attribute information of the three-dimensional model, and assigns the attribute information to the three-dimensional model. Claim 2 An information processing method according to claim 1, characterized by determining the line of sight direction based on the sum of the surface areas of the faces facing the same range of directions among each surface of the three-dimensional model. Claim 3 An information processing method according to claim 2, characterized in determining the line of sight direction based on the direction of the first component vector of the normal vector corresponding to the face with the largest sum of surface areas oriented in the same range for each surface of the three-dimensional model. Claim 4 An information processing method according to claim 3, characterized in that the direction along the first component vector, starting from the center of the three-dimensional model, is determined as the line of sight direction. Claim 5 An information processing method according to claim 1, characterized by adding a surface area weight to a normal vector for each surface of the three-dimensional model and determining the line of sight direction based on the value of the weight of the normal vector. Claim 6 An information processing method according to claim 5, characterized by determining the line of sight direction based on the sum of the weights of the normal vectors whose direction is within a predetermined range. Claim 7 An information processing method according to claim 6, characterized by determining the line of sight direction based on the peak of the sum value for the angle. Claim 8 An information processing method according to claim 1, characterized by outputting attribute information of a specific part when a specific part of the three-dimensional model to which attributes are assigned is selected. Claim 9 A program stored on a medium for executing an information processing method described in any one of paragraphs 1 through 8 on a computer. Claim 10 An information processing device for assigning attribute information to a three-dimensional model of an object, comprising a processing unit, wherein the processing unit calculates the surface area of ​​each surface of the three-dimensional model and the normal direction of each surface, determines a viewing direction based on the normal direction to which a weight is assigned by the size of the calculated surface area for each surface, acquires an image of the three-dimensional model viewed from the viewing direction, inputs the image viewed from the viewing direction to a learning completed model that outputs attribute information from an input image to obtain attribute information of the three-dimensional model, and assigns the attribute information to the three-dimensional model.

Citation Information

Patent Citations

  • Building model reconstruction method and system

    CN115661378A

  • Image output device and image output program

    JP2017167933A

  • Data processing method

    KR100167699B1