A semantic real - scene three - dimensional modeling method for urban buildings
By using OSM data and street scene image data to impart semantic description information to the three-dimensional model of urban buildings, the problem of lack of semantic description in the existing models is solved, and high-precision smart city application needs are achieved.
Patent Information
- Application Number
- CN202510145729.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-10
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2045-02-10
AI Technical Summary
The existing three-dimensional model of urban buildings generated using mass source data lacks semantic description information and cannot meet the high-precision needs of smart city applications.
By constructing a three-dimensional model of each building in the city based on OSM data, and using street view image data to obtain building feature description information and text information, determining the building function type based on a predefined building function description dictionary, and assigning values to the three-dimensional model to provide semantic description.
It realizes the semantic description information for the three-dimensional model of urban buildings, so that the model can meet the high-precision needs of smart city applications, and improves the usability and reference value of the model.
Smart Images

Figure CN119600608B_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the technical field of 3D modeling, and particularly relates to a semantic real - scene 3D modeling method for urban buildings. Background Art
[0002] With the rapid advancement of the global urbanization process, urban building planning, management, and construction are facing increasingly complex challenges. Especially in dynamic and complex urban environments, refined 3D models of urban buildings have become an important data foundation for supporting the development of smart cities. 3D models of urban buildings can not only intuitively display the structure of urban space and the distribution of buildings, but also provide data support for emergency response, urban planning, navigation services, environmental monitoring, etc. Traditional construction of 3D models of urban buildings usually relies on expensive equipment and time - consuming surveying processes, such as LiDAR (Light Detection and Ranging) scanning, satellite remote sensing, and aerial photogrammetry. This makes the construction and update costs of urban models high, and it is difficult to keep up with the rapid pace of urban changes. The rise of Internet crowdsourced data has provided a new direction for solving this problem. Crowdsourced data refers to data spontaneously provided by Internet users or devices (such as mobile terminals, Internet of Things devices). This type of data has advantages such as large scale, low cost, and real - time nature. Specifically, crowdsourced data can include geographical location information, geometric structure information, etc. of geographical entities. For example, the crowdsourced annotation data of OpenStreetMap (OSM) contains information such as building geometric outlines. These rich Internet data provide an efficient and dynamic data source for the construction and update of 3D models of urban buildings, thus breaking through the limitations of traditional surveying methods.
[0003] However, currently, the 3D models of buildings generated using crowdsourced data lack semantic description information, resulting in the 3D models of buildings not meeting the high - precision requirements of smart city applications. Summary of the Invention
[0004] The embodiments of this application provide a semantic real - scene 3D modeling method for urban buildings, which can solve the problem that the 3D models of buildings cannot meet the high - precision requirements of smart city applications.
[0005] The embodiments of this application provide a semantic real - scene 3D modeling method for urban buildings, including:
[0006] Construct 3D models of each building in the city according to the OSM data of the city;
[0007] Based on the road network data of the city, determine multiple street - view image sampling positions, and collect street - view image data in multiple directions at each street - view image sampling position;
[0008] For each street view image data, obtain the building feature description information and the text information set of the street view image data, and determine the building function type corresponding to the street view image data according to the building feature description information, the text information set, and a predefined building function description dictionary; the building feature description information is used to describe the relationship between the buildings in the street view image data and other entities, the text information set includes the text information on the building facades in the street view image data, and the building function description dictionary includes the function description information corresponding to multiple building function types;
[0009] According to all the determined building function types and multiple street view image sampling positions, assign semantic description information to the 3D models of each building in the city to obtain the semantic real-scene 3D models of each building in the city; the semantic description information is used to describe the attributes of the buildings.
[0010] Optionally, determine multiple street view image sampling positions based on the road network data of the city, including:
[0011] Set multiple street view image sampling positions on each road in the road network data; the distance between two adjacent street view image sampling positions on the same road is a preset distance.
[0012] Optionally, collect street view image data in multiple directions at each street view image sampling position, including:
[0013] For each street view image sampling position, calculate the azimuth of the road where the street view image sampling position is located, and crawl from the Internet the street view image data in the front, back, left, and right directions taken at the street view image sampling position based on the azimuth.
[0014] Optionally, obtain the building feature description information and the text information set of the street view image data, including:
[0015] Use a text and image multi-modal model to analyze and process the street view image data to obtain the building feature description information of the street view image data;
[0016] Use OCR technology to perform text detection and recognition on the street view image data to obtain the text information set of the street view image data.
[0017] Optionally, determine the building function type corresponding to the street view image data according to the building feature description information, the text information set, and a predefined building function description dictionary, including:
[0018] Fuse the building feature description information and the text information set through string splicing to obtain a fused text description;
[0019] Calculate the similarity between the fused text description and each function description information in the building function description dictionary;
[0020] Use the building function type corresponding to the function description information with the highest similarity to the fused text description as the building function type corresponding to the street view image data.
[0021] Optionally, according to all the determined building function types and multiple street view image sampling positions, assign semantic description information to the 3D models of each building in the city, including:
[0022] For each street view image sampling position respectively, take the area centered on the street view image sampling position with a preset buffer distance as the buffer area of the street view image sampling position, and assign the building function types corresponding to all the street view image data corresponding to the street view image sampling position as semantic description information to the 3D models of all the buildings within the buffer area.
[0023] Optionally, the semantic real - scene 3D modeling method further includes:
[0024] If there is a building in each building in the city that is not located within any buffer area, determine the target buffer area closest to the building from all the buffer areas, and assign the building function type corresponding to the target buffer area as semantic description information to the 3D model of the building.
[0025] Optionally, other entities include roads, trees, sidewalks, road signs, and vehicles.
[0026] Optionally, the semantic real - scene 3D modeling method further includes:
[0027] When the update cycle arrives, update the OSM data of the city and the street view image data corresponding to each street view image sampling position, and update the semantic real - scene 3D models of each building in the city based on the updated OSM data and street view image data.
[0028] The above - mentioned solution of this application has the following beneficial effects:
[0029] In the embodiments of this application, by constructing 3D models of each building in the city based on Internet crowd - sourced data, innovatively obtaining building feature description information and text information on the building facades from street view image data, then combining the building feature description information and the text information on the building facades to determine the semantic description information of each building in the city, and by assigning the semantic description information to the 3D models of the corresponding buildings, a building 3D model with semantic description information can be obtained, so that the building 3D model can meet the high - precision requirements of smart city applications.
[0030] Other beneficial effects of this application will be described in detail in the subsequent detailed implementation section. Brief Description of the Drawings
[0031] To more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0032] Figure 1 It is a flowchart of a semantic real - scene three - dimensional modeling method for urban buildings provided by an embodiment of the present application;
[0033] Figure 2 It is a schematic diagram of three - dimensional models of each building in the city provided by an embodiment of the present application;
[0034] Figure 3 It is a schematic diagram of a buffer area in the city provided by an embodiment of the present application;
[0035] Figure 4 It is a schematic diagram of semantic real - scene three - dimensional models of each building in the city provided by an embodiment of the present application. Detailed implementation manners
[0036] In the following description, for the purpose of illustration rather than limitation, specific details such as specific system structures and technologies are presented to thoroughly understand the embodiments of the present application. However, those skilled in the art should clearly understand that the present application can also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well - known systems, devices, circuits, and methods are omitted to avoid unnecessary details from interfering with the description of the present application.
[0037] It should be understood that when used in the specification of the present application and the appended claims, the term "comprising" indicates the presence of the described features, wholes, steps, operations, elements, and / or components, but does not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components, and / or their combinations.
[0038] It should also be understood that the term "and / or" used in the specification of the present application and the appended claims refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations.
[0039] As used in the specification of this application and the appended claims, the term "if" may be construed, depending on the context, as "when", or "once", or "in response to determining", or "in response to detecting". Similarly, the phrase "if determined" or "if [the described condition or event] is detected" may be construed, depending on the context, to mean "once determined", or "in response to determining", or "once [the described condition or event] is detected", or "in response to detecting [the described condition or event]".
[0040] In addition, in the description of the specification of this application and the appended claims, the terms "first", "second", "third", etc. are only used for distinguishing descriptions and cannot be construed as indicating or implying relative importance.
[0041] Reference to "one embodiment" or "some embodiments" or the like described in the specification of this application means that a specific feature, structure, or characteristic described in connection with that embodiment is included in one or more embodiments of this application. Thus, statements such as "in one embodiment", "in some embodiments", "in other some embodiments", "in still other embodiments", etc. that appear in different places in this specification do not necessarily all refer to the same embodiment, but mean "one or more but not all embodiments", unless otherwise specifically emphasized in another way. The terms "comprising", "including", "having" and their variants all mean "including but not limited to", unless otherwise specifically emphasized in another way.
[0042] In view of the problem that the current three-dimensional models of buildings cannot meet the high-precision requirements of smart city applications, the embodiments of this application provide a semantic real-scene three-dimensional modeling method for urban buildings. By constructing three-dimensional models of each building in the city based on Internet crowd-sourced data, innovatively obtaining building feature description information and text information of building facades from street view image data, then determining the semantic description information of each building in the city by combining the building feature description information and the text information of building facades, and by assigning the semantic description information to the three-dimensional models of the corresponding buildings, a three-dimensional model of a building with semantic description information can be obtained, so that the three-dimensional model of the building can meet the high-precision requirements of smart city applications.
[0043] The semantic real-scene three-dimensional modeling method for urban buildings provided by this application will be exemplarily described below in conjunction with specific embodiments.
[0044] As Figure 1 shown, the semantic real-scene three-dimensional modeling method for urban buildings provided by the embodiments of this application includes the following steps:
[0045] Step 11, constructing three-dimensional models of each building in the city according to the OSM data of the city.
[0046] In the related art, Open Street Map (OSM) data has been widely used to construct 3D white models of urban buildings (i.e., the 3D models of the above-mentioned buildings). OSM data provides open-source geographical data worldwide, especially detailed building contour data, which provides a basis for quickly generating 3D geometric models of buildings. For example, Ma (2024) et al. constructed 3D models of cities through OSM data. The construction of 3D models based on OSM data has advantages such as low cost and fast response. However, since the main composition of OSM data is two-dimensional building contour information, the directly generated 3D models lack semantic description information. Among them, the construction results of 3D models of each building in the city can be as Figure 2 shown.
[0047] Specifically, the construction of 3D models of buildings based on OSM data is a complex process involving data extraction, preprocessing, geometric modeling, and model optimization. As a global open-source map project, OSM provides rich geographical information data, including basic information such as the location, shape, and height of buildings, which provides a rich data source for the construction of 3D models of buildings. When constructing 3D models of each building in the city, it is first necessary to extract information related to buildings from the OSM data of the city. OSM data is usually stored in XML format, which contains a large number of elements such as tags, nodes, and ways. For the construction of 3D models of buildings, the elements with the "building" tag are mainly concerned, which usually contain information such as the geometric shape, height, and use of buildings.
[0048] The data extraction process is mainly implemented by writing Python (Python is a computer programming language) scripts and using specialized OSM data processing toolkits networkx and geopandas (networkx and geopandas are two different Python libraries, which are used for network analysis and geospatial data processing respectively). The extracted data needs to be preprocessed to ensure its quality and consistency. The preprocessing steps include data cleaning and format conversion. The purpose of data cleaning is to remove duplicate or incorrect data entries to ensure the uniqueness of the identifier (ID) of each building. Format conversion is to convert OSM data into a format suitable for 3D modeling. An example of building data obtained from OSM is shown in Table 1.
[0049] Table 1 Attribute fields of buildings obtained from OSM
[0050]
[0051] The geographical entities in the above table can be buildings. The unit of the number of floors is floors, and the unit of height is meters.
[0052] Building a 3D model of a building involves the following key steps: geometric shape determination, height information application, and visualization. Geometric shape determination is to determine the basic geometric shape of the building based on the geometry attribute in the OSM data. Height information application is to apply the building height information provided by the "building height" tag in the OSM data to the corresponding building geometry and perform stretching to form a three-dimensional building model. The 3D model of a building is determined by the geometry and height of the building and can be defined by a mathematical formula. For example, the 3D model of a rectangular building can be defined by the coordinates of its four corner points:
[0053]
[0054] where is the 3D model of the building, , , , are the coordinates of the four corner points of the building.
[0055] Step 12: Determine multiple street view image sampling positions based on the road network data of the city and collect street view image data in multiple directions at each street view image sampling position.
[0056] In some embodiments of the present application, multiple street view image sampling positions can be set on each road in the road network data. Among them, the distance between two adjacent street view image sampling positions on the same road is a preset distance, and this preset distance can be set according to the actual situation, for example, set to 50 meters.
[0057] After determining each street view image sampling position, for each street view image sampling position, calculate the azimuth angle of the road where the street view image sampling position is located, and crawl from the Internet the street view image data in the four directions of front, back, left, and right at the street view image sampling position to ensure that the collected street view image data can accurately reflect the actual perspective of the road.
[0058] In practical applications, the process of determining the street view image sampling position is as follows: preprocess the road network data of the city (which can be OSM road network data) to generate street view image sampling positions suitable for street view image data collection. First, select the Web Mercator coordinate system to perform projection conversion on the road network data, converting it from the geographic coordinate system to the projection coordinate system. Then, screen out the main road types, such as main roads, secondary roads, and residential roads, while removing irrelevant categories such as footpaths and bicycle lanes, and split multi-component elements, break all intersecting roads, and simplify the roads to reduce unnecessary complexity.
[0059] The generation of street view image sampling locations is accomplished through the Point Along Line tool provided by the Arcpy package in Python (the Arcpy package is a toolkit in Python), with a sampling interval of 50 meters. The longitude and latitude coordinates of the sampling points are added to the attribute table of the points through the Add XY Coordinates tool for subsequent processing. Finally, a CSV file containing street view image sampling location information is obtained.
[0060] Using the CSV file containing street view image sampling location information (which can be X, Y coordinates), street view image data is automatically collected from the Internet using Python web scraping technology. The web scraping program can send requests through the Baidu Street View Application Programming Interface (API) to obtain street view images at specified street view image sampling locations. Street view image data in four directions, namely front, back, left, and right, is collected for each street view image sampling location. Specifically, for each street view image sampling location, the azimuth angle of the road where it is located is calculated, and the horizontal view angle (heading) parameter in the street view API request is adjusted according to this angle to ensure that the collected street view image data can accurately reflect the actual view of the road. If we set as the set of street view image sampling locations, as the set of roads, then the relationship between the street view image sampling location and the road can be expressed as:
[0061]
[0062] where, is the sampling function and interval is the sampling interval. For each street view image sampling location , street view image data in four directions is collected, which can be expressed as:
[0063]
[0064] where, is the set of street view image data in four directions collected at the street view image sampling location , , , and respectively represent the street view image data in the front, back, left, and right directions of the street view image sampling location .
[0065] It should be noted that as a kind of crowdsourced visual information data with wide sources and frequent updates, compared with remote sensing data and aerial images, street view image data has the advantages of low cost, high detail, and real-time update, and is an important data source for obtaining building exterior details and semantic information extraction. Through street view image data, the exterior facade information of buildings can be captured, including door numbers, advertising signs, etc., to supplement the semantic description of the 3D building models generated from OSM data. In addition, street view image data also provides rich environmental context information, which helps to improve the fineness of building models and the accuracy of scene understanding.
[0066] Specifically, street view image data can reflect the relative relationship between buildings and surrounding entities and their spatial arrangement characteristics. In addition to text information, these characteristics include the distribution of buildings, the relationship with the surrounding environment, etc., which can provide important references for the semantic classification of buildings. For example, in residential areas, buildings are usually arranged more compactly and have similar heights, forming a neat street view layout; while around educational buildings such as schools, there may be obvious visual elements such as playgrounds and students carrying schoolbags, which constitute the specific "scene" characteristics of this area, thus helping to judge the functional use of the building.
[0067] It can be understood that street view image data and OSM data, as crowdsourced data, can be obtained through channels such as mobile devices, social networks, and Internet of Things sensors. For the convenience of data processing, after obtaining these data, they are usually processed through formatting, denoising, and normalization to ensure the effectiveness and consistency of the data, providing a solid foundation for the geometric and semantic construction of urban building models.
[0068] Step 13: For each street view image data respectively, obtain the building feature description information and text information set of the street view image data, and determine the building function type corresponding to the street view image data according to the building feature description information, text information set, and predefined building function description dictionary.
[0069] The above-mentioned building feature description information is used to describe the relationship between the building in the street view image data and other entities. Other entities include roads, trees, sidewalks, road signs, and vehicles, etc., and the relationship is a spatial or functional relationship. Exemplarily, the building feature description information can be: there are multiple modern buildings beside the road, the building has many signs, there are multiple cars on the roadside, and the pedestrians are relatively dense.
[0070] In some embodiments of the present application, specifically, a cross-modal model (CLIP, Contrastive Language-Image Pre-Training) can be utilized to analyze and process street view image data to obtain the building feature description information of the street view image data. Specifically, by deeply analyzing the street view image data through the cross-modal model, the semantic information of the building and its surrounding environment can be comprehensively understood, and the building feature description information of the street view image data can be output.
[0071] The above cross-modal model includes an image encoder and a text encoder . Before using the cross-modal model to analyze and process the street view image data, it is necessary to first extract image-text pairs from a large-scale dataset collected from the Internet, denoted as , where represents an image, represents the associated text description. Using the principle of contrastive learning, a dual-encoder model, namely the image encoder and the text encoder , is trained to learn to extract rich semantic features from the data pairs.
[0072] The image encoder and the text encoder are jointly trained by minimizing the contrastive loss function , which measures the distance difference between the positive sample pairs and the negative sample pairs , represents the negative sample text. Specifically, for each image-text pair, the image feature vector and the text feature vector are calculated, and their matching degree is measured by the cosine similarity. The contrastive loss function can be expressed as:
[0073]
[0074] where, represents the cosine similarity function, is the temperature parameter used to control the smoothness of the probability distribution, B is the set of all image-text pairs in the batch, is the negative sample text 's feature vector.
[0075] In the test phase, given a target image , the trained image encoder is used to extract its feature vector . Then, by calculating and all candidate text feature vectors The cosine similarity between them is used to retrieve the most matching text description. Finally, the overall semantic perception of the built environment can be expressed as:
[0076]
[0077] where is the predicted text description most relevant to the most relevant text description.
[0078] The above text information set includes the text information on the building facades in the street view image data, that is, the text information such as billboards, signs, and house numbers in the street view image data. The text information such as billboards, signs, and house numbers usually directly indicates the function or use of the building. The text information may include signs, names, function descriptions, etc. (such as "bank", "school", "supermarket", etc.), providing direct clues about the use and characteristics of the building.
[0079] In some embodiments of the present application, optical character recognition (OCR) technology can be used to perform text detection and recognition on the street view image data to obtain the text information set of the street view image data.
[0080] In practical applications, image processing and natural language processing technologies can be combined to identify the function types of buildings in the street view image data. Use OCR technology to perform text detection and recognition on the street view image data to obtain the text information set , where represents the detected th text string, such as bank, school, etc., , is the number of text strings.
[0081] The above building function description dictionary includes function description information corresponding to various building function types. Among them, the function description information is used to describe the function of the building, and the various building function types include commercial services, corporate enterprises, public services, residential service function types, etc.
[0082] In some embodiments of the present application, the specific implementation method for determining the building function type corresponding to the street view image data according to the building feature description information, the text information set, and the predefined building function description dictionary is as follows:
[0083] Step 13.1, fuse the building feature description information and the text information set through string splicing to obtain the fused text description. The fused text description D' can be expressed as: , is the fusion operation, is the building feature description information, is the text information set.
[0084] Step 13.2, calculate the similarity between the fused text description and each functional description information in the building function description dictionary.
[0085] The above building function description dictionary can be denoted as , indicating the functional description information corresponding to the j-th building function type in the dictionary, , , where N is the number of building function types. In some embodiments of the present application, the similarity between the fused text description and can be calculated by the cosine similarity metric method :
[0086]
[0087] where · represents the dot product of vectors, and ||·|| represents the Euclidean norm of vectors.
[0088] Step 13.3, take the building function type corresponding to the functional description information with the highest similarity to the fused text description as the building function type corresponding to the street view image data.
[0089] Exemplarily, for a certain street view image data, using the text-image multi-modal model to analyze and process it, the obtained building feature description information is: This building is a small company. Using OCR technology to detect and recognize the text, the obtained text information set is: A certain limited company, a certain dealer. Fuse "This building is a small company" and "A certain limited company, a certain dealer", and then calculate the similarity between the fused text description and each functional description information in the building function description dictionary. Finally, determine that the building function type corresponding to the functional description information with the highest similarity is company enterprise, and take company enterprise as the building function type of this street view image data.
[0090] It is worth mentioning that the key text information in the street view image data of the building provides a crucial basis for inferring the building function (such as the text information on billboards, signs, doorplates, etc. in the street view image data usually directly indicates the function or use of the building (such as "bank", "school", "supermarket", etc.)), which can effectively improve the efficiency and accuracy of building function inference. This process first involves in-depth analysis of the street view image data, using the open-source image-to-text CLIP model to generate the overall semantic description of the building. At the same time, text information is extracted from the same street view image through OCR technology. Finally, compare the description and text information with the predefined building type text description template to infer the semantic function of the building.
[0091] Step 14: According to all the determined building function types and multiple street view image sampling positions, assign semantic description information to the 3D models of each building in the city to obtain the semantic real-scene 3D models of each building in the city.
[0092] The above semantic description information is used to describe the attributes of buildings, such as shopping malls, banks, companies, etc. The above semantic real-scene 3D model refers to: on the basis of the traditional 3D geometric model of urban buildings (i.e., the 3D model constructed based on OSM data above), combined with the building attribute information from multi-source data, attach semantic information to the model, so that it can not only represent the geometric structure of the building, but also reflect the attributes such as the function and type of the building. The semantic real-scene 3D model has a wide range of application scenarios in the construction of smart cities, such as traffic management, building energy consumption analysis, disaster emergency response, urban planning and simulation, etc.
[0093] In some embodiments of the present application, in the above step 14, the specific implementation method of assigning semantic description information to the 3D models of each building in the city according to all the determined building function types and multiple street view image sampling positions can be: for each street view image sampling position, take the area centered on the street view image sampling position with a preset buffer distance (which can be set according to the actual situation, for example, set to 50 meters) as the buffer area of the street view image sampling position, and assign the building function types corresponding to all the street view image data of the street view image sampling position as the semantic description information to the 3D models of all the buildings in the buffer area.
[0094] Of course, if there are buildings in the city that are not located in any buffer area, determine the target buffer area closest to the building from all the buffer areas, and assign the building function type corresponding to the target buffer area as the semantic description information to the 3D model of the building.
[0095] Through the method of buffer areas, it is possible to effectively associate the function information of buildings extracted from street view image data with predefined function types, so as to realize the automatic construction of semantic real-scene 3D models. Among them, the results of buffer areas in the city can be as Figure 3 shown. Figure 3 The OSM buildings in it are the 3D models constructed based on OSM data as mentioned above. The results of the semantic real-scene 3D models of each building in the city can be as Figure 4 shown.
[0096] In some embodiments of the present application, by utilizing the real-time acquisition characteristics of street view image data, the model can be dynamically updated according to changes in the appearance of urban buildings. For example, the replacement of shops, changes in the facade information of buildings, or functional conversions can all be detected and recognized through new street view image data, thereby updating the semantic description information in the model. Combining OCR and image recognition algorithms to ensure the geometric update of the model and perform dynamic adjustments at the semantic level.
[0097] Specifically, when the update cycle arrives, the OSM data of the city and the street view image data corresponding to the sampling positions of each street view image can be updated, and the semantic real-world three-dimensional model of each building in the city can be updated based on the updated OSM data and street view image data. That is, when the update cycle arrives, the above steps 11, 13, and 14 are re-executed according to the updated OSM data and street view image data to obtain the updated semantic real-world three-dimensional model.
[0098] It is worth mentioning that the semantic real-world three-dimensional modeling method of the present application constructs three-dimensional models of each building in the city based on Internet crowd-sourced data, innovatively obtains building feature description information and text information on the building facade from street view image data, then determines the semantic description information of each building in the city by combining the building feature description information and the text information on the building facade, and by assigning the semantic description information to the three-dimensional model of the corresponding building, a three-dimensional model of the building with semantic description information can be obtained, thereby enabling the three-dimensional model of the building to meet the requirements for the functional attributes of buildings in smart city applications, making the model more accurately reflect the actual use of the building, and enhancing the usability and reference value of the model.
[0099] The above is the preferred implementation manner of the present application. It should be noted that for those of ordinary skill in the art in the technical field, without departing from the principle described in the present application, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present application.
Claims
1. A semantic real scene 3D modeling method for urban buildings, characterized in that: include: Constructing a three-dimensional model of each building in the city based on the OSM data of the city; Determine a plurality of street view image sampling locations based on the road network data of the city, and collect street view image data in multiple directions at each street view image sampling location; For each of the street view image data, respectively, obtain building feature description information and a text information set of the street view image data, and determine the building function type corresponding to the street view image data according to the building feature description information, the text information set and a predefined building function description dictionary; the building feature description information is used to describe the relationship between the building and other entities in the street view image data, the text information set includes text information of the building facade in the street view image data, and the building function description dictionary includes function description information corresponding to multiple building function types; Assigning semantic description information to the three-dimensional model of each building in the city according to the determined functional types of all buildings and the sampling locations of the plurality of street view images, so as to obtain a semantic real-scene three-dimensional model of each building in the city; the semantic description information is used to describe the attributes of the building; The determining the building function type corresponding to the street view image data according to the building feature description information, the text information set and a predefined building function description dictionary includes: The building feature description information and the text information set are merged by string concatenation to obtain a merged text description; Calculating the similarity between the fused text description and each functional description information in the building functional description dictionary; The building function type corresponding to the function description information having the highest similarity to the fused text description is used as the building function type corresponding to the street view image data.
2. The semantic real scene 3D modeling method according to claim 1, characterized in that: The determining of a plurality of street view image sampling locations based on the road network data of the city comprises: A plurality of street view image sampling positions are set on each road in the road network data; and a distance between two adjacent street view image sampling positions on the same road is a preset distance.
3. The semantic real scene 3D modeling method according to claim 2, characterized in that: The collecting of street view image data in multiple directions at each street view image sampling position includes: For each of the street view image sampling positions, the azimuth of the road where the street view image sampling position is located is calculated, and based on the azimuth, street view image data in four directions of front, back, left, and right taken at the street view image sampling position is obtained from an Internet crawler.
4. The semantic real scene 3D modeling method according to claim 1, characterized in that: The step of obtaining the building feature description information and text information set of the street view image data includes: Analyzing and processing the street view image data using a graphic-text multimodal model to obtain building feature description information of the street view image data; The OCR technology is used to perform text detection and recognition on the street view image data to obtain a text information set of the street view image data.
5. The semantic real scene 3D modeling method according to claim 1, characterized in that: The step of assigning semantic description information to the three-dimensional model of each building in the city according to the determined functional types of all buildings and the sampling locations of the plurality of street view images includes: For each street view image sampling position, an area in the city with the street view image sampling position as the center and a preset buffer distance as the radius is used as a buffer area of the street view image sampling position, and building function types corresponding to all street view image data corresponding to the street view image sampling position are assigned as semantic description information to the three-dimensional models of all buildings in the buffer area.
6. The semantic real scene 3D modeling method according to claim 5, characterized in that: The semantic real scene 3D modeling method also includes: If there is a building in the city that is not located in any buffer area, the target buffer area closest to the building is determined from all buffer areas, and the building function type corresponding to the target buffer area is assigned to the three-dimensional model of the building as semantic description information.
7. The semantic real scene 3D modeling method according to claim 1, characterized in that: The other entities include roads, trees, sidewalks, road signs, and vehicles.
8. The semantic real scene 3D modeling method according to claim 1, characterized in that: The semantic real scene 3D modeling method also includes: When the update cycle arrives, the OSM data of the city and the street view image data corresponding to each street view image sampling position are updated, and the semantic real scene three-dimensional model of each building in the city is updated based on the updated OSM data and street view image data.
Citation Information
Patent Citations
Many source data-driven LOD2-level city building model enhancement modeling algorithm
CN115482355A
Building instance function identification method and device and storage medium
CN118332499A