Method, device and storage medium for comprehensive landscape representation based on street images
By extracting and calculating features from street image sets, the problem of comprehensive urban landscape representation in existing technologies has been solved, achieving high-precision representation of urban microstructure and landscape features, and making up for the shortcomings of remote sensing images under complex weather conditions.
Patent Information
- Application Number
- CN202510142438.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-10
- Publication Date
- 2025-12-16
- Estimated Expiration
- 2045-02-10
AI Technical Summary
Existing technologies struggle to effectively represent complex urban landscapes across multiple cities using a single street image, and remote sensing imagery is limited in application under complex weather conditions, making it difficult to acquire detailed information about the city's interior.
By acquiring a set of street images and spatially connecting them based on geographic location, image features are extracted using a pre-trained scene feature extractor. The mean and standard deviation of the image embedding features are calculated and stitched together to form a comprehensive landscape representation.
It enables an intuitive and comprehensive representation of the microstructure and landscape features within the city, breaking through the limitations of single images and improving the accuracy and precision of urban landscape representation.
Smart Images

Figure CN119597956B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of information technology and urban geographic technology, and in particular to a comprehensive landscape representation method based on street images, a device and a storage medium. BACKGROUND
[0002] Urban landscape research is an important part of urban planning, environmental protection, and smart city construction. With the acceleration of global urbanization, the study of urban landscape not only helps to understand and analyze the current development of the city, but also provides a scientific basis for future urban planning and policy making. Through the study of urban landscape, information about the spatial structure, land use patterns, ecological environment conditions, and social and economic activities of the city can be revealed. These research results can help city managers develop more reasonable and sustainable development strategies and improve the quality of life of urban residents. The main technical means for studying urban landscape is remote sensing image technology. Remote sensing image technology can capture the overall view of the city quickly and widely by satellite or aircraft. This method can obtain high-resolution image data for accurate spatial analysis and change monitoring. Using remote sensing image technology, researchers can analyze information about urban land use changes, green space distribution, and building density. However, remote sensing image technology also has certain limitations, such as difficulty in obtaining detailed information within the city and limited application in complex weather conditions.
[0003] In addition to remote sensing image technology, existing technologies also use street images to study urban scenes. Street images are images taken on the ground by street view cars or personal devices, which can provide detailed visual information about the city, such as building facades, street facilities, and vegetation conditions. These image data can help researchers better understand the microstructure and landscape characteristics of the city. Existing technologies still use single street view images to represent a single scene, but urban landscapes are often a collection of multiple different urban scenes, and the number of urban scenes varies greatly in different spatial scales and spatial units. How to comprehensively represent urban landscapes based on complex multiple urban scenes is still a difficult problem to solve. SUMMARY
[0004] To make up for the above-mentioned defects, the present application proposes a comprehensive landscape representation method based on street images, a device and a storage medium.
[0005] The method technical solution is: a comprehensive landscape representation method based on street images, comprising:
[0006] obtaining a set of street images of a landscape representation area, each image in the set of street images containing the geographical location of the image taken by the image;
[0007] acquire a vector geospatial range of the to-be-landscape-represented region;
[0008] spatially connect the vector geospatial range and the street image set based on geographical positions to acquire a spatial mapping relationship between the vector geospatial range and the street image set;
[0009] extract a street image set located within the vector geospatial range from the street image set according to the spatial mapping relationship;
[0010] extract features of each image in the street image set using a pre-trained scene feature extractor to obtain image embedding features of each image, the image embedding features being feature representations of scenes of the images;
[0011] calculate a mean value and a standard deviation of the image embedding features of all images in the street image set to obtain a mean value embedding feature and a standard deviation embedding feature respectively, and splice the mean value embedding feature and the standard deviation embedding feature to obtain spliced features;
[0012] use the spliced features as a comprehensive landscape representation of the to-be-landscape-represented region.
[0013] Preferably, the process of extracting a street image set located within the vector geospatial range from the street image set according to the spatial mapping relationship further comprises dividing the vector geospatial range into a plurality of spatial units according to geographical features, extracting a sub-street image set corresponding to each spatial unit from the street image set according to the spatial mapping relationship, and constructing the street image set from the plurality of sub-street image sets.
[0014] Preferably, the process of extracting features of each image in the street image set using a pre-trained scene feature extractor to obtain image embedding features of each image comprises extracting features of each image in each sub-street image set using a pre-trained scene feature extractor to obtain image embedding features of each image in the corresponding sub-street image set, and using the corresponding image embedding features as feature representations of scenes of images in the corresponding spatial unit of the sub-street image set.
[0015] Preferably, the process of calculating a mean value and a standard deviation of the image embedding features of all images in the street image set to obtain a mean value embedding feature and a standard deviation embedding feature respectively, and splicing the mean value embedding feature and the standard deviation embedding feature to obtain spliced features comprises calculating a mean value and a standard deviation of image embedding features in each sub-street image set based on the corresponding sub-street image set of each spatial unit to obtain a mean value embedding feature and a standard deviation embedding feature of the sub-street image set, and splicing the mean value embedding feature and the standard deviation embedding feature to obtain spliced features of the corresponding sub-street image set.
[0016] Preferably, the comprehensive landscape representation process of the to-be-landscape-represented region with the spliced features comprises: taking the spliced features of the sub-street view image set as the comprehensive landscape representation of the corresponding space unit, and taking the comprehensive landscape representations of all the space units as the comprehensive landscape representation of the to-be-landscape-represented region.
[0017] Preferably, before feature extraction is performed on each image in the street view image set using the pre-trained scene feature extractor, each image in the street view image set is also preprocessed to unify the pixels and sizes.
[0018] Preferably, the preprocessing process comprises unifying the pixel sizes of each image.
[0019] The present application also proposes an electronic device, which comprises a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor, and the processor implements the method according to any one of the above when executing the computer program.
[0020] The present application also proposes a computer readable storage medium, which comprises a stored computer program, wherein the computer readable storage medium controls the device where the computer readable storage medium is located to execute the method according to any one of the above when the computer program runs.
[0021] The present application can more intuitively reflect the microstructure and landscape features of the city interior by using street view images. The present application breaks through the limitation of the prior art that only relies on a single street image for landscape representation, and can more comprehensively and accurately represent the landscape features of the city by comprehensively processing multiple street images. The present application uses a pre-trained deep learning model to extract features from street images, can accurately capture important features in each image, and realizes high-precision landscape representation of space units by calculating the average value and standard deviation of the features. BRIEF DESCRIPTION OF DRAWINGS
[0022] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the drawings needed to be used in the embodiments or prior art description will be briefly introduced. Obviously, the drawings in the following description only constitute some embodiments of the present application, and for those skilled in the art, other drawings can also be obtained without creative labor.
[0023] Figure 1 Flowchart of an embodiment of the method of the present application. DETAILED DESCRIPTION
[0024] The technical solutions of the present application will be clearly and completely described below with reference to the drawings. Obviously, the described embodiments are only some of the embodiments of the present application, but not all of them. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative work fall within the scope of protection of the present application.
[0025] In the description of the present application, it should be noted that the terms "center", "upper", "lower", "left", "right", "vertical", "horizontal", "inner", "outer" and the like indicate the orientation or positional relationship shown in the drawings, and are only for the convenience of describing the present application and simplifying the description, and do not indicate or imply that the devices or elements referred to must have a particular orientation, be constructed and operated in a particular orientation, and therefore cannot be understood as limiting the present application. In addition, the terms "first", "second", "third" are only for descriptive purposes and cannot be understood as indicating or implying relative importance.
[0026] In the description of the present application, it should be noted that unless otherwise explicitly specified and limited, the terms "mounting", "connection", "connection" should be understood in a broad sense, for example, it can be fixedly connected, or it can be detachably connected, or integrally connected; it can be mechanically connected, or it can be electrically connected; it can be directly connected, or it can be indirectly connected through an intermediate medium; it can be the communication inside two elements. For those of ordinary skill in the art, the specific meaning of the above terms in the present application can be understood according to the specific circumstances
[0027] The present application will be further described in detail below with reference to the drawings.
[0028] In combination Figure 1 , an embodiment of the present application. A comprehensive landscape representation method based on street images, comprising:
[0029] obtaining a set of street images of a region to be represented, each image in the set of street images containing the geographical position of the image; obtaining a vector geospatial range of the region to be represented; connecting the vector geospatial range and the set of street images based on the geographical position to obtain a spatial mapping relationship between the vector geospatial range and the set of street images; filtering and extracting a set of street images located within the vector geospatial range from the set of street images according to the spatial mapping relationship; using a pre-trained scene feature extractor to extract features from each image in the set of street images, respectively, to obtain image embedding features of each image, the image embedding features being a feature representation of the scene of the image; calculating the average and standard deviation of all image embedding features in the set of street images to obtain average embedding features and standard deviation embedding features, respectively; splicing the average embedding features and the standard deviation embedding features to obtain spliced features; and taking the spliced features as the comprehensive landscape representation of the region to be represented.
[0030] Further, based on the above-mentioned embodiments, the process of filtering and extracting the street view image set located in the vector geospatial range from the street view image set according to the spatial mapping relationship further comprises: dividing the vector geospatial range into a plurality of spatial units according to geographical features, filtering and extracting a sub-street view image set corresponding to each spatial unit from the street view image set according to the spatial mapping relationship, and the plurality of sub-street view image sets constitute the street view image set.
[0031] Further, based on one or more of the above-mentioned embodiments, the process of using the pre-trained scene feature extractor to extract features from each image in the street view image set to obtain image embedding features of each image respectively comprises: using the pre-trained scene feature extractor to extract features from each image in each sub-street view image set to obtain image embedding features of each image in the corresponding sub-street view image set, and taking the corresponding image embedding features as the feature representation of the image scene of the corresponding spatial unit of the sub-street view image set.
[0032] Further, based on one or more of the above-mentioned embodiments, the process of calculating the mean and standard deviation of all image embedding features in the street view image set to obtain mean embedding features and standard deviation embedding features respectively, and splicing the mean embedding features and the standard deviation embedding features to obtain spliced features comprises: based on the sub-street view image set corresponding to each spatial unit, calculating the mean and standard deviation of the image embedding features in the sub-street view image set to obtain the mean embedding features and the standard deviation embedding features of the sub-street view image set, and splicing the mean embedding features and the standard deviation embedding features to obtain the spliced features of the corresponding sub-street view image set.
[0033] Further, based on one or more of the above-mentioned embodiments, the process of taking the spliced features as the comprehensive landscape representation of the to-be-landscape-represented region comprises: taking the spliced features of the sub-street view image set as the comprehensive landscape representation of the corresponding spatial unit, and taking the comprehensive landscape representations of all spatial units as the comprehensive landscape representation of the to-be-landscape-represented region.
[0034] On the basis of one or more of the above embodiments, further, the present application further takes the city as the research area, and illustrates the comprehensive landscape of the street scene. In this embodiment, 36 cities in China with different economic development levels and geographical locations are selected as the research area, including obtaining the street view image set of Tencent Map of 36 cities in China, which contains a total of 8189058 street view images, wherein each image contains the latitude and longitude information of the shooting geographical location information; obtaining the administrative district vector spatial data of the 36 cities, which is the vector geographical spatial range data, in the rectangular coordinate system, the data using each coordinate axis data to represent the position and shape of the map graph or geographical entity, in this embodiment, the vector geographical spatial range data of the 36 cities is obtained according to the administrative district vector spatial data of the 36 cities by software input form. Subsequently, according to the vector geographical spatial range and the street view image set based on the geographical location, the spatial connection is carried out, that is, according to the latitude and longitude information of the 36 cities and the latitude and longitude geographical information of the shooting street view image set, the spatial connection is carried out, so that they are mapped together according to the geographical location. In the mapping process, the checking process is also included, that is, whether the latitude and longitude information of the 36 cities and the latitude and longitude geographical information of the shooting street view image set exist corresponding relationship, if a city or a region of a city has no corresponding image, the region is deleted, and the region with mapping relationship is reserved. Then the street view images with spatial mapping relationship are used for feature extraction by using the pre-trained scene feature extractor. In this example, the ResNet-18 model based on Places365 dataset is used for feature extraction, and the image embedding feature of each image is 512 dimensions, which is the feature representation of each scene. Then, the average value and the standard deviation of the image embedding features are calculated, and the average value embedding feature and the standard deviation embedding feature of 512 dimensions are obtained. The average value embedding feature and the standard deviation embedding feature are spliced to obtain the embedding feature of 1024 dimensions. The deep feature is the comprehensive landscape feature representation. In this embodiment, in order to improve the operation speed, the 36 cities can be further screened and expressed by dividing the spatial unit corresponding to the sub-street view image set, such as based on the kilometer grid as the basic spatial unit, the administrative district vector spatial data of each city is divided into multiple kilometer grids, and each city is divided into a spatial unit. Similarly, according to the spatial mapping relationship, the sub-street view image set corresponding to the spatial unit of the city is extracted from the street view image set, and the images of the sub-street view image set are further mapped to each basic spatial unit. If there is no street view image in the kilometer grid, the kilometer grid is not included in the following analysis process. In this way, each kilometer grid corresponds to several street view images, and the whole area to be expressed is divided into separate kilometer grids for expression by this way, which reduces the operation amount and improves the operation efficiency.Similarly, in the process of calculating the comprehensive landscape feature representation, the basic spatial unit is still used, and according to the street image corresponding to each kilometer grid, the average value and the standard deviation of the image embedding features are calculated to obtain the average value embedding features and the standard deviation embedding features of each 512 dimensions respectively. The average value embedding features and the standard deviation embedding features are spliced to obtain the embedding features of 1024 dimensions. The deep features are the comprehensive landscape feature representation of the kilometer grid. The embodiment aims at the defects of the existing remote sensing image technology that mainly provides a bird's eye view and is difficult to obtain detailed information inside the city. By using street images from the perspective of pedestrians, the microstructure and landscape features inside the city, such as building facades, street facilities and vegetation conditions, can be more intuitively reflected. The application breaks through the limitation of the prior art that only relies on a single street image for landscape representation, and through the comprehensive processing of multiple street images, the city landscape features can be more comprehensively and accurately represented. The embodiment uses a pre-trained deep learning model to extract features from street images, which can accurately capture important features in each image, and through calculating the average value and the standard deviation of these features, high-precision landscape representation of the spatial unit is realized.
[0035] On the basis of one or more of the above embodiments, further, before each image in the street image set is subjected to feature extraction using the pre-trained scene feature extractor, each image in the street image set is also subjected to row preprocessing of uniform pixels and sizes.
[0036] On the basis of one or more of the above embodiments, further, the preprocessing process includes uniformizing the pixel size of each image. In the process of the embodiment, since the images used in the application can come from multiple shooting devices, there is a pixel gap between the shooting devices. In order to facilitate the neural network to extract the features of the images, the pixels of the original images of the application are uniformized.
[0037] An electronic device includes a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor, and the processor implements the method according to any one of the above embodiments when executing the computer program.
[0038] A computer readable storage medium includes a stored computer program, wherein the computer readable storage medium controls the device where the computer readable storage medium is located to execute the method according to any one of the above embodiments when the computer program runs.
[0039] The landscape representation method proposed in the application is described in detail above. The principles and implementation modes of the present scheme are described in the above examples, but only to help understand the method and core idea of the present application. At the same time, according to different application scenarios, there will be changes in the specific implementation mode. In view of the above, the above example content should not be understood as a limitation of the present scheme.
Claims
1. A comprehensive landscape representation method based on street images, characterized in that, include: Obtain a set of street images of the area to be represented, wherein the set of street images is a street view image from the perspective of a pedestrian, and each image in the set contains the geographical location where the image was taken; Obtain the vector geospatial extent of the area to be represented by the landscape; By spatially connecting the vector geospatial extent with the street image set based on geographical location, the spatial mapping relationship between the vector geospatial extent and the street image set can be obtained; The spatial connection process includes spatially connecting latitude and longitude information with the latitude and longitude geographic information of the captured street image set. If there is no corresponding relationship between the latitude and longitude information and the latitude and longitude geographic information of the captured street image set, the area is deleted, and the areas with a mapping relationship are retained. Based on the spatial mapping relationship, a set of street view images located within the vector geospatial range is extracted from the street image set. The specific process includes dividing the vector geospatial range into multiple spatial units according to geographical features. The spatial unit division process includes dividing the administrative region vector spatial data of each city into multiple kilometer grids based on kilometer grids as the basic spatial unit, and converting all the kilometer grids of each city into a spatial unit. Based on the spatial mapping relationship, sub-street view image sets corresponding to each spatial unit are extracted from the street view image set. The images of the sub-street view image sets are mapped together with each basic spatial unit. Multiple sub-street view image sets constitute a street view image set. Each image in the street view image set is processed using a pre-trained scene feature extractor to extract features, resulting in image embedding features for each image. These image embedding features represent the scene features of the image. Calculate the mean and standard deviation of the embedding features of all images in the street view image set to obtain the mean embedding feature and the standard deviation embedding feature respectively. Then, concatenate the mean embedding feature and the standard deviation embedding feature to obtain the concatenated feature. The splicing features are used to represent the comprehensive landscape of the area to be represented.
2. The method according to claim 1, characterized in that, The process of extracting features from each image in the street view image set using a pre-trained scene feature extractor to obtain the image embedding features of each image includes: extracting features from each image in each sub-street image set using a pre-trained scene feature extractor to obtain the image embedding features of each image in the corresponding sub-street image set, and using the corresponding image embedding features as the feature representation of the scene of the corresponding spatial unit image in the sub-street image set.
3. The method according to claim 2, characterized in that, The process of calculating the mean and standard deviation of the image embedding features of all images in the street view image set, obtaining the mean embedding feature and the standard deviation embedding feature respectively, and then concatenating the mean embedding feature and the standard deviation embedding feature, includes: calculating the mean and standard deviation of the image embedding features in the sub-street view image set based on the sub-street view image set corresponding to each spatial unit, obtaining the mean embedding feature and the standard deviation embedding feature of the sub-street view image set, and concatenating the mean embedding feature and the standard deviation embedding feature to obtain the concatenated feature of the corresponding sub-street view image set.
4. The method according to claim 3, characterized in that, The comprehensive landscape representation process for the area to be represented by the splicing features includes using the splicing features of the sub-street view image set as the comprehensive landscape representation of the corresponding spatial unit, and using the comprehensive landscape representation of all spatial units as the comprehensive landscape representation of the area to be represented.
5. The method according to claim 1 or 2, characterized in that, Before using a pre-trained scene feature extractor to extract features from each image in the street view image set, each image in the street view image set is preprocessed to have uniform pixels and size.
6. The method according to claim 5, characterized in that, The preprocessing process includes standardizing the size of each image pixel.
7. An electronic device, characterized in that, The method includes a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor, wherein the processor, when executing the computer program, implements the method as claimed in any one of claims 1 to 6.
8. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes a stored computer program, wherein, when the computer program is executed, it controls the device on which the computer-readable storage medium is located to perform the method as described in any one of claims 1 to 6.
Citation Information
Patent Citations
Reverse geocoding method and device and electronic equipment
CN115525642A
Hyperspectral image feature extraction method based on multilevel variational automatic encoder
CN116416441A
Method and apparatus for encoding geographic location region as well as method and apparatus for establishing encoding model
US20240177469A1