Method for quantitatively analyzing spatial elements of urban high-intensity district based on image data

By extracting shallow and deep features of street scene panoramas, building and optimizing topology maps, combining with deep learning classification models, the problem of insufficient classification accuracy and consistency in multi-factor mixed scenarios is solved, and more efficient and accurate spatial element classification and distribution analysis is achieved.

CN120107685AActive Publication Date: 2025-06-06BEIJING UNIV OF CIVIL ENG & ARCHITECTURE
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510234418.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-28
Publication Date
2025-06-06
Estimated Expiration
2045-02-28

AI Technical Summary

Technical Problem

In the scene of multi-factor mixed, it is difficult for the prior art to ensure the classification accuracy and consistency of classification results.

Method used

By extracting shallow feature and deep feature extraction of street view panoramas, the topology map of spatial elements is constructed in combination with topology analysis, and the structure of the topology map is optimized through preset logical rules. Finally, the topology map is input into the pre-trained deep learning classification model to output the classification results and spatial distribution map of spatial elements.

Benefits of technology

It improves the reliability and accuracy of classification results, enhances the interpretability and consistency of classification results, and is suitable for the analysis of spatial structures in complex and diverse urban high-intensity areas.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120107685A_ABST
    Figure CN120107685A_ABST
Patent Text Reader

Abstract

The invention discloses a quantitative analysis method for spatial elements of an urban high-intensity district based on image data, and relates to the technical field of image element identification, and the method comprises the steps: obtaining a street view panorama from an observation point of a selected urban high-intensity district; on the basis of a visual detection technology, obtaining shallow layer features of the streetscape panorama, and on the basis of a deep learning model, obtaining deep layer features of the streetscape panorama; on the basis of topology analysis, registering the shallow layer features and the deep layer features with a set logic rule to construct a topological graph; and inputting the topological graph into a pre-trained classification model, and finally outputting a classification result and a spatial distribution graph of each space element in the street scene image. According to the method, the shallow features and the deep features of the image are utilized at the same time, multi-dimensional feature fusion is achieved, the classification accuracy is improved, and the adaptability to multi-class elements in a complex scene is enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image element recognition, and in particular to a method for quantitatively analyzing spatial elements of urban high-intensity areas based on image data. Background Art

[0002] With the rapid development of computer technology, the classification of spatial elements in urban area images has gradually shifted from manual interpretation to automated processing. At present, automated classification methods based on image processing technology have gradually become more diverse, mainly including supervised classification, unsupervised classification and object-oriented classification.

[0003] Among them, supervised classification relies on high-quality pre-labeled samples, and guides the classification model to classify the features of unknown areas by learning from known samples. However, the results of supervised classification are highly dependent on the representativeness and accuracy of the samples. When the number of samples is insufficient or the labeling quality is poor, the classification accuracy may drop significantly, and classification bias is prone to occur. Unsupervised classification automatically clusters image pixels through algorithms without pre-labeling samples. However, unsupervised classification has high computing resource requirements, which limits its application efficiency in image processing in large-scale urban areas.

[0004] Object-oriented classification is a method based on image segmentation, which divides the image into a series of small homogeneous regions (i.e., "objects") and then classifies these objects. This method can effectively combine spectral information with geometric information, but the classification results are easily affected by the quality of the image segmentation algorithm. Over-segmentation or under-segmentation will directly affect the final classification accuracy, causing the classification results to deviate from the actual situation.

[0005] In contrast, deep learning-based classification methods can deeply mine deep features in images and show significant advantages in urban image classification. Through multi-scale feature extraction and context information fusion technology, deep learning models can effectively distinguish between spatial elements such as buildings, vegetation, and roads in complex scenes. However, deep learning methods mainly focus on deep mining of the characteristics of spatial elements themselves, and make insufficient use of the knowledge of associations between elements and specific identification features. In multi-factor mixed scenes, it is easy to lead to poor classification accuracy and inconsistent classification results. Summary of the invention

[0006] 1) Technical issues solved The present invention provides a method for quantitatively analyzing spatial elements of urban high-intensity areas based on image data, which can solve the problems of poor classification accuracy and inconsistency of previous and subsequent classification results in scenes with mixed multiple elements.

[0007] 2) Technical solution To achieve the above object, the present invention provides the following technical solution: a method for quantitatively analyzing spatial elements of urban high-intensity areas based on image data, comprising the following steps: Obtain street view panoramas from selected observation points in high-intensity urban areas; Based on edge detection, color analysis and texture analysis, shallow feature extraction is performed on the street view panorama, and spatial elements in the street view panorama are preliminarily identified in combination with predefined shallow feature templates; For each of the spatial elements initially identified, the deep features thereof are obtained through a deep learning model; wherein the deep features include the geometric shape and spatial position of each spatial element, and the semantic association relationship between different spatial elements; Based on topological analysis, a topological graph between spatial elements is constructed, wherein each node of the topological graph represents each spatial element that is initially identified, and each of the nodes carries its shallow features and deep features, and the edges of the nodes represent the semantic association relationship between the spatial elements, and the structure of the topological graph is optimized through preset logical rules; The constructed topological map is input into a pre-trained classification model, and the classification model outputs the classification result of each spatial element in the street view panorama and its spatial distribution map according to the node characteristics and the topological structure.

[0008] Furthermore, before shallow feature extraction is performed on the acquired street view panorama, image preprocessing is performed on the street view panorama, and the image preprocessing includes denoising, image enhancement and image segmentation operations, corresponding to: removing random noise in the street view panorama and retaining edge information; The brightness and contrast of the street view panorama after noise removal are analyzed, histogram equalization is performed on the local area of ​​the image, and contrast enhancement is performed on the low-contrast area.

[0009] Furthermore, edge detection, color analysis and texture analysis are performed on the street view panorama to extract the shallow features, specifically: Perform Canny edge detection on the acquired street view panorama, calculate the intensity of grayscale changes in the street view panorama, and use gradient information to locate the boundaries of objects in the street view panorama; Extract the color distribution characteristics of each object located by HSV color space; The gray level co-occurrence matrix is ​​used to extract the surface texture information of each located object.

[0010] Furthermore, each of the objects located in the street view panorama and the shallow features extracted therefrom are matched with the preset shallow feature template, and the spatial elements represented by each object are preliminarily identified through a matching algorithm.

[0011] Furthermore, for each of the spatial elements that are initially identified, the corresponding deep features are extracted through a deep learning semantic segmentation model, including the geometric shape and spatial position of each spatial element, the relative distance between different spatial elements, and the contextual semantic association relationship between different spatial elements.

[0012] Furthermore, the preset logic rules include: The weight value of the edge is adjusted based on the geometric shape and spatial position of the spatial element preliminarily identified, and the relative distance between different spatial elements; wherein the projection distance between different spatial elements is calculated, a distance threshold is set, and the weight is adjusted proportionally; The weight value of the edge is adjusted according to the semantic association relationship between different spatial elements; if two spatial elements have a functional association in the real scene, the weight is increased; if the two spatial elements have no functional association, the edge of the node represented by the two spatial elements is deleted.

[0013] Furthermore, by optimizing the preset logic rules, the weights of the edges are updated, and a graph convolution operation is performed on the topological graph, thereby optimizing the node features and spatial distribution in the topological graph.

[0014] Furthermore, the pre-trained classification model classifies the spatial elements of the high-intensity area based on deep learning technology. During the training process of the classification model, the topological map is input and the corresponding feature vector of each node is generated according to the shallow features and deep features of each node. The classification model aggregates the features of the target node's neighbor nodes through the message passing mechanism according to the adjacency relationship and edge weight of the target node to be identified, and fuses them with the current features of the target node to generate the target node features containing context information; The feature vector of each node is input into the defined classifier to output the classification result of each node. The classification result of each node is combined with the spatial position relationship in the topological map to generate a spatial distribution map of the spatial elements in the street view panorama.

[0015] III) Beneficial effects: Compared with the prior art, the invention has the following beneficial effects: The present invention preliminarily determines the spatial elements in the image by jointly extracting shallow features and deep features from the street view panorama, constructs a topological map of the spatial elements, and adjusts the edge weights of the topological map through preset logical rules. This enhances the model's adaptability to the complex correlations between spatial elements in actual street scenes and is suitable for complex and diverse spatial structures in high-intensity areas of cities.

[0016] Through the pre-trained deep learning classification model, the characteristics of spatial elements are combined with their topological relationships, and the message passing mechanism of adjacent node features and weights is used to optimize the node feature expression, so that the classification model can effectively capture contextual information, thereby improving the reliability and accuracy of the classification results. Finally, the classification results are combined with the spatial distribution map to intuitively display the distribution of each spatial element. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] Figure 1 A schematic diagram of a flow chart of a method for quantitatively analyzing spatial elements of high-intensity urban areas based on image data provided by an embodiment of the present invention; Figure 2 A schematic diagram of multiple target identification spatial elements of a high-intensity area in a method for quantitative analysis of spatial elements of a high-intensity area in an urban area based on image data provided by an embodiment of the present invention; Figure 3 A schematic diagram of a topological map of multiple spatial elements in an image constructed using topological technology in a method for quantitatively analyzing spatial elements in high-intensity urban areas based on image data provided by an embodiment of the present invention; Figure 4 A schematic diagram of a topological map optimized based on preset logical rules in a quantitative analysis method of spatial elements of high-intensity urban areas based on image data provided in an embodiment of the present invention. DETAILED DESCRIPTION

[0018] The technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention.

[0019] In the description of the present invention, it should be understood that the terms "longitudinal", "lateral", "length", "width", "thickness", "up", "down", "front", "back", "left", "right", "vertical", "horizontal", "top", "bottom", "inside" and "outside" etc., indicating orientations or positional relationships, are based on the orientations or positional relationships shown in the accompanying drawings, and are only for the convenience of describing the present invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be understood as a limitation on the present invention.

[0020] In addition, the terms “first”, “second”, etc., if used, are merely used to distinguish between the descriptions and should not be understood as indicating or implying relative importance.

[0021] It should be noted that, in the absence of conflict, the features in the embodiments of the present invention may be combined with each other.

[0022] With the rapid development of computer technology, the classification of spatial elements in urban area images has gradually shifted from manual interpretation to automated processing. The methods of automated classification using image processing technology are becoming increasingly diverse, mainly including supervised classification, unsupervised classification, and object-oriented classification. These methods have their own characteristics in practical applications.

[0023] Among them, supervised classification is one of the most widely used methods, which relies on high-quality pre-labeled samples, that is, by learning samples of known spatial elements, to guide the classification model to classify the elements of unknown areas. However, the results of supervised classification are highly dependent on the representativeness and accuracy of the samples, and classification bias is prone to occur when the sample quality is inaccurate or the number of samples is insufficient.

[0024] Unsupervised classification automatically clusters image pixels through algorithms without pre-labeling samples. It is a relatively independent classification method. However, unsupervised classification algorithms require the number of categories to be set in advance and rely heavily on the clustering quality of the results. In addition, due to the complex calculation process, slow convergence speed, and long computer operation time of unsupervised classification, its application efficiency in large-scale urban area image processing is limited.

[0025] In addition to the above two classification methods, object-oriented classification is a commonly used method based on image segmentation, which divides the image into a series of small homogeneous areas (i.e., "objects") and then classifies these objects. This method can effectively combine the spectral information and geometric information of the image, but the classification results are easily affected by the quality of the image segmentation algorithm. Excessive or insufficient image segmentation will directly affect the final classification accuracy, making the classification results deviate from the actual situation. Therefore, the object-oriented classification method is also not suitable for large-scale urban area image processing.

[0026] With the rapid development of artificial intelligence, traditional image processing techniques are gradually being replaced by machine learning methods in the classification of spatial elements in images of urban areas.

[0027] Among them, the methods based on traditional machine learning mainly use algorithms such as decision trees, random forests or maximum likelihood methods. These methods are mainly based on the shallow features of the image, that is, they use the direct observation features such as color, texture, shape, etc. of the region in the image to complete the classification. However, such methods have limitations in extracting and utilizing deep features in the image, resulting in insufficient adaptability of the classification results to complex scenes, especially when there are fuzzy boundaries or complex relationships between spatial elements.

[0028] In contrast, deep learning-based methods can fully exploit the deep features in images and have been widely used in urban image classification in recent years. By constructing a multi-layer network structure, deep learning can analyze the high-order semantic information of images and significantly improve classification accuracy. Among them, network structures such as U-Net, Pyramid Scene Parsing Network (PSP Net), and DeepLabv3+ are widely used in spatial feature classification tasks. Through multi-scale feature extraction, context information fusion and other technologies, they can effectively distinguish between elements such as buildings, vegetation, and roads in complex scenes. However, deep learning methods mainly focus on the deep mining of the features of the elements themselves, and lack the use of the knowledge of the association between specific elements and their specific identification features, which may lead to insufficient interpretability and consistency of classification results in scenes with mixed multiple elements.

[0029] In summary, the evolution of spatial element classification technology in urban area images from traditional image processing to artificial intelligence methods has greatly improved the automation and accuracy of classification. However, various methods still have different limitations and challenges. How to further improve classification efficiency, accuracy and applicability is still an important direction of current research.

[0030] In order to solve the above problems, the inventors proposed a method combining Figures 1 to 4 The method shown in the figure is a quantitative analysis method of spatial elements of urban high-intensity areas based on image data. This method not only improves the accuracy and efficiency of quantitative analysis of spatial elements of urban high-intensity areas, but also has strong applicability and scalability, providing more intelligent technical support for urban planning and management.

[0031] Specifically, first, perform S1: obtain a street view panorama from an observation point in a selected high-intensity area of ​​the city.

[0032] According to the characteristics of high-intensity urban areas, such as high building density, large population density, and large traffic flow, the observation points are selected based on the following factors: Geographical location distribution: Priority will be given to representative areas such as commercial areas, transportation hubs, and densely populated areas.

[0033] Scene coverage: Observation points need to be evenly distributed to ensure that the street view panorama covers the main roads, squares and key facilities in the area.

[0034] Functional requirements: Based on specific analysis needs, select areas that can fully reflect various spatial elements such as buildings, roads, plants, and street lights.

[0035] For example, in a city's commercial center, subway station entrances and exits, main road intersections, and commercial plazas are selected as observation points to ensure that crowded areas and key traffic nodes are covered, and to obtain street view panoramic views of the corresponding observation points.

[0036] Before extracting features from the acquired street view panorama, image preprocessing is required.

[0037] Specifically, random noise (such as Gaussian noise and salt and pepper noise) in the street view panorama is removed to avoid affecting the accuracy of subsequent feature extraction while retaining edge information. In some embodiments of the present invention, noise is removed based on a weighted average method by calculating local image similarity while retaining edge and texture information.

[0038] For example, for the street view image of a commercial center mentioned above, the non-local mean filter is applied, and the clarity of the edge area is improved after noise removal, thereby retaining the main details of the building outline.

[0039] The image enhancement operation is to improve the brightness, contrast and visual clarity of the image, improve the visibility of low-light or local low-contrast areas, and provide higher quality input for subsequent feature extraction. In some embodiments of the present invention, global histogram equalization is performed on the entire image to redistribute the grayscale and improve the image contrast, especially the details of the highlights and shadows. Alternatively, adaptive histogram equalization is used to process local areas of the image to avoid over-enhancement that may be caused by global equalization. For areas affected by shadows in street view images, gamma correction is used to adjust the brightness. Laplace filtering or high-pass filtering is performed on the enhanced image to highlight the edge features in the image and improve the clarity of structures such as building outlines and road boundaries.

[0040] For example, after CLAHE processing of a street view panorama along a road, the street lights and road signs that were originally unclear in the shadows become visible, and the contrast is improved, providing more usable information for the classification model input.

[0041] In some embodiments of the present invention, the street view panorama after noise filtering and enhancement processing is divided into multiple small blocks to ensure that each small block contains local information, which is convenient for subsequent feature extraction and classification analysis of specific spatial elements. Among them, the street view panorama is cut into multiple small blocks according to a fixed size (such as 128×128 pixels) using a grid division method, and each small block is used as a basic unit for feature extraction. Alternatively, the K-means clustering algorithm is applied to the street view image to cluster different areas based on pixel values ​​(color, brightness) and spatial distribution to ensure that the pixels in each area are similar and have specific local features.

[0042] In summary, it can be understood that through the above preprocessing and segmentation operations, the processed street view panorama has high quality and clarity, and the local segmentation blocks can accurately represent the spatial elements in the image, providing reliable input data for subsequent feature extraction and classification models.

[0043] After preprocessing the street view panorama, S2 is performed: shallow feature extraction is performed on the street view panorama based on edge detection, color analysis and texture analysis, and the spatial elements in the street view panorama are preliminarily identified in combination with predefined shallow feature templates. And S3: for each preliminarily identified spatial element, its deep features are obtained through a deep learning model; wherein the deep features include the geometric shape and spatial position of each spatial element, and the semantic association relationship between different spatial elements.

[0044] First, edge detection, color analysis, and texture analysis are performed on the street view panorama to extract shallow features. The shallow features mainly extract basic visual information related to objects from the pixel level of the street view panorama, including edges, colors, textures, etc. These shallow features can reflect the contours and local structural information of objects in the image, and are an important basis for subsequent spatial element classification.

[0045] Specifically, by detecting the sudden change area of ​​gray value in the image, the contour information reflecting the boundary and shape of the object is extracted. In some embodiments of the present invention, the Canny algorithm is used to find the area with the largest gray value change by calculating the gradient amplitude and direction of each pixel, and non-maximum suppression is applied to retain the local extreme value of the edge, and double threshold segmentation (high threshold and low threshold) is used to connect the edges, thereby extracting the edges of the object elements in the image.

[0046] Taking the above example, after applying Canny edge detection to the street view image of the commercial street, the building outlines, street light poles, traffic light poles and road boundaries were successfully extracted. The edge detection results are visually clear and less affected by noise.

[0047] By extracting color distribution features, the color information of objects in the street view panorama is described to provide support for classification. In some embodiments of the present invention, the image is converted from the RGB color space to the HSV color space to separate the hue (H), saturation (S) and brightness (V) of the image. Among them, the HSV space is more in line with human perception of color and is easy to process. Then, the color histogram of each area of ​​the image is calculated, and the distribution of different color components is counted. Regarding the histogram dimension, each channel can be divided into 256 gray levels. Then, the sliding window technology is used to extract the color distribution features of the local area.

[0048] For example, HSV analysis is performed on the green belt area in a street scene to identify the main hue (green) of the plant area, and the saturation and brightness differences of different plant types are quantified through histograms.

[0049] By analyzing the local pixel arrangement pattern in the image, texture features reflecting information such as surface roughness and structural complexity of the object are extracted. In some embodiments of the present invention, the spatial relationship of gray value pairs in the image is statistically analyzed using a gray level co-occurrence matrix (GLCM). For example, when texture analysis is performed on the road area in a street view image, the GLCM extracts that the road surface texture has a high uniformity (high homogeneity), while the sidewalk texture has a high contrast (brick structure).

[0050] In summary, it can be understood that the acquired street view panorama is subjected to denoising and enhancement operations, followed by edge detection, color analysis and texture analysis, to extract the contours of objects such as buildings, plants, roads, street lights, traffic lights, motor vehicles, non-motor vehicles, etc.; to extract the color distribution characteristics of objects such as buildings, plants, roads, street lights, traffic lights, motor vehicles, non-motor vehicles, etc.; and to extract the texture patterns of objects such as building surfaces and road surfaces.

[0051] Through the above steps, shallow features reflecting the contours, colors and local structures of objects in the street view panorama are obtained, laying the foundation for subsequent feature registration and classification based on topological analysis.

[0052] After extracting the shallow features, the deep features of the street view panorama are extracted based on the semantic segmentation model. The deep features are the expression of high-level semantic information in the image, which usually include: Geometry: The overall outline and three-dimensional structure of an object, such as the shape of a building or the outline of a motor vehicle; Spatial position: the specific position of an object in the image and its relative position relationship with other objects; Contextual relationship: The semantic connection between an object and its surrounding environment, such as the relationship between traffic lights and roads, buildings and people.

[0053] Deep features are automatically learned from data through deep learning models. Deep learning models usually rely on convolutional neural networks (CNNs) and their extended architectures, such as semantic segmentation models, to extract deep features. Common semantic segmentation model architectures include, but are not limited to, U-Net, DeepLabv3+, and PSP Net.

[0054] It can be understood that the geometry, spatial location, and contextual relationships of deep features provide the input of the classification model for subsequent topological analysis and logical rule alignment.

[0055] The following is a table for reference Figure 2, the target spatial elements of high-intensity urban areas will be integrated with shallow features, deep features and set logical rules.

[0056] Then S4 is performed: based on topological analysis, a topological map between spatial elements is constructed, where each node of the topological map represents each spatial element that is initially identified, and each node carries its shallow and deep features. The edges of the nodes represent the semantic association relationship between the spatial elements, and the structure of the topological map is optimized through preset logical rules.

[0057] Specifically, first, a topological graph is constructed, where each node represents an object element in the street view panorama. Figure 3 and Figure 4 ,It should be noted that the types of object elements are more ,than the types of target spatial elements to be identified.,Nodes contain shallow features (such as edges, colors, ,textures, etc.) and deep features (geometric shapes, spatial positions, ,contextual relationships) of objects, which are represented by specific data ,structures.

[0058] The edges between nodes represent the spatial relationship between the elements of the object. The edge weight reflects the strength of the topological relationship and is calculated by combining the following identified features: Spatial position: the distance between adjacent object elements (such as the shortest distance between a building and a road); Color similarity: measures the closeness of colors by color space distance (such as HSV or Lab space); Semantic relevance: obtained through semantic matching of object categories by semantic segmentation models. For example, buildings are usually adjacent to roads, and street lights are usually close to roads.

[0059] In addition, in some embodiments of the present invention, the relative distance between different spatial elements can also be calculated. The relative distance generally refers to the geometric distance between different spatial elements or the distance under a certain metric. The deep learning semantic segmentation model can further calculate the relative distance between each spatial element by accurately locating the position of each spatial element.

[0060] For example, if there is an obvious gap between the building and the road in the image, or there is a specific layout relationship between the building and the street lamp, the relative distance between the building and the road can be calculated by extracting the geometric information and spatial position of the spatial elements, and it can serve as an important reference information in topological analysis.

[0061] In summary, after extracting the shallow and deep features between spatial elements to construct the nodes of the topological graph, the preset logical rules are used to define the connection relationship between spatial elements in the topological graph, and these rules are used to optimize the structure of the graph. Specifically, the core of the preset logical rules is to adjust the edge weights between nodes in the topological graph based on the geometric shape, spatial position, relative distance and semantic association relationship of spatial elements. You can refer to the above table here, and the summarized logical rules are as follows.

[0062] First, the projection distance between different spatial elements is calculated based on the geometric shape and spatial position of the spatial elements extracted by the deep learning model. The projection distance refers to the distance between two spatial element objects on a plane, usually considering the distance between the closest edges or center points of the two objects.

[0063] After obtaining the projection distance between the spatial elements, in some embodiments, a distance threshold is considered to be set to control whether to adjust the edge weight. For example, if the projection distance between two spatial elements is less than the threshold, it means that they have a strong spatial association, and the edge weight will increase; conversely, if the distance is far, it means that the association between the two spatial elements is weak, and the corresponding edge weight will decrease.

[0064] In addition to spatial location, the semantic association between spatial elements is also an important factor in adjusting edge weights. Semantic association refers to the functional connection between different spatial elements in the actual environment. For example, referring to the above table, there is usually a close semantic association between buildings and roads, because buildings are usually distributed along roads, and roads provide traffic channels for buildings. There is also a semantic association between street lights and roads, because street lights are generally installed on both sides of the road to provide night lighting.

[0065] For spatial elements with functional associations, the weight of their connecting edges should be increased. For spatial elements without functional associations, such as buildings and street lights, if there is no direct functional connection, the edge between the two spatial elements should be deleted. Figure 3 and Figure 4 ,There are 13 nodes in the initial topological graph, which represent different spatial elements, such as buildings, roads, green spaces, street lights, etc. The goal of optimization is to adjust the weights of edges and delete the edges between nodes that have no functional association, so as to simplify the graph structure and improve the practical significance and classification efficiency of the topological graph.

[0066] Assume that the initial topology is as shown in the following table.

[0067] During the topology optimization process, the functional context relationship is determined: There is a functional relationship between the building and the road. The building needs to provide transportation connection through the road, so this edge is maintained.

[0068] There is usually no direct functional relationship between buildings and streetlights (except in special cases, such as building lighting relying on streetlights), so this side should be deleted.

[0069] The road has a functional relationship with the parking lot, and the parking lot needs to access the road, so this edge is maintained.

[0070] The streetlight has a functional relationship with the square (it usually illuminates the square), so this side is retained.

[0071] The bus stop has a functional relationship with the road, and the bus stop is usually located next to the road, keeping this side.

[0072] The topology nodes after optimization are shown in the following table.

[0073] In some feasible embodiments of the present invention, a rule set may be established in advance to guide the adjustment of the weights. For example, the rule set stipulates that the connection between buildings, roads and street lights has a high weight, while the connection between vehicles and street lights has no significant functional association, so its weight is zero or low.

[0074] It is understandable that under the action of preset logical rules, the edge weights in the topological map will be adjusted to reflect the real relationship between spatial elements. On this basis, the graph convolution network (GCN) operation is performed to further optimize the node features and spatial distribution in the topological map. The graph convolution operation updates the features of each node by weighted aggregation of the features of its neighboring nodes. This process helps to improve the representation ability of the topological map, especially in the spatial element relationships of large-scale high-intensity urban areas.

[0075] Specifically, the edge weight is updated by the feature correlation of neighbor nodes, so that the edges with stronger semantic and spatial correlation have higher weights. The update formula is: in, For the Layer Node The eigenvector of For Node The set of neighbor nodes of is the learnable weight matrix of the graph convolutional layer; For Node and The degree of is a non-linear activation function, such as ReLU.

[0076] For each node , by aggregating its neighbor nodes The features of the node are updated based on the features of the node. Aggregation methods usually include weighted average, weighted sum, etc. For each edge, the weight of the edge will affect the degree of aggregation of the neighbor node features. The edge with a larger weight will have a greater impact on the neighbor node features, while the edge with a smaller weight will reduce the impact of the neighbor node features. According to the aggregation results, combined with the original node features, the node features are updated through a nonlinear activation function. The updated node features can better capture the relationship and topological structure between spatial elements.

[0077] The result of graph convolution will update the node features in the topological graph, thereby optimizing the structure of the topological graph. Specifically, after the graph convolution operation, the update of node features can make the topological graph more consistent with the actual spatial relationship. For example, the connection between spatial elements such as buildings, roads, and street lights will become clearer, and the adjustment and update of edge weights ensure that the actual semantic associations between these spatial elements are reasonably reflected.

[0078] The optimized topological map obtained after graph convolution can provide more accurate input features for subsequent classification tasks, help the classification model better identify spatial elements, and generate accurate spatial distribution maps.

[0079] In summary, it can be understood that by adjusting the weights of edges based on geometric shapes, spatial positions, semantic associations, etc., and optimizing the topological map in combination with graph convolution operations, the modeling of the relationships between spatial elements can be effectively improved, ensuring that the topological map can accurately reflect the spatial structure in the actual scene. The optimized topological map not only improves the classification accuracy of spatial elements, but also enhances the interpretability and consistency of the classification results, providing strong support for subsequent urban street scene analysis and spatial element distribution prediction.

[0080] Finally, S5 is performed: the constructed topological map is input into the pre-trained classification model, and the classification model outputs the classification results of each spatial element in the street view panorama and its spatial distribution map according to the node features and topological structure. The core goal of this step is to use the pre-trained classification model. The model classifies each node by learning the relationship between node features and topological structure, combining the feature space and the contextual semantic information of the spatial elements, and finally outputs the category label of each spatial element in the street view panorama and their spatial distribution in the image.

[0081] The pre-trained classification model combines deep learning technology and topological map information to efficiently classify spatial elements based on their shallow and deep features. Through a multi-layer neural network and message passing mechanism, the classification model can accurately identify and classify spatial elements in street view panoramas and generate a spatial distribution map of spatial elements.

[0082] Regarding the training process of the classification model, it starts with the input topology map. Through the deep learning classification model, the feature vector of each node is extracted from the original spatial element description. This feature vector is passed to the classification model as the representation vector of the node.

[0083] Specifically, referring to the above-mentioned graph neural network, the adjacency relationship and node features in the graph structure are used to aggregate information layer by layer. The output of each layer updates the node features. In the process of message passing, the weight of the edge is used to adjust the contribution of the neighbor node to the target node, gradually integrating shallow information into high-level semantics.

[0084] The feature vector of each node is then input into a predefined classifier. The classifier determines the category of the spatial element based on the feature vector of each node, and outputs the probability distribution of the category to which the node belongs through the fully connected layer and the Soft Max function, thereby outputting the category label to which the node belongs.

[0085] Once the classification results of all nodes are output, the classification model will combine the classification label and spatial position relationship of each node to generate a spatial distribution map of spatial elements in the street view panorama. The spatial distribution map not only reflects the category of each spatial element, but also shows their distribution in the geographic space.

[0086] In summary, it can be understood that through this deep learning-based classification model, combined with node features and adjacency relationships in the topological map, spatial element classification can not only accurately identify individual elements, but also efficiently classify them according to spatial location and context information. In this process, the message passing mechanism and feature fusion technology can effectively capture the relationship between spatial elements and ultimately output a more accurate spatial distribution map. This method not only improves the classification accuracy, but also enhances the spatial consistency and interpretability of the classification results, providing strong technical support for urban street scene analysis.

[0087] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. The patent protection scope of the present invention shall be based on the claims. All equivalent structural changes made using the contents of the description and drawings of the present invention should also be included in the protection scope of the present invention.

Claims

1. A quantitative analysis method of spatial elements of urban high-intensity areas based on image data, characterized in that: The following steps are involved: Obtain street view panoramas from selected observation points in high-intensity areas of the city; Based on edge detection, color analysis and texture analysis, shallow feature extraction is performed on the street view panorama, and spatial elements in the street view panorama are preliminarily identified in combination with predefined shallow feature templates; For each of the spatial elements initially identified, the deep features thereof are obtained through a deep learning model; wherein the deep features include the geometric shape and spatial position of each spatial element, and the semantic association relationship between different spatial elements; Based on topological analysis, a topological graph between spatial elements is constructed, wherein each node of the topological graph represents each spatial element that is initially identified, and each node carries its shallow and deep features, and the edges of the nodes represent the semantic association relationship between the spatial elements, and the structure of the topological graph is optimized through preset logical rules; The constructed topological map is input into a pre-trained classification model, and the classification model outputs the classification result of each spatial element in the street view panorama and its spatial distribution map according to the node characteristics and the topological structure.

2. The method for quantitative analysis of spatial elements of urban high-intensity areas based on image data according to claim 1 is characterized in that: Before performing shallow feature extraction on the acquired street view panorama, image preprocessing is performed on the street view panorama. The image preprocessing includes denoising, image enhancement and image segmentation operations, corresponding to: removing random noise in the street view panorama and retaining edge information; The brightness and contrast of the street view panorama after noise removal are analyzed, histogram equalization is performed on the local area of ​​the image, and contrast enhancement is performed on the low-contrast area.

3. The method for quantitative analysis of spatial elements of urban high-intensity areas based on image data according to claim 1 is characterized in that: Perform edge detection, color analysis and texture analysis on the street view panorama to extract the shallow features, specifically: Perform Canny edge detection on the acquired street view panorama, calculate the intensity of grayscale changes in the street view panorama, and use gradient information to locate the boundaries of objects in the street view panorama; Extract the color distribution characteristics of each object located by HSV color space; The gray level co-occurrence matrix is ​​used to extract the surface texture information of each located object.

4. The method for quantitative analysis of spatial elements of urban high-intensity areas based on image data according to claim 3 is characterized in that: Each of the objects located in the street view panorama and the shallow features extracted therefrom are matched with the preset shallow feature template, and the spatial elements represented by each object are preliminarily identified through a matching algorithm.

5. The method for quantitative analysis of spatial elements of urban high-intensity areas based on image data according to claim 1 is characterized in that: For each of the spatial elements initially identified, the corresponding deep features are extracted through a deep learning semantic segmentation model, including the geometric shape and spatial position of each spatial element, the relative distance between different spatial elements, and the contextual semantic association relationship between different spatial elements.

6. The method for quantitative analysis of spatial elements of urban high-intensity areas based on image data according to claim 5 is characterized in that: The preset logic rules include: The weight value of the edge is adjusted based on the geometric shape and spatial position of the spatial element preliminarily identified, and the relative distance between different spatial elements; wherein the projection distance between different spatial elements is calculated, a distance threshold is set, and the weight is adjusted proportionally; The weight value of the edge is adjusted according to the semantic association relationship between different spatial elements; if two spatial elements have a functional association in the real scene, the weight is increased; if the two spatial elements have no functional association, the edge of the node represented by the two spatial elements is deleted.

7. The method for quantitative analysis of spatial elements of urban high-intensity areas based on image data according to claim 6 is characterized in that: By optimizing the preset logic rules, updating the edge weights, and performing graph convolution operations on the topological graph, the node features and spatial distribution in the topological graph are optimized.

8. The method for quantitative analysis of spatial elements of urban high-intensity areas based on image data according to claim 1 is characterized in that: The pre-trained classification model is based on deep learning technology to classify the spatial elements of the high-intensity area. During the training process of the classification model, the topological map is input and the corresponding feature vector of each node is generated according to the shallow features and deep features of each node; The classification model aggregates the features of the target node's neighbor nodes through the message passing mechanism according to the adjacency relationship and edge weight of the target node to be identified, and fuses them with the current features of the target node to generate the target node features containing context information; The feature vector of each node is input into the defined classifier to output the classification result of each node. The classification result of each node is combined with the spatial position relationship in the topological map to generate a spatial distribution map of the spatial elements in the street view panorama.

Citation Information

Patent Citations

  • Urban and rural landscape image quantification and improvement method based on community differentiation

    CN117079124A

  • Historical urban style integrity evaluation method and system based on deep learning

    CN119477804A

  • Method for updating road signs and markings on basis of monocular images

    US20230135512A1