A method for identifying urban illegal areas based on collaborative multi-view images of air, ground and space

Through the multi-view image collaborative recognition method, a VGG building discrimination model is constructed using remote sensing images, drone images and ground video sensor images, which solves the high labor cost and occlusion problems in urban illegal areas detection, and achieves efficient and accurate identification of illegal areas.

CN117218524BActive Publication Date: 2025-08-12HANGZHOU DIANZI UNIV +2
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310972334.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-08-03
Publication Date
2025-08-12
Estimated Expiration
2043-08-03

AI Technical Summary

Technical Problem

The prior art has problems such as high labor and calculation costs, low identification efficiency, and high-density buildings and greening shading in urban violation areas.

Method used

A multi-view collaborative recognition method of remote sensing images, drone images and ground video sensor images is adopted to construct a VGG building discriminant model through feature extraction and fusion, and a drone image containing ground image building features is generated using a conditional generation adversarial network, and a VGG model is used to identify illegal areas.

Benefits of technology

The efficiency and accuracy of illegal areas detection have been improved, the identification difficulties caused by high-density buildings and greening shading have been solved, and the rapid and accurate identification of illegal areas has been achieved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117218524B_ABST
    Figure CN117218524B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for identifying urban illegal areas based on the collaboration of multi-view images of air, space, and land. The method extracts and fuses the features of multi-view image information to obtain urban land use data and urban illegal areas pre-detected by time-series remote sensing. The remote sensing images, ground images, and drone images of the urban illegal areas are obtained, and the features of the urban illegal areas are generated by a conditional generative adversarial network. The images are input into a VGG model for training. Finally, the building discrimination model based on the VGG pre-trained model is used to identify the urban illegal areas. The present invention integrates air, space, and land image and video information to detect and discriminate urban illegal areas, solving the problems of traditional methods that only use remote sensing images / ground images or drone images, the problem that performance cannot be guaranteed when facing occlusion by buildings and urban greenery, and the problem of limited drone acquisition range, thereby achieving rapid and accurate detection and discrimination of urban illegal areas.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of land management and urban and rural planning, and relates to a method for identifying urban illegal areas based on collaborative multi-perspective images of air, ground and sky. Background Art

[0002] Illegal areas refer to various buildings and structures that have been constructed without the permission of the competent authorities, including illegal buildings, illegal occupation of public areas and streets, etc. (such as buildings, stalls, shops and garbage dumps that illegally occupy public roads). Such buildings usually encroach on safe passages and illegally occupy resources such as arable land, causing negative impacts on urban public spaces and the ecological environment.

[0003] Currently, the detection and identification of illegal areas primarily relies on manual visits, ground-based video sensors, primarily those used in urban surveillance systems, drone-generated images, and remote sensing monitoring. Each of these methods has its own unique characteristics. Manual visits are labor-intensive, time-consuming, and economically expensive, and their efficiency is relatively low. Ground-based video sensors, such as those used in urban surveillance systems, are currently deployed at a high rate and over a wide area in most cities, enabling long-term monitoring. However, the volume of recorded video data is substantial, and extracting information about illegal areas from this data requires significant computational effort. Remote sensing imagery provides a wider coverage area than urban surveillance systems, but its spatiotemporal granularity is not as high. Furthermore, some urban illegal areas are often concealed. For example, in densely populated urban environments, obstruction by buildings and urban greenery makes it difficult for urban surveillance systems and satellites to obtain image information beneath these obstructions, limiting the effectiveness of monitoring methods based on ground-based video sensors and satellite remote sensing images. This is where drones' flexibility and maneuverability in capturing image information beneath these obstructions becomes crucial. Leveraging the strengths of these methods can improve timeliness and accuracy, while reducing computational and labor costs. Therefore, it is particularly important to obtain multi-view image data of accomplices and collaboratively identify illegal areas from multiple perspectives.

[0004] In summary, in terms of urban illegal area identification, there is an urgent need for a multi-perspective collaborative identification method to achieve low-cost and high-recognition rate illegal area discrimination. Summary of the Invention

[0005] The present invention aims to address the shortcomings of existing urban illegal area identification technologies by proposing a method for identifying illegal areas in cities based on collaborative, multi-perspective imagery from air, ground, and space. This method utilizes remote sensing imagery, real-time drone images, and images recorded by ground-based video sensors. By extracting and fusing multi-perspective image information, it constructs a model for identifying illegal areas based on a visual geometry group network. This method addresses the challenges of multi-perspective collaborative identification of illegal areas in urban environments with occlusion caused by tall buildings and greenery, thereby improving the efficiency and accuracy of illegal area detection.

[0006] The specific steps include:

[0007] Step 1: Abnormal area detection, positioning, and drone path planning:

[0008] Acquire multi-temporal remote sensing images of the target city, use land use planning data and time-series remote sensing images to pre-detect areas of land cover anomalies and obtain their location information. Use a path planning method based on a Rapidly Exploring Random Tree algorithm to plan drone routes in these anomaly areas.

[0009] Step 2: UAV and ground-view image acquisition: Based on the path planned in step 1, use a drone to obtain a multi-angle UAV-view image sequence of the abnormal area; based on the abnormal area positioning information obtained in step 1, use a ground-based video sensor to obtain a ground-view street view image sequence of the abnormal area;

[0010] Step 3: Generate drone images containing building features in ground images: Input the drone images and ground images into a conditional generative adversarial network to generate drone images containing building features in the ground images and construct a sample image library of abnormal areas. The sample image library of abnormal areas is divided into a training set and a test set, and the training set is preprocessed.

[0011] Step 4: Visual Geometry Group Network (VGG) building recognition model training: The VGG model is trained based on the preprocessed training set until the VGG model is trained to the pre-set target parameters. The trained model is named the building recognition model based on the VGG pre-trained model.

[0012] Step 5: Identify illegal areas: Use the trained network model to identify the abnormal area sample library, and determine whether there are illegal areas in the abnormal area based on the model.

[0013] The method uses land use data and time-series remote sensing images to pre-detect areas of abnormal land cover. It then obtains ground and drone images of the abnormal areas, and uses a conditional adversarial network to generate drone images containing building features from the ground images. The image dataset is divided into a training set and a test set. The training set is pre-processed, and a VGG network model with relatively reasonable parameter weights is trained using selected drone images similar to the ground images. This network model is then used to identify whether buildings exist in the selected pre-selected areas. If a building exists in the abnormal area, it is considered a violation.

[0014] The present invention has the following technical effects: 1) at the city scale, it uses land use data and time-series remote sensing images to pre-detect areas of abnormal land cover; 2) in abnormal areas, it combines drone-viewed images with ground-viewed images to collaboratively determine whether there are illegal areas within the abnormal areas, thereby improving recognition accuracy; 3) it solves the problem of illegal areas being unable to be identified due to obstruction by high-density buildings and greenery in cities, thereby improving recognition rate;

[0015] This invention, for the first time, proposes integrating air-space, ground-level, and video information to detect and identify urban traffic violations. This method addresses the shortcomings of traditional methods that rely solely on remote sensing imagery, ground-based imagery, or drone imagery. It overcomes the performance issues faced by traditional methods when obscured by buildings and urban greenery, as well as the limited range of drone data. It enables rapid and accurate detection and identification of urban traffic violations. The ability to monitor and identify urban traffic violations plays a vital role in urban management, planning, and development, as well as in the construction of smart cities and enhancing urban appearance. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] Figure 1 Flowchart of the present invention. DETAILED DESCRIPTION

[0017] The present invention will be fully described below in detail with reference to the accompanying drawings.

[0018] like Figure 1 As shown in FIG, a method for identifying urban traffic violation areas based on collaborative multi-view images of air, ground and space includes the following steps:

[0019] Step 1: Obtain multi-temporal remote sensing images of the target city. Leverage existing measured land use data and time-series remote sensing images to pre-detect areas of land cover anomalies and obtain their location information. Plan the drone's aerial photography route within the anomaly area using a path planning method based on a Rapidly Exploring Random Tree algorithm.

[0020] Step 2: Based on the path planned in step 1, use the drone to obtain a multi-angle drone view image sequence S of the abnormal area aerial-view Based on the abnormal area positioning information obtained in step 1, the ground video sensor is used to obtain the ground perspective street view image sequence S of the abnormal area street-view ;

[0021] Step 3: Input the drone images and ground images of the abnormal area into the generative adversarial network. Using the drone images as input and the ground images as labels, the network generates drone images containing architectural features of the ground images, and constructs a sample image library of abnormal areas. The sample image library of abnormal areas is then preprocessed, including training set annotation and data augmentation. The dataset consisting of the preprocessed images is divided into training and test sets according to a preset ratio. The training and test sets are divided according to the object categories in the labels of the original data in a ratio of 8:2.

[0022] The goal of the constructed conditional generative adversarial network is to synthesize drone images while replicating the content of the reference ground image. That is, using the conditional generative adversarial network as an image synthesis model, using the drone image as a condition and the ground image as a label, it generates drone images that contain architectural features from the ground image. The specific steps are as follows:

[0023] Step 1: Input the drone image and the ground image into the conditional generative adversarial network;

[0024] Step 2: Introduce the generator network. During the training process, use the drone image as a conditional input to obtain building condition information based on the ground image. Use the least squares structure generator to learn the architectural features of the drone perspective image, and then generate the drone image containing the architectural features of the ground image.

[0025] Step 3: Introduce a discriminator network based on the PatchGAN classifier. This network determines image differences by comparing the patch sizes between the drone image containing ground-level building features and the ground image. After identification by the discriminator, the residuals generated during the generation of the drone image containing ground-level building features are instance-normalized. Each convolutional layer in the generation of the drone image containing ground-level building features is spectrally normalized. Determine whether the generated drone image containing ground-level building features differs from the ground image. If so, return to Step 2; otherwise, proceed to Step 4.

[0026] Step 4: Calculate the loss function and perform backpropagation to obtain the drone image containing the building features of the ground image.

[0027] The loss function is:

[0028] L=L_G+L_D+L_perceptual

[0029]

[0030]

[0031] L_perceptual=λ+E[|ω(G(z))-ω(x)| 2 ]

[0032] Where L_G is the generator loss function, L_D is the discriminator loss function, L_perceptual is the perceptual loss function, G represents the generator, D represents the discriminator, x represents the real ground image, z represents the input noise of the generator, ω represents the pre-trained feature network, and λ represents the weight of the loss function.

[0033] Step 4. Obtain the parameters of the VGG model pre-trained on the ImageNet dataset (which can be downloaded from the relevant website) and initialize the VGG model based on the parameters; input the data in the training set and the prior box, perform forward propagation of the VGG model, calculate the loss value between the predicted result and the true label according to the loss function of the VGG model, and name the trained model as the building discrimination model based on the VGG pre-trained model.

[0034] In this embodiment, in order to ensure that the VGG model (Visual Geometry Group Network) can learn architectural features, the training photos need to be labeled; because only one object feature in the violation area needs to be identified, only one color needs to be defined for labeling; in order to ensure that the VGG model can learn architectural features, the training set data also needs to be enhanced, that is, a preset number of pictures in the training set are selected each time, and the pictures are added with noise, flipped, rotated, and scaled.

[0035] Step 5: Use the trained VGG building recognition model to identify the abnormal area sample library and convert the aerial image sequence S of the abnormal area into aerial-view and ground image sequence S street-view Input into the VGG building recognition model for prediction analysis to obtain the normalized probability P aerial-view 、P street-view .

[0036] According to P aerial-vie,,street-vie, The value of the total probability formula can be used to obtain the building abnormality probability P 0uildin4 The value of:

[0037] Among them, P aerial-view is the probability that the aerial image is judged as a building by the recognition model, P street-view is the probability that the ground image is judged as a building by the recognition model, ∑S aerial-vie, is the total number of aerial images, ∑S street-vie, is the total number of ground images.

[0038] Determine the probability of building abnormality P ille4al-0uildin4 ∈{0,1}, where 1 indicates a violation area and 0 indicates a non-violation area:

Claims

1. A method for identifying urban traffic violation areas based on collaborative multi-view images of air, ground, and space, characterized by: The specific steps include: Step 1: Abnormal area detection, positioning, and drone path planning: Acquire multi-temporal remote sensing images of the target city, use land use data and time-series remote sensing images to pre-detect areas of land cover anomalies and obtain their location information; plan drone aerial photography routes in anomaly areas based on a path planning method using a rapidly expanding random tree sampling algorithm; Step 2: UAV and ground-view image acquisition: Based on the path planned in step 1, use a drone to obtain a multi-angle UAV-view image sequence of the abnormal area; based on the abnormal area positioning information obtained in step 1, use a ground-based video sensor to obtain a ground-view street view image sequence of the abnormal area; Step 3: Generate drone images containing building features in ground images: Input the drone images and ground images into the conditional generative adversarial network to generate drone images containing building features in the ground images and build a sample image library of abnormal areas; After preprocessing the abnormal area sample image library, the preprocessed abnormal area sample image library is divided into a training set and a test set according to a preset ratio; The UAV image and the ground image are input into a conditional generative adversarial network to generate a UAV image containing the building features of the ground image; specifically: Step 1: Input the drone image and the ground image into the conditional generative adversarial network; Step 2: Introduce a generator network. During training, the UAV image is used as a conditional input. Building condition information is obtained based on the ground image. The least squares structure generator is used to learn the building features of the UAV image. Then, the UAV image containing the building features of the ground image is generated. Step 3: Introduce a discriminator network based on the PatchGAN classifier. The image difference is determined by comparing the patch size between the drone image containing the ground image building features and the ground image. After the discriminator determines the difference, the residual generated in the process of generating the drone image containing the ground image building features is instance-normalized, and each convolutional layer of the drone image containing the ground image building features is spectrally normalized. Determine whether there is a difference between the drone image containing the ground image building features and the ground image. If so, return to step 2; otherwise, proceed to step 4. Step 4: Calculate the loss function and perform backpropagation to obtain the drone image containing the building features of the ground image; Step 4: VGG building recognition model training: train the VGG model according to the preprocessed training set until the VGG model is trained to the pre-set target parameters, and name the trained model as the building recognition model based on the VGG pre-trained model; Step 5: Identify illegal buildings: Use the trained network model to identify the abnormal area sample library and determine whether there are illegal areas in the abnormal area based on the model.

2. The urban traffic violation area identification method based on air-ground multi-view image collaboration as claimed in claim 1, characterized in that: In the step three, the abnormal area sample image library is preprocessed, and the preprocessing includes: training set annotation and data enhancement.

3. The urban traffic violation area identification method based on air-ground multi-view image collaboration as claimed in claim 1, characterized in that: The step 5 is specifically as follows: using the trained VGG building recognition model to identify the abnormal area sample library, and converting the aerial image sequence S of the abnormal area into aerial-view and ground image sequence S street-view Input into the VGG building recognition model for prediction analysis to obtain the normalized probability P aerial-view 、P street-view ; According to P aerial-view 、P street-view The value of , through the total probability formula, can be obtained as the probability of building abnormality The value of: ; Among them, P aerial-view is the probability that the aerial image is judged as a building by the recognition model, P street-view is the probability that the ground image is judged as a building by the recognition model, is the total number of aerial images, is the total number of ground images; Determine the probability of building abnormality , where 1 indicates a violation area and 0 indicates a non-violation area: .

Citation Information

Patent Citations

  • Automatic illegal building identification method based on full convolutional neural network

    CN112307873A

  • Intelligent patrol system of unmanned aerial vehicle based on parking apron cluster

    CN112422783A