Law enforcement image unsupervised detection method based on edge computing power

By integrating edge computing devices and fusing adaptive cross-layer optimization networks on drones, the high latency and labeling dependency issues of drone law enforcement image processing are resolved, enabling real-time, unsupervised illegal building detection by drones, and improving the real-time and intelligent level of detection.

CN120808120AActive Publication Date: 2025-10-17TIANJIN SURVEYING & MAPPING INST CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202511315703.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-15
Publication Date
2025-10-17
Estimated Expiration
2045-09-15

AI Technical Summary

Technical Problem

In existing technologies, drone law enforcement image processing relies on centralized cloud processing, resulting in high latency, high bandwidth pressure, concentrated computing pressure, and high operation and maintenance costs. In addition, the supervised learning model is highly dependent on labeled data, making it difficult to effectively identify targets such as illegal buildings in complex scenarios.

Method used

An unsupervised detection method based on edge computing power is adopted, and OrangePi AIPro devices are used for real-time image processing. Combined with the fusion of adaptive cross-layer optimization network and scale-adaptive asymmetric convolution module, real-time image segmentation and recognition during drone flight are achieved, and feature extraction and segmentation are performed through unsupervised optimization of the objective function.

Benefits of technology

The drone can complete the real-time segmentation and identification of illegal buildings and other targets during flight, reducing latency and bandwidth usage, improving the level of intelligent detection, adapting to complex scenarios, and reducing dependence on labeled data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120808120A_ABST
    Figure CN120808120A_ABST
Patent Text Reader

Abstract

The invention discloses a law enforcement image unsupervised detection method based on edge computing power. The law enforcement image unsupervised detection method comprises the following steps of 1) directly inputting collected image data into edge computing equipment and performing preprocessing to obtain a to-be-processed image; 2) inputting the to-be-processed image into a fusion adaptive cross-layer optimization network for feature extraction and segmentation, wherein the fusion adaptive cross-layer optimization network comprises a double-path cross-layer feature extractor and an unsupervised optimization objective function formed by the weighted sum of an in-graph comparison loss function and a local reconstruction loss function; and 3) the edge computing device generates a segmentation result in real time according to the finally output feature map. According to the invention, the cost pressure of the system on bandwidth and computing power is reduced, the autonomy and intelligent level of the unmanned aerial vehicle in cruising are improved, the efficient closed loop of collection, processing and feedback can be realized during urban law enforcement monitoring, and the method has significant application value and social benefit.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of video fusion, in particular to a law enforcement image unsupervised detection method based on edge computing. BACKGROUND

[0002] In the prior art, the unmanned aerial vehicle is usually only used as a data acquisition platform to transmit law enforcement images to the cloud or a local server for centralized processing. For example, a city house illegal construction intelligent inspection method, device, system and equipment are disclosed in Chinese patent CN115984694A. The city house illegal construction intelligent inspection method disclosed in the embodiments of the present application acquires images of a target geographical area collected by an unmanned aerial vehicle or a satellite at different times by means of the unmanned aerial vehicle or the satellite to inspect and photograph the target geographical area. The images of the target geographical area at different times are spliced, and a pre-trained semantic segmentation model for identifying a house area is used to identify house area images of the target geographical area. Whether there are newly built houses or demolished houses in the target geographical area is determined by means of the house area images. If there are newly built houses or demolished houses, a warning of illegal construction is issued.

[0003] However, the above-mentioned patent still has the following disadvantages: High latency, the cloud centralized processing method relies heavily on network transmission and cannot guarantee real-time performance; Large bandwidth pressure, high-resolution real-time image transmission occupies a large amount of bandwidth, which is not conducive to large-scale application; High computing pressure, high server load and high operation and maintenance cost; Strong dependence on labeling, supervised learning model training requires a large amount of labeled data, and the effect is limited when migrating to a new scene. The recognition effect is limited in complex scenes such as similar building appearance, light change and serious occlusion. SUMMARY

[0004] Therefore, the technical problem to be solved by the present application is to overcome the defects in the prior art, and a law enforcement image unsupervised detection method based on edge computing is proposed. The method has no dependence on the network environment, can realize real-time image processing and analysis during flight, and can complete the segmentation and identification of illegal buildings and other targets without relying on external network transmission.

[0005] To solve the above technical problems, the present application provides a law enforcement image unsupervised detection method based on edge computing, comprising: A law enforcement image unsupervised detection method based on edge computing, comprising the following steps: 1) inputting the collected image data directly into an edge computing device and performing preprocessing to obtain a to-be-processed image; 2) inputting the image to be processed into a fusion adaptive cross-layer optimization network for feature extraction and segmentation, wherein the fusion adaptive cross-layer optimization network comprises a dual-path cross-layer feature extractor and an unsupervised optimization objective function composed of a weighted sum of an intra-graph contrast loss function and a local reconstruction loss function, and the dual-path cross-layer feature extractor comprises an initial feature encoding module, a main feature extraction path composed of a plurality of identical feature encoding modules connected in series, and a feature fusion path connected in parallel with the main feature extraction path and performing cross-layer feature fusion; wherein each of the feature encoding modules is a scale-adaptive asymmetric convolution module, and the scale-adaptive asymmetric convolution module is implemented by the following sub-steps, 2.1.1) inputting an input image or a feature map as an original input and performing preliminary feature extraction to generate shared features after batch normalization and activation function processing, 2.1.2) inputting the shared features into three parallel asymmetric convolution branches with different receptive field scales to extract short-range features, medium-range features and long-range features, 2.1.3) element-wise multiplying the short-range features and the medium-range features to generate an attention map, and weighting the long-range features with the attention map to obtain gated features, 2.1.4) aggregating the gated features, the original input copy features and the long-range features to output a feature map; 3) the edge computing device generates a segmentation result in real time according to the final output feature map.

[0006] As one of the preferred schemes, in step 1), the edge computing device is an OrangePi AIPro device, and the collected image data does not need to rely on external network transmission and is directly connected to the OrangePi AIPro device.

[0007] As one of the preferred schemes, in step 1), the preprocessing includes at least one of the following operations: size standardization, illumination and color correction, noise filtering.

[0008] As one of the preferred schemes, the three branches are a short-range feature branch, a medium-range feature branch and a long-range feature branch, each branch is composed of a depth separable convolution with a depth separable convolution connected in series to capture features in different receptive field ranges, is a size parameter of a convolution kernel in the depth separable convolution.

[0009] As one of the preferred schemes, the segmentation result is output to a user terminal in the form of a contour label and a binary segmentation map.

[0010] ​​As one of the preferred schemes, in step 3), the segmentation result comprises contour labeling and binary segmentation of the target region, and the target region comprises illegal buildings, facilities operating on the road or garbage dumping areas.

[0011] As one of the preferred schemes, the unsupervised optimization objective function is specifically implemented as, 2.2.1) Calculate the intra-graph contrast loss: randomly sample anchor blocks, define positive sample blocks that are spatially adjacent and visually similar, and negative sample blocks that are significantly different, and optimize the loss by minimizing the feature distance between the anchor blocks and the positive sample blocks and maximizing the feature distance between the anchor blocks and the negative sample blocks; 2.2.2) Calculate the local image block reconstruction loss: input the feature embedding of the anchor block into the lightweight reconstruction module after the encoder to reconstruct the original image block, and use the mean square error between the original image block and the reconstructed block as the loss; 2.2.3) Weighted sum of intra-graph contrast loss and local image block reconstruction loss according to the preset weight, as the final unsupervised optimization objective function.

[0012] The technical scheme of the present application has the following advantages: To solve the above problems, the present application provides an embedded edge computing unmanned aerial vehicle city law enforcement real-time image unsupervised detection method. The scheme embeds a high-performance edge computing device OrangePi AIPro into the unmanned aerial vehicle, so that it can realize real-time image processing and analysis during flight, without relying on external network transmission to complete the segmentation and identification of illegal buildings and other targets. The unsupervised optimization objective function (unsupervised detection algorithm) is free from the dependence on a large amount of labeled data, and can realize efficient detection using the image itself features, so that the unmanned aerial vehicle can collect and process in real time during cruising, and the results are fed back to the user end in real time, thereby significantly reducing the delay, reducing the bandwidth occupation, and improving the intelligent level of detection and the practical application value.

[0013] The technology can be widely applied to illegal building automatic segmentation detection, smart city comprehensive law enforcement, emergency management rapid response and other scenes; and can also be extended to other city management tasks, such as road occupation operation identification, exposed garbage dumping monitoring, public facility damage inspection and the like, to provide efficient and intelligent technical support for city management. BRIEF DESCRIPTION OF DRAWINGS

[0014] In order to more clearly illustrate the specific embodiments of the present application or the technical solutions in the prior art, the following will briefly introduce the drawings needed to be used in the specific embodiments or prior art description. Obviously, the drawings in the following description are some embodiments of the present application, and those skilled in the art can also obtain other drawings according to these drawings without creative labor.

[0015] Figure 1 The structural schematic diagram of the edge-computing-based law enforcement image unsupervised detection method provided in the first embodiment of the present application. DETAILED DESCRIPTION

[0016] The technical solutions of the present application will be described clearly and completely below with reference to the drawings. Obviously, the described embodiments are part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor fall within the scope of protection of the present application.

[0017] The edge-computing-based law enforcement image unsupervised detection method of the present application comprises the following steps: 1) The collected image data is directly input into the onboard edge computing device and preprocessed to obtain a to-be-processed image. Specifically, the edge computing device such as OrangePi AIpro (the edge computing device is integrated into a UAV; the UAV is equipped with a high-resolution camera, which collects a high-definition video stream of a target area in real time during cruising. The video stream, i.e. image data, is transmitted to the onboard OrangePi AIPro device in real time, which decodes it in real time to convert the continuous video signal into a complete image frame sequence (i.e. frame-by-frame conversion into an image). Subsequently, the OrangePi AIpro uses its built-in multi-core CPU and NPU with up to 8TOPS AI computing power to quickly preprocess each frame of image in the sequence. The collected real-time video stream data is directly input into the onboard OrangePi AIPro device without relying on external transmission. At the same time, the edge computing device data preprocessing, such as OrangePi AIPro, directly pre-processes the image after receiving it, and the preprocessing steps include: size standardization; illumination and color correction; noise filtering. This step ensures that the input image still has high quality in complex environments (strong light, shadow, low illumination).

[0018] 2) The to-be-processed image is input into a fusion adaptive cross-layer optimization network for feature extraction and segmentation. The fusion adaptive cross-layer optimization network comprises a dual-path cross-layer feature extractor and an unsupervised optimization objective function composed of a weighted sum of an intra-graph contrast loss function and a local reconstruction loss function. The dual-path cross-layer feature extractor comprises an initial feature encoding module, a main feature extraction path and a feature fusion path respectively composed of a plurality of serially connected same feature encoding modules, The main feature extraction path and the feature fusion path receive and process the feature maps output from the previous depth level at each depth level feature encoding module. The feature fusion path also introduces feature maps output from different depth levels in the main feature extraction path through a plurality of adapter layers.

[0019] Conventional deep network encoders usually adopt a one-way linear structure with layer-by-layer deepening. Such a structure can capture global semantic information in deep features, thus judging regional categories, but loses accurate local boundary information due to the step-by-step decrease in spatial resolution; on the contrary, shallow features retain pixel-level spatial details such as edges, but lack high-level semantic guidance. This split between semantic and spatial information often leads to problems such as boundary blurring and small target missing in segmentation results.

[0020] To solve the problem of split between deep semantic information and shallow spatial detail information in traditional encoders, as shown in Figure 1 , the present application designs a dual-path cross-layer feature extractor. The core of the architecture is to realize the collaborative optimization of multi-level features through parallel backbone feature extraction paths (left path) and feature fusion paths (right path) and dense cross-layer connections between them. Specifically, given an input image , it is first mapped to shallow feature maps by an initial feature encoding module, which serves as the input to the entire dual-path structure. The backbone feature extraction path is composed of N serial feature encoding modules, where N can be selected from 3 to 10 according to actual conditions. In the embodiments of the present application, three will be used as an example for demonstrative illustration, to extract hierarchical features from low-order information such as edges and textures to high-order semantic information such as targets and scenes layer by layer from bottom to top. The present application represents the feature maps output by the i-th deep level of the backbone feature extraction path as . In parallel, the feature fusion path is also composed of N serial feature encoding modules. At each i-th deep level, the feature encoding module not only receives and processes the feature maps output from the previous deep level, but more importantly, it actively introduces feature maps from different deep levels in the backbone feature extraction path (such as and shallower , ) through multiple adapter layers (such as implemented by 1x1 convolution). These feature maps from different paths and deep levels are fused in an element-wise addition manner after completing channel dimension unification and feature alignment through the adapter layers, forming a fusion feature rich in multi-scale contextual information, which is then input into the feature encoding module at the current deep level for deep integration, so that deep semantics can guide the selection of shallow details, and shallow details can also provide accurate spatial positioning for deep semantics, thus generating final features with both high-level semantic discriminability and high-precision spatial details without any spatial downsampling, maintaining full resolution .

[0021] Let represent the initial feature encoding module, and respectively represent the feature encoding module of the i-th layer of the left path (main feature extraction path) and the right path (feature fusion path), wherein each feature encoding module is of the same network structure. represent an adapter layer (linear transformation of unified channel dimension and aligned features, realized by 1x1 convolution), and the entire feature extraction process can be accurately described by the following formula: 1. Initial feature extraction: ; 2. Main feature extraction path (left path): ; represent a feature encoding module series operator, and the feature encoding modules of the depth levels from =1 to are executed in turn; 3. Feature fusion path (right path): the feature fusion process of the path is calculated by the following depth levels, which is consistent with the network architecture diagram: ; ; ; The final feature is obtained by element-level addition of the feature maps output by the main feature extraction path and the feature fusion path at the last depth level (N=3 in this embodiment), is the final feature map, and is specifically represented as:

[0022] When processing real-time images of unmanned aerial vehicle law enforcement, unique challenges are faced due to high-altitude view, variable lighting and complex urban ground objects. The target (such as illegal construction, operating facilities occupying the road) presents great diversity in form, scale and direction. For example, the outline of the building can be a regular rectangle, or an irregular polygon; the key edge features (such as eaves, walls) in the aerial image can present an arbitrary angle. In addition, the target texture is often highly intertwined with the surrounding environment (such as vegetation, shadows, roads), which poses strict requirements on the robustness of feature extraction.

[0023] To address the above technical problems, the feature encoding module designed by the application is a scale-adaptive asymmetric convolution module (hereinafter referred to as SAACM). The core idea of this module is: through parallel asymmetric convolution paths with different receptive fields, the special perception of different scale and direction features is simulated, and a dynamic gating mechanism is introduced, so that the SAACM module can adaptively focus on the long-range dependency relationship with the most information quantity according to the internal correlation of the features, thereby providing high-quality and high-recognizability feature representation for subsequent unsupervised learning tasks.

[0024] The overall architecture of SAACM is shown in the figure, and its specific process is as follows: Assume that the input feature map of the module is ,in are the number of channels, height and width of the feature map respectively. As the primary feature encoding module, its input image is the image to be processed. The input image of other feature encoding modules is the feature image output by the previous depth level or the fused feature image. In order to realize residual learning and retain the original input information, the module introduces an identity path, which inputs the feature map Directly passed to the subsequent fusion stage, the input feature map or the image to be processed is the original input of the scale-adaptive asymmetric convolution module, and the output of the identity mapping can be expressed as , is the original input copy feature.

[0025] 2.1.1) Extracting shared features Input feature map First, a 5×5 depth wise-separable convolution (DWConv) is performed to perform preliminary feature extraction. This operation aims to effectively expand the receptive field at a lower computational cost and capture local spatial context information to generate a shared feature. .

[0026]

[0027] in, represents a 5×5 depthwise separable convolution, BN represents batch normalization, and RELU represents the activation function.

[0028] 2.1.2) Multi-scale asymmetric feature extraction In order to efficiently capture the different scale ranges and directional slender structural features, the present invention will share the features It is fed into three parallel asymmetric convolution branches. Each branch consists of a of and a of Series structure, is the size parameter of the convolution kernel in the depthwise separable convolution. This design is similar to the simulation While convolutionally expanding the receptive field, the number of parameters is significantly reduced and the sensitivity to horizontal and vertical structural features is enhanced.

[0029] Short-range feature branches: ; Mid-range feature branches: ; Long-range feature branch: ; Through these three paths, the module can simultaneously analyze information within the receptive field ranges of 7x7, 11x11 and 21x21, achieving parallel perception of multi-scale targets, for short-range features, for medium-range features, for long-range features.

[0030] 2.1.3) Feature gating and fusion To achieve dynamic interaction and information filtering between features of different scales, the present application designs a feature gating mechanism. First, the present application performs element-level multiplication on short-range features and medium-range features to capture the co-occurrence patterns of features at two scales. Then, a spatial Softmax function is used to convert the result into a normalized attention map whose value reflects the importance of features at different spatial positions.

[0031] ; wherein denotes element-level multiplication. The attention map is then used to weight long-range features so that the model can adaptively enhance or suppress long-distance context information, achieving dynamic filtering of information.

[0032] ; 2.1.4) Multi-path residual aggregation Finally, to ensure the richness of information flow and promote gradient propagation, the present application aggregates the gated features with two adjusted residual paths. The first path comes from the original input features of the identity path, and the second path comes from the ungated long-range features , which are respectively adjusted in channel and transformed in feature space through a 1x1 convolution to match the feature dimensions of the main feature extraction path.

[0033]

[0034] The final output features The core features filtered by the scale-adaptive gating, the original low-level information and the complete long-range context information are fused to form a comprehensive and robust feature description of the UAV image. Through its unique design, SAACM can efficiently deal with the challenge of variable target size and arbitrary direction in the UAV law enforcement scene, laying the foundation for precise unsupervised recognition tasks.

[0035] To achieve efficient feature extraction under the condition of limited computing power of edge computing devices, the present application discards the traditional single-path convolution stacking structure in network encoder design and proposes an unsupervised fusion adaptive cross-layer optimization network. The architecture is composed of a "scale-adaptive asymmetric convolution module" and a "two-way cross-layer feature extractor", aiming to capture rich multi-scale and multi-directional texture and structure information in real-time images with low computational overhead.

[0036] Unsupervised feature extraction and image segmentation form the core of the technical solution of the present application. An unsupervised fusion adaptive cross-layer optimization network based on an unsupervised light-weight segmentation model designed for UAV law enforcement real-time images is deployed on edge computing devices. Through unique structural design and self-supervised learning paradigm, the network can automatically learn the differences between law enforcement objects such as illegal buildings, roadside businesses and garbage dumps and the background environment without human annotation. To achieve unsupervised training, the network optimization process of the present application does not rely on human-annotated segmentation masks. The model learns pixel-level feature embedding space with high discriminability and local detail preservation by jointly optimizing two cooperative self-supervised learning tasks, namely intra-image contrastive loss function and local image block reconstruction loss function. This dual constraint mechanism enhances the semantic discrimination ability of the features and ensures the integrity of the visual details. Specifically, 2.2.1 Intra-Image Contrastive Loss The core idea of this loss function is to compare different image blocks (patches) within the same image to automatically cluster similar features and distinguish dissimilar features. The specific process is as follows: 1) Randomly sample anchor patch with its center located in the local region of pixel .

[0037] 2) Define positive patch : These image blocks are spatially adjacent to the anchor patch and have high similarity in visual attributes (such as structure and texture). Similarity is measured by indicators such as structural similarity index. For example, different regions of the same building roof can be positive samples for each other.

[0038] 3) Define image patches that are structurally or texturally significantly different from the anchor patch as negative sample patches (NegativePatches)

[0039] 4) The goal is to minimize the distance between the anchor patch and positive sample patches while maximizing the distance between the anchor patch and negative sample patches in the high-dimensional feature space.

[0040] The contrastive loss function is defined as follows:

[0041] where, denotes the feature embedding generated by the two-path cross-layer feature extractor, denotes the cosine similarity, and τ is the temperature hyperparameter. By optimizing this loss model, we can learn that semantically similar pixels (e.g., both belonging to illegal building areas) are close to each other in the feature space, while pixels with large semantic differences (e.g., illegal buildings and vegetation) are far apart.

[0042] 3.2.2 Local image patch reconstruction loss To prevent the model from losing local detail information during contrastive learning, a local reconstruction loss is introduced as a supplementary constraint. The basic idea is that the feature embedding must contain enough information to reconstruct the original image patch . The specific steps are as follows.

[0043] 1) Attach a lightweight reconstruction module after the encoder .

[0044] 2) This module takes the feature embedding of the anchor patch as input and outputs the reconstructed patch .

[0045] 3) The reconstruction loss is defined as the mean squared error (MSE) between the original image patch and the reconstructed patch .

[0046] By minimizing this loss, we can preserve key visual information such as texture, color, and structure in the original image while maintaining feature discriminability.

[0047] 3.2.3 Final objective function, i.e., unsupervised optimization objective function The final optimization objective of the network, i.e., the unsupervised optimization objective function, is composed of the intra-graph contrastive loss and the local reconstruction loss, and is adjusted by the balancing coefficient , where = 0.001: ; Through the unsupervised optimization objective function of joint optimization, the model can obtain high-quality pixel-level feature embedding under unsupervised conditions, thereby laying a solid foundation for subsequent implementation of real-time image region segmentation with clear boundaries and high precision.

[0048] The network structure of the application mainly receives input unmanned aerial vehicle images, and efficiently converts them into a high-dimensional and information-rich feature map through its network structure (scale adaptive asymmetric convolution module, double-path fusion), wherein each pixel position corresponds to a high-dimensional feature vector. The pros and cons of the network architecture design determine whether the feature can capture sufficient and useful information. However, in the absence of manual annotation, how to guide the extractor to learn 'high-quality' features that can both distinguish different ground objects and preserve fine boundaries is the core challenge of unsupervised methods. Therefore, instead of using traditional supervised learning, the application draws on and improves the ideas in the self-supervised learning field and designs a dual optimization objective consisting of an intra-graph contrast loss and a local image block reconstruction loss. The objective function serves as the fundamental basis for training and jointly constrains and optimizes the learning process of the feature extractor from two dimensions, enabling it to automatically learn high-quality feature representations suitable for urban law enforcement scenarios on unlabeled data.

[0049] The intra-graph contrast loss itself is a mature self-supervised learning technique. Generally, without labels, a positive and negative sample pair is constructed to enable the model to learn a meaningful feature representation space. The application has the following characteristics compared with it: 1. Traditional contrast learning often requires a huge external data set for pre-training. However, the 'intra-graph contrast' strategy of the application completely mines positive and negative samples within a single unmanned aerial vehicle image. This method greatly reduces the dependence on external data, enabling the model to perform highly adaptive feature learning for each real-time acquired image.

[0050] 2. Unlike contrast learning for classification (focusing on the features of the entire graph), the contrast in the application is performed between pixel-level image blocks (patches). This changes the learning goal from 'identifying what this graph is' to 'distinguishing what this small area in the graph belongs to'.

[0051] The local image block reconstruction loss is a classic component of self-supervised learning models, and its main role is to serve as an information constraint. By enabling the encoded features to be restored to the original input by the reconstruction module, it ensures that the encoded features do not lose too much original information. The application has the following characteristics compared with it: There is a potential risk of using contrastive learning alone, that the model may only learn the simplest features to distinguish positive and negative samples (e.g., only learn to distinguish "gray roof" and "green trees"), and ignore the tile texture of the roof, the leaf shape of the tree, and a large amount of detailed information.

[0052] By increasing the reconstruction loss, the present application forces the model to generate feature embeddings that are not only "discriminable" but also "invertible". The feature embeddings must retain enough rich local detail information to successfully reconstruct the original image block. Similarly: for reading comprehension, the contrastive loss requires "summarizing the main idea of the paragraph", while the reconstruction loss requires "repeating the key details of the original text".

[0053] In summary, the reconstruction loss and the contrastive loss are cleverly designed to form a complementary constraint system specifically for the unsupervised high-precision segmentation task. The contrastive loss is responsible for building macro semantic discrimination, while the reconstruction loss is responsible for preserving micro visual fidelity.

[0054] 3. Real-time segmentation and result output After the fusion of the adaptive cross-layer optimization network completes feature learning, the network already has strong image understanding capabilities. In order to effectively convert this capability into the final segmented image, the present application designs a lightweight segmentation head at the back end of the network, which is used to efficiently generate the segmentation result of the target region. After the feature extractor generates a high-dimensional feature map through encoding, it is passed to the segmentation head to complete the final pixel-level semantic decoding task.

[0055] The design goal of the segmentation head is to convert the abstract feature representation output by the feature extractor into a pixel-level segmentation map that can be directly understood by humans. In this scheme, the backbone network does not perform spatial downsampling during feature encoding, so the segmentation head does not need a complex upsampling module to restore the image resolution, thereby maintaining a more lightweight and efficient structure. The core modules of the segmentation head include the following: first is the feature refinement and channel dimension reduction module, which is composed of several lightweight convolutional layers (3x3 convolution), the main function of which is to refine the information and reduce the channel dimension of the input high-dimensional feature map, i.e., the final feature map (e.g., 2048 channels) while maintaining the original spatial resolution. Through this step, the feature map is mapped to a more compact and task-relevant feature space, providing more efficient feature representation for subsequent classification and decoding.

[0056] Next, the refined feature map is passed to a pixel-level classification layer, which is a 1x1 convolutional layer responsible for outputting C channels of prediction results at each pixel position, where C is the number of predefined semantic classes (such as buildings, roads, vegetation, background, etc.). The value of each channel represents the original score of the pixel belonging to the corresponding class. The output of the classification layer is processed by an activation function, usually using the Softmax activation function to normalize the C scores of each pixel pixel by pixel, converting them into a probability distribution. This process enables the model to ultimately generate a segmentation map in which each pixel is assigned a class label that is most likely to be the correct one, thereby completing accurate semantic segmentation, including the outline labeling and binary segmentation map of the target area, including illegal buildings, roadside business facilities, or garbage disposal areas.

[0057] In terms of real-time result synthesis and output, the segmentation map generated by model inference will be visualized in real time on the device end and immediately fed back to the ground control terminal. Specifically, the model will superimpose the final segmentation results in the form of outline labeling or pseudo-color segmentation map on the original video frames in real time, forming a dynamic labeled video stream. This video stream that integrates the segmentation results is immediately returned to the ground user terminal (such as a tablet or monitoring center screen) through the UAV's image transmission system. This "what you see is what you get" closed-loop system not only provides on-site law enforcement personnel with real-time structured understanding of the scene, but also provides strong visual support for accurate and efficient on-site decision-making.

[0058] The technical solution integrates a high-performance edge computing device OrangePi AIPro in the UAV, enabling law enforcement images to be preprocessed and unsupervised detection analyzed during the collection process without relying on cloud-based centralized processing, thereby significantly improving the real-time performance of illegal behavior monitoring and reducing the dependence on network environment, effectively avoiding delays and bandwidth occupation caused by image transmission. At the same time, the unsupervised feature learning method proposed by the present invention eliminates the dependence on large-scale manually labeled data, and can maintain strong transferability and adaptability under different building forms, complex lighting conditions, and occlusion environments; combined with the structural design of the scale-adaptive asymmetric convolution module and the dual-path cross-layer feature extractor, the model can fully fuse global semantic and local detail information, ensuring the boundary clarity and detection accuracy of the segmentation results, effectively overcoming the problems of boundary blurring and small target missed detection in the prior art.

[0059] Therefore, the application not only reduces the cost pressure of the system on bandwidth and computing power, but also improves the autonomy and intelligent level in the unmanned aerial vehicle cruise, so that the city law enforcement monitoring can realize the efficient closed loop of collecting, processing and feeding back at the same time, and has significant application value and social benefits. The technology is not only suitable for illegal building detection, but also can be extended to smart city scenarios such as occupation operation monitoring, garbage stacking identification, public facility inspection and emergency management, and provides efficient and reliable technical support for city governance.

[0060] Obviously, the above embodiments are only examples for clearly illustrating, but not limitation to the embodiments. For those skilled in the art, other different forms of changes or variations can be made on the basis of the above description. Here, all the embodiments need not and cannot be exhausted. The obvious changes or variations derived therefrom are still within the protection scope of the present application.

Claims

1. An unsupervised detection method for law enforcement images based on edge computing, characterized in that: The following steps are involved: 1) The collected image data is directly input into the edge computing device and pre-processed to obtain the image to be processed; 2) Inputting the image to be processed into a fused adaptive cross-layer optimization network for feature extraction and segmentation, the fused adaptive cross-layer optimization network comprising a dual-path cross-layer feature extractor and an unsupervised optimization objective function consisting of a weighted sum of an intra-image contrast loss function and a local reconstruction loss function. The dual-path cross-layer feature extractor comprises an initial feature encoding module, a backbone feature extraction path consisting of multiple identical feature encoding modules connected in series, and a feature fusion path that runs parallel to the backbone feature extraction path and performs cross-layer feature fusion. Among them, the feature encoding modules are all scale-adaptive asymmetric convolution modules, and the implementation of the scale-adaptive asymmetric convolution modules includes the following sub-steps: 2.1.1) Take the input image or feature map as the original input and perform preliminary feature extraction, then generate shared features after batch normalization and activation function processing. 2.1.2) The shared features are fed into three parallel asymmetric convolution branches with different receptive field scales to extract short-range features, medium-range features, and long-range features respectively. 2.1.3) Multiply the short-range features and the medium-range features element-wise to generate an attention map, and use the attention map to weight the long-range features to obtain the gated features. 2.1.4) Aggregate the gated features with the original input replica features and long-range features and output the feature map; 3) The edge computing device generates a segmentation result in real time based on the final output feature map.

2. The method according to claim 1, characterized in that In step 1), the edge computing device is an OrangePiAIPro device, and the collected image data does not need to rely on external network transmission and is directly connected to the OrangePiAIPro device.

3. The method according to claim 1, characterized in that In step 1), the preprocessing includes at least one of the following operations: size normalization, illumination and color correction, and noise filtering.

4. The method according to claim 1, wherein The three branches are short-range feature branch, medium-range feature branch and long-range feature branch, each of which consists of a A depthwise separable convolution and a The depth of separable convolution is connected in series to capture the features in different receptive fields. It is the size parameter of the convolution kernel in depthwise separable convolution.

5. The method according to claim 1, wherein The segmentation results are output to the user terminal in the form of contour annotation and binary segmentation map.

6. The method according to claim 1, characterized in that In step 3), the segmentation result includes the contour annotation and binary segmentation map of the target area, and the target area includes illegal buildings, road-blocking business facilities or garbage dumping areas.

7. The method according to claim 1, characterized in that In step 2), the specific implementation of the unsupervised optimization objective function includes: 2.2.1) Calculating the intra-image contrast loss: We randomly sample anchor blocks and define spatially adjacent and visually similar positive blocks and significantly different negative blocks. We optimize the loss by minimizing the feature distance between the anchor blocks and the positive blocks and maximizing the feature distance between the anchor blocks and the negative blocks. 2.2.2) Calculating the local image block reconstruction loss: The lightweight reconstruction module after the encoder uses the feature embedding of the anchor block as input to reconstruct the original image block. The mean squared error between the original image block and the reconstructed block is used as the loss. 2.2.3) The intra-image contrast loss and the local image block reconstruction loss are weighted and summed according to the preset weights as the final unsupervised optimization objective function.

Citation Information

Patent Citations

  • Intelligent inspection method, device, system and equipment for illegal building of urban house

    CN115984694A

  • Dynamic scene blurred image blind restoration method based on multi-stream attention adversarial network

    CN110969589A

  • Road surface crack detection method based on edge reconstruction network

    CN118096672A

  • Remote sensing image semantic segmentation method of asymmetric double-branch coding network based on clustering mutual contrast loss

    CN118864857A

  • Three-dimensional point-cloud semantic segmentation method based on multi-level boundary enhancement for unstructured environment

    WO2024230038A1