Forest illegal felling identification system based on computer vision

By using an improved dual-stream change detection network and a spatiotemporal compliance verification mechanism based on a graph database, the problems of inconsistent multi-source data and unstable detection in the identification of illegal logging in forest areas have been solved. This has enabled high-precision identification of illegal logging in forest areas with low false alarms, meeting the regulatory needs under complex terrain and meteorological conditions.

CN121459162APending Publication Date: 2026-02-03泰安市泰山风景名胜区管理委员会彩石溪管理区
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202511598297.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-04
Publication Date
2026-02-03

AI Technical Summary

Technical Problem

Existing technologies for identifying illegal logging in forest areas suffer from problems such as inconsistent spatiotemporal labeling of multi-source remote sensing data, large cross-sensor registration errors, unstable change detection, lack of cross-modal fusion, and incomplete evidence chains. These issues lead to large fluctuations in detection results, high false alarm rates, and difficulty in balancing timeliness and accuracy.

Method used

A spatiotemporal compliance verification mechanism employing an improved dual-stream change detection network, graph database, and rule engine generates a fine-grained target set through unified spatiotemporal identification, cross-sensor registration, semantic segmentation and instance segmentation, and conditional random field refinement. This allows for rapid inference at the edge and verification in the cloud, resulting in a high-precision, low-false-positive adjudication result.

Benefits of technology

It enables automatic detection and precise location of illegal logging in forest areas, and has the advantages of high precision, low false alarm, cross-modal robustness, rapid response and full-process traceability, meeting the needs of continuous supervision and law enforcement evidence collection under complex terrain and weather conditions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121459162A_ABST
    Figure CN121459162A_ABST
Patent Text Reader

Abstract

The invention discloses a forest illegal felling identification system based on computer vision, and the system comprises the following modules: a multi-source collection module which is used for collecting data and generating a unified space-time identification data set and an index; the correction registration module is used for correcting data, performing cross-sensor registration and unified resampling, and outputting an aligned image and a geo-fence index; the double-flow change module is used for pairing different time phase sub-graphs according to spatial-temporal indexes, constructing an improved double-flow network to generate a continuous change thermodynamic diagram and extracting candidate change blocks; the vector segmentation module is used for carrying out segmentation, refinement and vectorization under geo-fence constraints to obtain a fine-grained target set; the compliance checking module is used for combining the graph database and the rule engine to carry out space-time compliance checking; and the alarm judgment module is used for edge rapid reasoning and alarming, cloud rechecking and multi-source verification to form a judgment result set. The system gives consideration to accuracy and timeliness, and is suitable for continuous supervision and law enforcement evidence collection of forest regions under complex weather and topographic conditions.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of remote sensing image processing and intelligent monitoring technology, and in particular to a computer vision-based system for identifying illegal logging in forest areas. Background Technology

[0002] With the increasing demand for forest resource protection, carbon sequestration assessment, and law enforcement evidence collection, multi-source remote sensing monitoring and computer vision recognition in forest areas are becoming increasingly important. Existing technologies mainly rely on temporal difference or simple threshold segmentation of single optical images, or use SAR and LiDAR as supplementary evidence without achieving deep collaboration. They are easily affected by clouds, fog, shadows, and seasonal changes. The spatiotemporal annotation and quality of data from different platforms are inconsistent, the geometric, radiometric, and terrain correction links are incomplete, and the accumulation of errors in cross-sensor registration and unified resolution resampling leads to unstable comparison benchmarks and large fluctuations in detection results.

[0003] In downstream processing, change detection generally lacks optical-radar feature alignment and cross-modal fusion, resulting in numerous artifacts in the generated differential response. Target generation relies mainly on pixel-level masks, lacking semantic / instance joint segmentation and conditional random field refinement. Raster-to-vector conversion and topology repair are insufficient, and the boundaries of logging areas are fragmented and difficult to map. Furthermore, there is a lack of spatiotemporal compliance verification based on graph databases and rule engines, as well as a targeted supplementary inspection mechanism driven by the "minimum conflict set." Edge alerts and cloud-based verification are disconnected, the evidence chain is incomplete, and it is difficult to balance timeliness and accuracy.

[0004] Therefore, how to provide a computer vision-based system for identifying illegal logging in forest areas is a problem that urgently needs to be solved by those skilled in the art. Summary of the Invention

[0005] One objective of this invention is to propose a computer vision-based system for identifying illegal logging in forest areas. This invention comprehensively employs an improved dual-stream change detection network, spatiotemporal compliance verification based on graph databases and rule engines, and achieves automatic discovery, precise location, and evidence chain generation of suspicious logging activities through a collaborative mechanism of edge fast inference and cloud verification. It has the advantages of high accuracy, low false alarm rate, cross-modal robustness, fast response, strong domain adaptability, and full-process traceability.

[0006] A computer vision-based system for identifying illegal logging in forest areas according to an embodiment of the present invention includes the following modules: The multi-source acquisition module is used to acquire point clouds of visible light, infrared, multispectral, synthetic aperture radar and lidar, and preprocess them to form a raw multi-source dataset with unified spatiotemporal identification and a spatiotemporal identification index table. The calibration and registration module is used to correct the original multi-source dataset and perform cross-sensor registration and uniform resolution resampling, outputting an aligned image set and a geofence index. The dual-stream change module is used to pair aligned subgraphs of different time phases in the same region based on the spatiotemporal identifier index table, construct an improved dual-stream change detection network to perform feature alignment and cross-modal fusion, obtain a continuous change heatmap and extract a candidate change block set; The segmentation vector module is used to jointly perform semantic segmentation and instance segmentation under the constraints of the geofence index. It refines the boundaries using conditional random fields, performs raster-to-vector conversion and topology repair, and generates a fine-grained target set. The compliance verification module is used to complete spatiotemporal compliance verification based on a fine-grained target set and relying on graph database knowledge constraints and rule engine. It extracts the minimum conflict set for targets that are determined to be unstable and sends it back to trigger targeted supplementary inspection and time window expansion. It outputs a list of verified targets, a confidence threshold table and a domain adaptation model. The alarm adjudication module is used to take the target list and domain adaptation model as input, perform rapid reasoning on the edge side and trigger alarms, and complete the review and multi-source verification in the cloud to form an adjudication result set.

[0007] A method for identifying illegal logging in forest areas based on computer vision according to an embodiment of the present invention includes the following steps: Point clouds from visible light, infrared, multispectral, synthetic aperture radar, and lidar are collected and preprocessed to form a raw multi-source dataset with unified spatiotemporal labels and a spatiotemporal label index table. Geometric correction, radiometric correction and topographic correction are performed sequentially on the original multi-source dataset, and cross-sensor registration and uniform resolution resampling are implemented to output aligned image sets and geofence indexes; Using the aligned image set as input, and based on the spatiotemporal identifier index table, the aligned sub-images of different time phases in the same region are paired, and an improved dual-stream change detection network is constructed to perform feature alignment and cross-modal fusion, thereby obtaining a continuous change heatmap and extracting a candidate change block set; Based on the candidate change block set, semantic segmentation and instance segmentation are jointly performed under the constraints of geofence index. The boundary is refined using conditional random field, and raster to vector conversion and topology repair are performed to generate a fine-grained target set. Based on a fine-grained target set, spatiotemporal compliance verification is completed by relying on graph database knowledge constraints and rule engine. For targets that are determined to be unstable, the minimum conflict set is extracted and returned to trigger targeted supplementary inspection and time window expansion until the result is stable. The output includes a list of verification targets, a confidence threshold table and a domain adaptation model. Using the target list and domain adaptation model as input, the system performs rapid reasoning and triggers alarms at the edge, and completes review and multi-source verification in the cloud to form a set of adjudication results.

[0008] Optionally, the acquisition of visible light, infrared, multispectral, synthetic aperture radar, and lidar point clouds, and the preprocessing to form an original multi-source dataset with unified spatiotemporal identification and a spatiotemporal identification index table, specifically includes: Collect point cloud data from visible light, infrared, multispectral, synthetic aperture radar, and lidar to obtain the original set of entries and record imaging time and spatial positioning information; The original set of entries is unified with time reference and time synchronization is performed, and time alignment is completed according to the preset time window to obtain the time-aligned set of entries; Spatial reference normalization and projection unification are applied to the time-aligned item set, and regional gridding is performed to obtain the spatially aligned item set. Preprocessing is performed on the spatially aligned entry set by channel: For visible light, infrared, and multispectral images, cloud and fog detection and removal are performed by combining dark channel prior with brightness temperature thresholds and morphological opening and closing operations; shadow detection and compensation are performed by combining chromaticity saturation thresholds and shadow indices with neighborhood reflectance ratio recovery; and random noise suppression is performed by block matching 3D filtering and bilateral filtering. For synthetic aperture radar images, speckle noise suppression is performed by Lee adaptive filtering and nonlocal mean; amplitude normalization and phase unwrapping are performed for amplitude and phase consistency. For lidar point clouds, outlier removal is performed by statistical filtering and radius filtering; and initial ground point segmentation is performed by stepwise morphological filtering. The result is a preprocessed data set. The preprocessed data set is used to generate a unified spatiotemporal identifier according to a fixed coding rule. The region code, time code, sensor type, spatial resolution, observation geometry and quality level are written into the data to obtain a unified spatiotemporal identifier set and to establish a spatiotemporal identifier index table. The preprocessed data set is bound one-to-one with the unified spatiotemporal identifier set, and entries below the preset quality threshold are removed to form the original multi-source dataset with unified spatiotemporal identifiers. The spatiotemporal identifier index table is then archived as a version baseline.

[0009] Optionally, the step of sequentially performing geometric correction, radiometric correction, and topographic correction on the original multi-source dataset, and implementing cross-sensor registration and uniform resolution resampling to output an aligned image set and a geofence index specifically includes: The visible light, infrared and multispectral images in the original multi-source dataset are registered according to the proportional polynomial and orthorectified in combination with the digital elevation model. The synthetic aperture radar images are subjected to imaging geometric correction and terrain projection correction. The lidar point cloud is subjected to trajectory calculation, attitude correction and strip stitching, and the geometric correction data set is output. Perform radiometric correction, atmospheric correction and radar calibration on the geometric correction dataset: convert the digital values ​​of the optical channels into apparent reflectivity, perform atmospheric correction based on the atmospheric transmission model, calibrate the radar channel backscatter intensity to the terrain-normalized intensity, perform range compensation and incident angle compensation on the point cloud echo intensity, and output the radiometric correction dataset. Cross-sensor registration is performed on the radiometric correction dataset as the operation object: corner detection and edge extraction are used to construct a scale space descriptor, and initial transformation parameters are obtained by combining random consistency elimination. Pyramid hierarchical search and local affine fine-tuning are performed by using mutual information and phase correlation as joint measures. Fine registration is completed according to the region window defined by the spatiotemporal identifier index table, and the fine registration dataset is output. The finely registered dataset is resampled with high fidelity to a unified pixel grid, cropped according to the regional grid number to generate aligned sub-images, and then aggregated to output an aligned image set. By combining the aligned image set with the forest compartment map and administrative boundaries, coordinate consistency and topological restoration operations are performed to generate fence faces and unique numbers, establish a one-to-one mapping and spatial inclusion relationship between the fence and the aligned submap, and output the geofence index.

[0010] Optionally, the step of using an aligned image set as input, pairing aligned sub-images of different time phases within the same region according to a spatiotemporal identifier index table, constructing an improved dual-stream change detection network for feature alignment and cross-modal fusion, obtaining a continuous change heatmap, and extracting a candidate change block set, specifically includes: The aligned image set is retrieved, and aligned sub-images of different time phases are paired within the region and time window given by the spatiotemporal identification index table. Quality thresholds are set for cloud cover ratio, viewing angle difference, and registration residual. Only pairs that meet the thresholds are retained. For pairs that meet the thresholds, histogram matching and intensity standardization are performed on the optical channel to unify brightness and contrast. Amplitude normalization and phase quality check are performed on the radar channel to suppress differences in imaging power and interference quality. The pixel grid and pyramid scale are also unified to form standardized aligned sub-image pairs that can be directly compared. An improved dual-stream change detection network is constructed at the encoding end. The optical branch uses multi-scale residual blocks and deformable convolutions with channel attention superimposed, while the radar branch uses amplitude-phase decoupled convolutions and speckle robust residual blocks with channel attention superimposed. Cross-modal correlation fields are calculated on features of the same scale and cross-modal offset fields are regressed. Feature alignment is completed by sub-pixel resampling. At the same time, brightness and contrast are recalibrated for optical features and amplitude and phase are normalized for radar features. Alignment consistency score is output. Cross-modal attention and gated residual fusion are introduced at the bottleneck layer. Multi-scale features after two branches are aligned and superimposed are selected by scale-adaptive weight selection and superimposed. Pyramid aggregation compression is used to compress them into a fused semantic tensor. At the decoding end, spatial resolution is restored step by step through jump connection and upsampling. During the training phase, a hybrid loss of cross-entropy and overlap is used and hard sample focus terms and temporal consistency regularization are superimposed to stabilize learning. During the inference phase, sliding window overlapping inference and boundary smoothing are implemented to generate a continuously changing heat map. The pixel value of the continuously changing heat map represents the perturbation probability of the current position between time phases. Multi-threshold connected domain growth and scale space extremum detection are performed on continuously changing heatmaps. Candidate responses are extracted by combining morphological refinement and minimum area constraints. Adaptive thresholds are set based on alignment consistency scores and cross-modal consistency scores to eliminate unstable responses. Geofencing indexes are used to delete responses outside the fence. Finally, a set of candidate change blocks is output, and the area, aspect ratio, boundary curvature, and confidence score are recorded for each candidate change block.

[0011] Optionally, the step of jointly performing semantic segmentation and instance segmentation around the candidate change block set under the constraints of the geofence index, refining the boundaries using a conditional random field, and performing raster-to-vector conversion and topology repair to generate a fine-grained target set, specifically includes: The target region is generated around the candidate change block set under the constraints of the geofence index. The cropping is completed by polygon overlay and difference operation. The region is expanded outward by a fixed buffer distance and outbound cells are removed. Semantic segmentation and instance segmentation are jointly implemented in the target region. The semantic segmentation adopts an encoder-decoder network and a multi-scale pyramid, with cross-entropy superimposed overlap loss. The instance segmentation adopts candidate block-guided mask region proposal and pixel-level mask regression, and uses boundary-aware loss to improve edge separability. The semantic segmentation and instance segmentation results are analyzed and refined. First, the maximum a posteriori selection is performed based on the category probability and mask overlap. Then, the nearest neighbors of the same category are merged based on the minimum gap threshold. A conditional random field is used to minimize the energy by using the category probability as the unary potential and the color difference, gradient difference and spatial distance as the binary potential. Finally, an insurmountable constraint is applied to the fence boundary to obtain the refined object mask set. The refined object mask set is deleted by removing the mask connected components according to the preset area threshold, and opening and closing operations are performed for smoothing, hole filling and narrow channel disconnection. Based on contour tracking, the raster to vector conversion is completed, vertex simplification and polyline smoothing are performed, and vertex snapping and distance limit control are performed on the fence boundary of the preset range to obtain the vector feature set. Perform topology repair and target generation on the vector feature set, including: detecting and resolving self-intersections and overlaps, correcting hole orientations, stitching across tiles along shared boundaries, merging or deleting narrow patches according to area and aspect ratio thresholds, outputting a fine-grained target set, and recording the category, instance number, area, perimeter, aspect ratio, boundary curvature, relative position to the fence, and overall confidence level for each fine-grained target.

[0012] Optionally, the step of performing spatiotemporal compliance verification based on a fine-grained target set, relying on graph database knowledge constraints and a rule engine, extracting the minimum conflict set for targets determined to be unstable and sending it back to trigger targeted supplementary checks and time window expansion until the results stabilize, and outputting a list of verified targets, a confidence threshold table, and a domain adaptation model, specifically including: External information is collected object by object based on a fine-grained target set, including spatial overlay with forest compartments and permitted areas, recording forest compartment number, permit number and whether it is within the permit, constructing fixed buffers for roads and water bodies and calculating minimum distance and intersection relationship, performing regional statistics on slope grid to obtain mean, maximum value and over-threshold ratio, associating historical operation records with candidate change blocks according to time windows, determining whether the time window and cooling period are satisfied, and summarizing to form a knowledge constraint set and a check item table; Perform spatiotemporal compliance verification according to the verification checklist, including applying threshold rules to slope, using buffer overlay for roads and water bodies, using spatial overlay for permitted areas, using time windows and cooling-off periods to determine historical operations, outputting a set of compliant targets and a set of suspected non-compliant targets, and generating verification evidence records; Stability indicators are calculated for suspected non-compliant sets, including cross-temporal response consistency, overlap with geofence, evidence path integrity, and source freshness. Unstable targets are marked, an unstable label table is generated, and stable suspected sets and stability metrics are obtained. Extract the minimum conflict set from the unstable marker table, remove redundancies from the rules that trigger failure, retain the minimum entry group that can cause failure on its own, issue a feedback instruction based on the minimum conflict set, implement targeted supplementary inspection and time window expansion, add aligned subgraphs and neighborhood samples, recalculate the continuously changing heatmap and candidate change block set, and incrementally update the fine-grained target set. Threshold adaptation and domain adaptation are implemented for the compliant target set and the stable suspected set. A confidence threshold table is generated according to the confidence quantile and the target false alarm rate. Pseudo-label semi-supervised training and adaptive normalization are used for small incremental training to obtain the domain-adapted model.

[0013] Optionally, the step of taking the target list and domain adaptation model as input, performing rapid inference and triggering alarms at the edge, and completing review and multi-source verification in the cloud to form a set of adjudication results specifically includes: Read the list of verification targets and the domain adaptation model, load a lightweight engine at the edge to perform batch fast inference, and output the edge alarm candidate set and the corresponding confidence and time stamp; Set thresholds for filtering based on geofencing and alignment consistency, generate a list of on-site alarms, and record the device identification and triggering reason; The on-site alarm list is sent back to the cloud, and the continuous change heat map and candidate change blocks are recalculated based on the aligned image set and fine-grained target set to form the cloud review result; A second spatiotemporal compliance check is performed by combining historical operations, permitted scope, road and water body buffer zones, and slope, and a review opinion and evidence snapshot are output. Consistency determination and conflict resolution are performed on the on-site alarm list and the cloud review results. The decision confidence and reason code are calculated, and the decision result set is generated. The decision result set is bound to the spatiotemporal identifier and written to the regulatory interface and log.

[0014] The beneficial effects of this invention are: This invention significantly improves the consistency and usability of multi-source data by introducing unified spatiotemporal identification, high-precision cross-sensor registration, and unified resolution resampling into the end-to-end link of "multi-source acquisition—correction and registration—dual-stream change—segmentation vector—compliance verification—alarm adjudication." Based on an improved dual-stream change detection network, it achieves optical and radar feature alignment and cross-modal fusion, effectively suppressing false differences caused by clouds, shadows, and seasonal variations, thus improving the stability and recall rate of continuously changing heatmaps. Under geofencing constraints, it jointly performs semantic segmentation and instance segmentation, and refines boundaries using conditional random fields, completes raster-to-vector conversion and topology repair, resulting in fine-grained target data that can be directly mapped and statistically analyzed. The standard set makes the logging area boundaries more complete and geometrically consistent; relying on graph databases and rule engines to carry out spatiotemporal compliance verification, and combining knowledge such as permit scope, time window, road / water body buffer and slope to generate a traceable evidence chain, extract the minimum conflict set for unstable targets and trigger targeted supplementary inspection and time window expansion, so that the conclusions are self-consistent and convergent and the false alarm rate is reduced; finally, the verification target list and domain adaptation model are used to quickly reason on the edge side and trigger alarms, and the cloud review and multi-source verification form an adjudication result set, realizing the edge-cloud collaboration of "low-latency early warning + high-reliability adjudication", taking into account timeliness, accuracy and interpretability, and meeting the needs of continuous supervision and law enforcement evidence collection under complex terrain and weather conditions. Attached Figure Description

[0015] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are used in conjunction with embodiments of the invention to explain the invention and do not constitute a limitation thereof. In the drawings:

[0016] Figure 1 This is a framework diagram of a computer vision-based forest illegal logging identification system proposed in this invention. Figure 2 This is a schematic diagram of a computer vision-based method for identifying illegal logging in forest areas proposed in this invention. Figure 3 This is a framework diagram of the improved dual-flow change detection network in a computer vision-based method for identifying illegal logging in forest areas proposed in this invention. Detailed Implementation

[0017] The present invention will now be described in further detail with reference to the accompanying drawings. These drawings are simplified schematic diagrams, illustrating only the basic structure of the invention, and therefore only show the components relevant to the invention.

[0018] refer to Figure 1 A computer vision-based system for identifying illegal logging in forest areas includes the following modules: The multi-source acquisition module is used to acquire point clouds of visible light, infrared, multispectral, synthetic aperture radar and lidar, and preprocess them to form a raw multi-source dataset with unified spatiotemporal identification and a spatiotemporal identification index table. The calibration and registration module is used to correct the original multi-source dataset and perform cross-sensor registration and uniform resolution resampling, outputting an aligned image set and a geofence index. The dual-stream change module is used to pair aligned subgraphs of different time phases in the same region based on the spatiotemporal identifier index table, construct an improved dual-stream change detection network to perform feature alignment and cross-modal fusion, obtain a continuous change heatmap and extract a candidate change block set; The segmentation vector module is used to jointly perform semantic segmentation and instance segmentation under the constraints of the geofence index. It refines the boundaries using conditional random fields, performs raster-to-vector conversion and topology repair, and generates a fine-grained target set. The compliance verification module is used to complete spatiotemporal compliance verification based on a fine-grained target set and relying on graph database knowledge constraints and rule engine. It extracts the minimum conflict set for targets that are determined to be unstable and sends it back to trigger targeted supplementary inspection and time window expansion. It outputs a list of verified targets, a confidence threshold table and a domain adaptation model. The alarm adjudication module is used to take the target list and domain adaptation model as input, perform rapid reasoning on the edge side and trigger alarms, and complete the review and multi-source verification in the cloud to form an adjudication result set.

[0019] refer to Figure 2-3 A computer vision-based method for identifying illegal logging in forest areas includes the following steps: Point clouds from visible light, infrared, multispectral, synthetic aperture radar, and lidar are collected and preprocessed to form a raw multi-source dataset with unified spatiotemporal labels and a spatiotemporal label index table. Geometric correction, radiometric correction and topographic correction are performed sequentially on the original multi-source dataset, and cross-sensor registration and uniform resolution resampling are implemented to output aligned image sets and geofence indexes; Using the aligned image set as input, and based on the spatiotemporal identifier index table, the aligned sub-images of different time phases in the same region are paired, and an improved dual-stream change detection network is constructed to perform feature alignment and cross-modal fusion, thereby obtaining a continuous change heatmap and extracting a candidate change block set; Based on the candidate change block set, semantic segmentation and instance segmentation are jointly performed under the constraints of geofence index. The boundary is refined using conditional random field, and raster to vector conversion and topology repair are performed to generate a fine-grained target set. Based on a fine-grained target set, spatiotemporal compliance verification is completed by relying on graph database knowledge constraints and rule engine. For targets that are determined to be unstable, the minimum conflict set is extracted and returned to trigger targeted supplementary inspection and time window expansion until the result is stable. The output includes a list of verification targets, a confidence threshold table and a domain adaptation model. Using the target list and domain adaptation model as input, the system performs rapid reasoning and triggers alarms at the edge, and completes review and multi-source verification in the cloud to form a set of adjudication results.

[0020] In this embodiment, the acquisition of visible light, infrared, multispectral, synthetic aperture radar, and lidar point clouds, and the preprocessing to form a raw multi-source dataset with a unified spatiotemporal identifier and a spatiotemporal identifier index table, specifically includes: Collect point cloud data from visible light, infrared, multispectral, synthetic aperture radar, and lidar to obtain the original set of entries and record imaging time and spatial positioning information; The original set of entries is unified with time reference and time synchronization is performed, and time alignment is completed according to the preset time window to obtain the time-aligned set of entries; Spatial reference normalization and projection unification are applied to the time-aligned item set, and regional gridding is performed to obtain the spatially aligned item set. Preprocessing is performed on the spatially aligned entry set by channel: For visible light, infrared, and multispectral images, cloud and fog detection and removal are performed by combining dark channel prior with brightness temperature thresholds and morphological opening and closing operations; shadow detection and compensation are performed by combining chromaticity saturation thresholds and shadow indices with neighborhood reflectance ratio recovery; and random noise suppression is performed by block matching 3D filtering and bilateral filtering. For synthetic aperture radar images, speckle noise suppression is performed by Lee adaptive filtering and nonlocal mean; amplitude normalization and phase unwrapping are performed for amplitude and phase consistency. For lidar point clouds, outlier removal is performed by statistical filtering and radius filtering; and initial ground point segmentation is performed by stepwise morphological filtering. The result is a preprocessed data set. The preprocessed data set is used to generate a unified spatiotemporal identifier according to a fixed coding rule. The region code, time code, sensor type, spatial resolution, observation geometry and quality level are written into the data to obtain a unified spatiotemporal identifier set and to establish a spatiotemporal identifier index table. The preprocessed data set is bound one-to-one with the unified spatiotemporal identifier set, and entries below the preset quality threshold are removed to form the original multi-source dataset with unified spatiotemporal identifiers. The spatiotemporal identifier index table is then archived as a version baseline.

[0021] This implementation method significantly reduces cloud and fog noise and multi-source differences by unifying the spatiotemporal identification, quality preprocessing and coding index of multi-source data, thereby improving the efficiency of subsequent registration and retrieval, and ensuring data traceability and version consistency.

[0022] In this embodiment, the step of sequentially performing geometric correction, radiometric correction, and topographic correction on the original multi-source dataset, and implementing cross-sensor registration and uniform resolution resampling to output an aligned image set and geofence index specifically includes: The visible light, infrared and multispectral images in the original multi-source dataset are registered according to the proportional polynomial and orthorectified in combination with the digital elevation model. The synthetic aperture radar images are subjected to imaging geometric correction and terrain projection correction. The lidar point cloud is subjected to trajectory calculation, attitude correction and strip stitching, and the geometric correction data set is output. Perform radiometric correction, atmospheric correction and radar calibration on the geometric correction dataset: convert the digital values ​​of the optical channels into apparent reflectivity, perform atmospheric correction based on the atmospheric transmission model, calibrate the radar channel backscatter intensity to the terrain-normalized intensity, perform range compensation and incident angle compensation on the point cloud echo intensity, and output the radiometric correction dataset. Cross-sensor registration is performed on the radiometric correction dataset as the operation object: corner detection and edge extraction are used to construct a scale space descriptor, and initial transformation parameters are obtained by combining random consistency elimination. Pyramid hierarchical search and local affine fine-tuning are performed by using mutual information and phase correlation as joint measures. Fine registration is completed according to the region window defined by the spatiotemporal identifier index table, and the fine registration dataset is output. The finely registered dataset is resampled with high fidelity to a unified pixel grid, cropped according to the regional grid number to generate aligned sub-images, and then aggregated to output an aligned image set. By combining the aligned image set with the forest compartment map and administrative boundaries, coordinate consistency and topological restoration operations are performed to generate fence faces and unique numbers, establish a one-to-one mapping and spatial inclusion relationship between the fence and the aligned submap, and output the geofence index.

[0023] This implementation method obtains grid-aligned images and geofence indexes by combining geometric, radiometric, and topographic corrections with multi-scale fine registration of pyramids and unified resampling, thereby reducing misalignment artifacts and improving the accuracy and stability of change detection.

[0024] In this embodiment, the step of using an aligned image set as input, pairing aligned sub-images of different time phases within the same region according to a spatiotemporal identifier index table, constructing an improved dual-stream change detection network for feature alignment and cross-modal fusion to obtain a continuous change heatmap and extract a candidate change block set, specifically includes: The aligned image set is retrieved, and aligned sub-images of different time phases are paired within the region and time window given by the spatiotemporal identification index table. Quality thresholds are set for cloud cover ratio, viewing angle difference, and registration residual. Only pairs that meet the thresholds are retained. For pairs that meet the thresholds, histogram matching and intensity standardization are performed on the optical channel to unify brightness and contrast. Amplitude normalization and phase quality check are performed on the radar channel to suppress differences in imaging power and interference quality. The pixel grid and pyramid scale are also unified to form standardized aligned sub-image pairs that can be directly compared. The encoding end of the improved dual-stream change detection network is constructed. The optical branch uses multi-scale residual blocks and deformable convolution with channel attention to enhance texture and boundary representation. The radar branch uses amplitude-phase decoupled convolution and speckle robust residual blocks with channel attention to improve robustness to speckles and geometric distortions. Cross-modal correlation fields are calculated on the same-scale features and cross-modal offset fields are regressed. Feature alignment is completed by sub-pixel resampling. At the same time, brightness and contrast are recalibrated for optical features and amplitude and phase are normalized for radar features. Alignment consistency score is output. Cross-modal attention and gated residual fusion are introduced at the bottleneck layer. Multi-scale features after two branches are aligned and superimposed are selected by scale-adaptive weight selection and superimposed. Pyramid aggregation compression is used to compress them into a fused semantic tensor. At the decoding end, spatial resolution is restored step by step through jump connection and upsampling. During the training phase, a hybrid loss of cross-entropy and overlap is used and hard sample focus terms and temporal consistency regularization are superimposed to stabilize learning. During the inference phase, sliding window overlapping inference and boundary smoothing are implemented to generate a continuously changing heat map. The pixel value of the continuously changing heat map represents the perturbation probability of the current position between time phases. Multi-threshold connected domain growth and scale space extremum detection are performed on continuously changing heatmaps. Candidate responses are extracted by combining morphological refinement and minimum area constraints. Adaptive thresholds are set based on alignment consistency scores and cross-modal consistency scores to eliminate unstable responses. Geofencing indexes are used to delete responses outside the fence. Finally, a set of candidate change blocks is output, and the area, aspect ratio, boundary curvature, and confidence score are recorded for each candidate change block.

[0025] This implementation constructs an improved dual-stream change detection network by pairing aligned subgraphs from different time phases within the same region and performing cross-modal feature alignment, attention fusion, and pyramid decoding to generate a continuous change heatmap and extract candidate change blocks. This significantly improves the detection rate of weak textures, small-scale burr marks, and occluded targets, reduces artifacts caused by viewpoint differences, speckle noise, and brightness inconsistencies, and ensures the stability and traceability of change response.

[0026] In this embodiment, the step of jointly performing semantic segmentation and instance segmentation around the candidate change block set under the constraints of the geofence index, refining the boundaries using a conditional random field, and performing raster-to-vector conversion and topology repair to generate a fine-grained target set specifically includes: The target region is generated around the candidate change block set under the constraints of the geofence index. The cropping is completed by polygon overlay and difference operation. The region is expanded outward by a fixed buffer distance and outbound cells are removed. Semantic segmentation and instance segmentation are jointly implemented in the target region. The semantic segmentation adopts an encoder-decoder network and a multi-scale pyramid, with cross-entropy superimposed overlap loss. The instance segmentation adopts candidate block-guided mask region proposal and pixel-level mask regression, and uses boundary-aware loss to improve edge separability. The semantic segmentation and instance segmentation results are analyzed and refined. First, the maximum a posteriori selection is performed based on the category probability and mask overlap. Then, the nearest neighbors of the same category are merged based on the minimum gap threshold. A conditional random field is used to minimize the energy by using the category probability as the unary potential and the color difference, gradient difference and spatial distance as the binary potential. Finally, an insurmountable constraint is applied to the fence boundary to obtain the refined object mask set. The refined object mask set is deleted by removing the mask connected components according to the preset area threshold, and opening and closing operations are performed for smoothing, hole filling and narrow channel disconnection. Based on contour tracking, the raster to vector conversion is completed, vertex simplification and polyline smoothing are performed, and vertex snapping and distance limit control are performed on the fence boundary of the preset range to obtain the vector feature set. Perform topology repair and target generation on the vector feature set, including: detecting and resolving self-intersections and overlaps, correcting hole orientations, stitching across tiles along shared boundaries, merging or deleting narrow patches according to area and aspect ratio thresholds, outputting a fine-grained target set, and recording the category, instance number, area, perimeter, aspect ratio, boundary curvature, relative position to the fence, and overall confidence level for each fine-grained target.

[0027] This implementation combines semantic segmentation and instance segmentation under the constraints of geofence indexing, and uses conditional random field edge refinement, raster to vector conversion and topology repair to generate a fine-grained target set with accurate boundaries and complete structure. This process transforms the coarse-grained response of candidate blocks into measurable object elements, which facilitates area and shape statistics, cross-map splicing and boundary consistency verification, thereby providing high-quality geometric evidence for subsequent compliance judgment and law enforcement evidence collection.

[0028] In this embodiment, the spatiotemporal compliance verification based on a fine-grained target set, relying on graph database knowledge constraints and a rule engine, is completed. For targets deemed unstable, a minimum conflict set is extracted and fed back to trigger targeted supplementary checks and time window expansion until the results stabilize. The output includes a list of verified targets, a confidence threshold table, and a domain adaptation model. Specifically, this includes: External information is collected object by object based on a fine-grained target set, including spatial overlay with forest compartments and permitted areas, recording forest compartment number, permit number and whether it is within the permit, constructing fixed buffers for roads and water bodies and calculating minimum distance and intersection relationship, performing regional statistics on slope grid to obtain mean, maximum value and over-threshold ratio, associating historical operation records with candidate change blocks according to time windows, determining whether the time window and cooling period are satisfied, and summarizing to form a knowledge constraint set and a check item table; Perform spatiotemporal compliance verification according to the verification checklist, including applying threshold rules to slope, using buffer overlay for roads and water bodies, using spatial overlay for permitted areas, using time windows and cooling-off periods to determine historical operations, outputting a set of compliant targets and a set of suspected non-compliant targets, and generating verification evidence records; Stability indicators are calculated for suspected non-compliant sets, including cross-temporal response consistency, overlap with geofence, evidence path integrity, and source freshness. Unstable targets are marked, an unstable label table is generated, and stable suspected sets and stability metrics are obtained. Extract the minimum conflict set from the unstable marker table, remove redundancies from the rules that trigger failure, retain the minimum entry group that can cause failure on its own, issue a feedback instruction based on the minimum conflict set, implement targeted supplementary inspection and time window expansion, add aligned subgraphs and neighborhood samples, recalculate the continuously changing heatmap and candidate change block set, and incrementally update the fine-grained target set. Threshold adaptation and domain adaptation are implemented for the compliant target set and the stable suspected set. A confidence threshold table is generated according to the confidence quantile and the target false alarm rate. Pseudo-label semi-supervised training and adaptive normalization are used for small incremental training to obtain the domain-adapted model.

[0029] This implementation method constructs knowledge constraints by collecting forest compartments, permits, roads, water bodies, slopes, and historical operation records around a fine-grained target set, and performs spatiotemporal compliance verification by a rule engine. It extracts the minimum conflict set for unstable samples and triggers targeted supplementary inspections and time window expansion. At the same time, it achieves domain adaptation by using quantile adaptive thresholds and pseudo-label semi-supervision. Finally, it outputs a list of verification targets and a threshold table, which significantly suppresses false alarms and false negatives and improves the generalization ability across forest types and seasonal phases.

[0030] In this embodiment, the process of taking the target verification list and domain adaptation model as input, performing rapid inference and triggering alarms at the edge, and completing review and multi-source verification in the cloud to form a set of adjudication results specifically includes: Read the list of verification targets and the domain adaptation model, load a lightweight engine at the edge to perform batch fast inference, and output the edge alarm candidate set and the corresponding confidence and time stamp; Set thresholds for filtering based on geofencing and alignment consistency, generate a list of on-site alarms, and record the device identification and triggering reason; The on-site alarm list is sent back to the cloud, and the continuous change heat map and candidate change blocks are recalculated based on the aligned image set and fine-grained target set to form the cloud review result; A second spatiotemporal compliance check is performed by combining historical operations, permitted scope, road and water body buffer zones, and slope, and a review opinion and evidence snapshot are output. Consistency determination and conflict resolution are performed on the on-site alarm list and the cloud review results. The decision confidence and reason code are calculated, and the decision result set is generated. The decision result set is bound to the spatiotemporal identifier and written to the regulatory interface and log.

[0031] This implementation method uses edge-side rapid inference to link cloud-based review and multi-source verification to form a decision result with confidence level and evidence snapshot, which shortens alarm response time and ensures the reliability of the final conclusion and full traceability of the process.

[0032] Example 1: To verify the feasibility of this invention in practice, it was applied to a key forest area characterized by alternating mountainous and valley terrain. The experimental area included mixed management units of state-owned and collectively owned forests, with a continuous monitoring area of ​​approximately 1,200 square kilometers. The experimental period covered the growing season from late spring to early autumn, a region with undulating terrain, frequent cloud cover, and historically, sporadic illegal logging and small-scale clearing along forest roads.

[0033] Data sources include airborne and ground-based lidar point clouds, all-weather synthetic aperture radar imagery, and visible and near-infrared multispectral imagery. All platforms employ unified time coding and regional grids for data acquisition. An edge node cluster consisting of forest ranger stations, lookout towers, and forest road checkpoints is deployed on-site, while a unified graph database and rule engine are deployed in the cloud, storing knowledge entries such as forest compartment boundaries, historical work permits, road and water buffer zones, slope, and topographic indices. After acquisition, the system automatically performs geometric, radiometric, and topographic corrections, performs cross-sensor registration and unified resolution resampling, and generates aligned image sets and geofence indexes. Subsequently, sub-maps of different time phases within the same region are paired according to spatiotemporal identifiers and fed into an improved dual-stream change detection network. The network aligns and fuses optical details and radar structural information at the feature level, outputting a continuously changing heatmap. Around high-response areas, the system combines semantic segmentation and instance segmentation under geofence constraints, refines boundaries using conditional random fields, completes raster-to-vector conversion and topology repair, and obtains fine-grained targets that can be directly mapped and statistically analyzed. All targets are sent to the graph database for spatiotemporal compliance verification. If the evidence is insufficient or there are conflicts, the system automatically extracts the minimum conflict set and triggers targeted supplementary inspection and time window expansion. After stabilization, a list of verification targets and a domain adaptation model are generated and sent to the edge for continuous inference and alarms. The cloud performs multi-source verification and forms a decision result, generating a traceable report composed of location, time, rule matching and evidence snapshot.

[0034] The experimental area used continuous monitoring data from the past three months, which was compared with two baseline methods: one was a combination of single optical temporal difference and threshold segmentation, and the other was change detection using a single-branch convolutional network. The evaluation focused on target-level detection and boundary representation, as well as the availability and processing timeliness of the evidence chain for law enforcement. Table 1 presents the comprehensive results of the core indicators. Cloud cover was statistically analyzed in stratified categories of less than 20% and greater than 60%. Area and boundary-related indicators were calculated only within the geofence, and latency was statistically analyzed per standardized submap pair.

[0035] Table 1. Summary and Comprehensive Evaluation of the Experimental Zone over Three Months

[0036] As shown in Table 1, the dual-stream change detection integrating optics and radar significantly reduced false differences, with the target-level F1 reaching 0.92 under clear sky conditions and remaining at 0.89 under cloudy conditions. Boundary representation, through joint segmentation, CRF refinement, and vector topology repair, improved the median IoU to 0.83, reduced the maximum boundary deviation from over 20 meters to the order of 10 meters, and significantly narrowed the area error of the logging area. Improved alignment consistency scoring ensured comparability between multi-source temporal phases. Spatiotemporal compliance verification based on graph databases and rule engines increased the compliance pass rate to over 80%, while the convergence rate of disputed samples approached 80% through the use of minimum conflict sets and targeted supplementary inspection strategies. Edge-cloud collaboration improved processing time, reducing edge alarm latency to the order of 10 seconds and the average cloud adjudication time to less than 10 minutes, facilitating rapid intervention and evidence collection for suspicious operations. In summary, this invention can still stably provide target-level, mappable, and traceable identification and adjudication results even in complex terrain and high cloud cover environments, supporting continuous monitoring and law enforcement applications in forest areas with higher accuracy and lower latency.

[0037] The above description is only a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any equivalent substitutions or modifications made by those skilled in the art within the scope of the technology disclosed in the present invention, based on the technical solution and inventive concept of the present invention, should be covered within the scope of protection of the present invention.

Claims

1. A computer vision-based system for identifying illegal logging in forest areas, characterized in that, Includes the following modules: The multi-source acquisition module is used to acquire point clouds of visible light, infrared, multispectral, synthetic aperture radar and lidar, and preprocess them to form a raw multi-source dataset with unified spatiotemporal identification and a spatiotemporal identification index table. The calibration and registration module is used to correct the original multi-source dataset and perform cross-sensor registration and uniform resolution resampling, outputting an aligned image set and a geofence index. The dual-stream change module is used to pair aligned subgraphs of different time phases in the same region based on the spatiotemporal identifier index table, construct an improved dual-stream change detection network to perform feature alignment and cross-modal fusion, obtain a continuous change heatmap and extract a candidate change block set; The segmentation vector module is used to jointly perform semantic segmentation and instance segmentation under the constraints of the geofence index. It refines the boundaries using conditional random fields, performs raster-to-vector conversion and topology repair, and generates a fine-grained target set. The compliance verification module is used to complete spatiotemporal compliance verification based on a fine-grained target set and relying on graph database knowledge constraints and rule engine. It extracts the minimum conflict set for targets that are determined to be unstable and sends it back to trigger targeted supplementary inspection and time window expansion. It outputs a list of verified targets, a confidence threshold table and a domain adaptation model. The alarm adjudication module is used to take the target list and domain adaptation model as input, perform rapid reasoning on the edge side and trigger alarms, and complete the review and multi-source verification in the cloud to form an adjudication result set.

2. The computer vision-based illegal logging identification system for forest areas according to claim 1, characterized in that, The modules are connected in the following way: Point clouds from visible light, infrared, multispectral, synthetic aperture radar, and lidar are collected and preprocessed to form a raw multi-source dataset with unified spatiotemporal identifiers and a spatiotemporal identifier index table. Geometric correction, radiometric correction, and topographic correction are performed sequentially on the original multi-source dataset, and cross-sensor registration and uniform resolution resampling are implemented to output aligned image sets and geofence indexes. Using the aligned image set as input, and based on the spatiotemporal identifier index table, the aligned sub-images of different time phases in the same region are paired, and an improved dual-stream change detection network is constructed to perform feature alignment and cross-modal fusion, thereby obtaining a continuous change heatmap and extracting a candidate change block set; Based on the candidate change block set, semantic segmentation and instance segmentation are jointly performed under the constraints of geofence index. The boundary is refined using conditional random field, and raster to vector conversion and topology repair are performed to generate a fine-grained target set. Based on a fine-grained target set, spatiotemporal compliance verification is completed by relying on graph database knowledge constraints and rule engine. For targets that are determined to be unstable, the minimum conflict set is extracted and returned to trigger targeted supplementary inspection and time window expansion until the result is stable. The output includes a list of verification targets, a confidence threshold table and a domain adaptation model. Using the target list and domain adaptation model as input, the system performs rapid reasoning and triggers alarms at the edge, and completes review and multi-source verification in the cloud to form a set of adjudication results.

3. The computer vision-based illegal logging identification system for forest areas according to claim 1, characterized in that, The acquisition of visible light, infrared, multispectral, synthetic aperture radar, and lidar point clouds, followed by preprocessing to form a raw multi-source dataset with unified spatiotemporal identification and a spatiotemporal identification index table, specifically includes: Collect point cloud data from visible light, infrared, multispectral, synthetic aperture radar, and lidar to obtain the original set of entries and record imaging time and spatial positioning information; The original set of entries is unified with time reference and time synchronization is performed, and time alignment is completed according to the preset time window to obtain the time-aligned set of entries; Spatial reference normalization and projection unification are applied to the time-aligned item set, and regional gridding is performed to obtain the spatially aligned item set. Preprocessing is performed on the spatially aligned entry set by channel: For visible light, infrared, and multispectral images, cloud and fog detection and removal are performed by combining dark channel prior with brightness temperature thresholds and morphological opening and closing operations; shadow detection and compensation are performed by combining chromaticity saturation thresholds and shadow indices with neighborhood reflectance ratio recovery; and random noise suppression is performed by block matching 3D filtering and bilateral filtering. For synthetic aperture radar images, speckle noise suppression is performed by Lee adaptive filtering and nonlocal mean; amplitude normalization and phase unwrapping are performed for amplitude and phase consistency. For lidar point clouds, outlier removal is performed by statistical filtering and radius filtering; and initial ground point segmentation is performed by stepwise morphological filtering. The result is a preprocessed data set. The preprocessed data set is used to generate a unified spatiotemporal identifier according to a fixed coding rule. The region code, time code, sensor type, spatial resolution, observation geometry and quality level are written into the data to obtain a unified spatiotemporal identifier set and to establish a spatiotemporal identifier index table. The preprocessed data set is bound one-to-one with the unified spatiotemporal identifier set, and entries below the preset quality threshold are removed to form the original multi-source dataset with unified spatiotemporal identifiers. The spatiotemporal identifier index table is then archived as a version baseline.

4. A computer vision-based system for identifying illegal logging in forest areas according to claim 1, characterized in that, The process involves sequentially performing geometric correction, radiometric correction, and topographic correction on the original multi-source dataset, and then implementing cross-sensor registration and uniform resolution resampling to output an aligned image set and a geofence index. Specifically, this includes: The visible light, infrared and multispectral images in the original multi-source dataset are registered according to the proportional polynomial and orthorectified in combination with the digital elevation model. The synthetic aperture radar images are subjected to imaging geometric correction and terrain projection correction. The lidar point cloud is subjected to trajectory calculation, attitude correction and strip stitching, and the geometric correction data set is output. Perform radiometric correction, atmospheric correction and radar calibration on the geometric correction dataset: convert the digital values ​​of the optical channels into apparent reflectivity, perform atmospheric correction based on the atmospheric transmission model, calibrate the radar channel backscatter intensity to the terrain-normalized intensity, perform range compensation and incident angle compensation on the point cloud echo intensity, and output the radiometric correction dataset. Cross-sensor registration is performed on the radiometric correction dataset as the operation object: corner detection and edge extraction are used to construct a scale space descriptor, and initial transformation parameters are obtained by combining random consistency elimination. Pyramid hierarchical search and local affine fine-tuning are performed by using mutual information and phase correlation as joint measures. Fine registration is completed according to the region window defined by the spatiotemporal identifier index table, and the fine registration dataset is output. The finely registered dataset is resampled with high fidelity to a unified pixel grid, cropped according to the regional grid number to generate aligned sub-images, and then aggregated to output an aligned image set. By combining the aligned image set with the forest compartment map and administrative boundaries, coordinate consistency and topological restoration operations are performed to generate fence faces and unique numbers, establish a one-to-one mapping and spatial inclusion relationship between the fence and the aligned submap, and output the geofence index.

5. A computer vision-based system for identifying illegal logging in forest areas according to claim 1, characterized in that, The process involves taking an aligned image set as input, pairing aligned sub-images from different time phases within the same region according to a spatiotemporal identifier index table, constructing an improved dual-stream change detection network for feature alignment and cross-modal fusion, obtaining a continuous change heatmap, and extracting a candidate change block set. Specifically, this includes: The aligned image set is retrieved, and aligned sub-images of different time phases are paired within the region and time window given by the spatiotemporal identification index table. Quality thresholds are set for cloud cover ratio, viewing angle difference, and registration residual. Only pairs that meet the thresholds are retained. For pairs that meet the thresholds, histogram matching and intensity standardization are performed on the optical channel to unify brightness and contrast. Amplitude normalization and phase quality check are performed on the radar channel to suppress differences in imaging power and interference quality. The pixel grid and pyramid scale are also unified to form standardized aligned sub-image pairs that can be directly compared. An improved dual-stream change detection network is constructed at the encoding end. The optical branch uses multi-scale residual blocks and deformable convolutions with channel attention superimposed, while the radar branch uses amplitude-phase decoupled convolutions and speckle robust residual blocks with channel attention superimposed. Cross-modal correlation fields are calculated on features of the same scale and cross-modal offset fields are regressed. Feature alignment is completed by sub-pixel resampling. At the same time, brightness and contrast are recalibrated for optical features and amplitude and phase are normalized for radar features. Alignment consistency score is output. Cross-modal attention and gated residual fusion are introduced at the bottleneck layer. Multi-scale features after two branches are aligned and superimposed are selected by scale-adaptive weight selection and superimposed. Pyramid aggregation compression is used to compress them into a fused semantic tensor. At the decoding end, spatial resolution is restored step by step through jump connection and upsampling. During the training phase, a hybrid loss of cross-entropy and overlap is used and hard sample focus terms and temporal consistency regularization are superimposed to stabilize learning. During the inference phase, sliding window overlapping inference and boundary smoothing are implemented to generate a continuously changing heat map. The pixel value of the continuously changing heat map represents the perturbation probability of the current position between time phases. Multi-threshold connected domain growth and scale space extremum detection are performed on continuously changing heatmaps. Candidate responses are extracted by combining morphological refinement and minimum area constraints. Adaptive thresholds are set based on alignment consistency scores and cross-modal consistency scores to eliminate unstable responses. Geofencing indexes are used to delete responses outside the fence. Finally, a set of candidate change blocks is output, and the area, aspect ratio, boundary curvature, and confidence score are recorded for each candidate change block.

6. A computer vision-based system for identifying illegal logging in forest areas according to claim 1, characterized in that, The process involves jointly performing semantic segmentation and instance segmentation around the candidate change block set under the constraints of the geofence index, refining the boundaries using a conditional random field, and performing raster-to-vector conversion and topology repair to generate a fine-grained target set. Specifically, this includes: The target region is generated around the candidate change block set under the constraints of the geofence index. The cropping is completed by polygon overlay and difference operation. The region is expanded outward by a fixed buffer distance and outbound cells are removed. Semantic segmentation and instance segmentation are jointly implemented in the target region. The semantic segmentation adopts an encoder-decoder network and a multi-scale pyramid, with cross-entropy superimposed overlap loss. The instance segmentation adopts candidate block-guided mask region proposal and pixel-level mask regression, and uses boundary-aware loss to improve edge separability. The semantic segmentation and instance segmentation results are analyzed and refined. First, the maximum a posteriori selection is performed based on the category probability and mask overlap. Then, the nearest neighbors of the same category are merged based on the minimum gap threshold. A conditional random field is used to minimize the energy by using the category probability as the unary potential and the color difference, gradient difference and spatial distance as the binary potential. Finally, an insurmountable constraint is applied to the fence boundary to obtain the refined object mask set. The refined object mask set is deleted by removing the mask connected components according to the preset area threshold, and opening and closing operations are performed for smoothing, hole filling and narrow channel disconnection. Based on contour tracking, the raster to vector conversion is completed, vertex simplification and polyline smoothing are performed, and vertex snapping and distance limit control are performed on the fence boundary of the preset range to obtain the vector feature set. Perform topology repair and target generation on the vector feature set, including: detecting and resolving self-intersections and overlaps, correcting hole orientations, stitching across tiles along shared boundaries, merging or deleting narrow patches according to area and aspect ratio thresholds, outputting a fine-grained target set, and recording the category, instance number, area, perimeter, aspect ratio, boundary curvature, relative position to the fence, and overall confidence level for each fine-grained target.

7. A computer vision-based system for identifying illegal logging in forest areas according to claim 1, characterized in that, The process, based on a fine-grained target set and relying on graph database knowledge constraints and a rule engine, completes spatiotemporal compliance verification. For targets deemed unstable, a minimum conflict set is extracted and fed back to trigger targeted re-inspection and time window expansion until the results stabilize. The output includes a list of verified targets, a confidence threshold table, and a domain adaptation model. Specifically, this includes: External information is collected object by object based on a fine-grained target set, including spatial overlay with forest compartments and permitted areas, recording forest compartment number, permit number and whether it is within the permit, constructing fixed buffers for roads and water bodies and calculating minimum distance and intersection relationship, performing regional statistics on slope grid to obtain mean, maximum value and over-threshold ratio, associating historical operation records with candidate change blocks according to time windows, determining whether the time window and cooling period are satisfied, and summarizing to form a knowledge constraint set and a check item table; Perform spatiotemporal compliance verification according to the verification checklist, including applying threshold rules to slope, using buffer overlay for roads and water bodies, using spatial overlay for permitted areas, using time windows and cooling-off periods to determine historical operations, outputting a set of compliant targets and a set of suspected non-compliant targets, and generating verification evidence records; Stability indicators are calculated for suspected non-compliant sets, including cross-temporal response consistency, overlap with geofence, evidence path integrity, and source freshness. Unstable targets are marked, an unstable label table is generated, and stable suspected sets and stability metrics are obtained. Extract the minimum conflict set from the unstable marker table, remove redundancies from the rules that trigger failure, retain the minimum entry group that can cause failure on its own, issue a feedback instruction based on the minimum conflict set, implement targeted supplementary inspection and time window expansion, add aligned subgraphs and neighborhood samples, recalculate the continuously changing heatmap and candidate change block set, and incrementally update the fine-grained target set. Threshold adaptation and domain adaptation are implemented for the compliant target set and the stable suspected set. A confidence threshold table is generated according to the confidence quantile and the target false alarm rate. Pseudo-label semi-supervised training and adaptive normalization are used for small incremental training to obtain the domain-adapted model.

8. A computer vision-based system for identifying illegal logging in forest areas according to claim 1, characterized in that, The process involves taking the target list and domain adaptation model as input, performing rapid inference and triggering alarms at the edge, and completing verification and multi-source confirmation in the cloud to form a set of adjudication results. Specifically, this includes: Read the list of verification targets and the domain adaptation model, load a lightweight engine at the edge to perform batch fast inference, and output the edge alarm candidate set and the corresponding confidence and time stamp; Set thresholds for filtering based on geofencing and alignment consistency, generate a list of on-site alarms, and record the device identification and triggering reason; The on-site alarm list is sent back to the cloud, and the continuous change heat map and candidate change blocks are recalculated based on the aligned image set and fine-grained target set to form the cloud review result; A second spatiotemporal compliance check is performed by combining historical operations, permitted scope, road and water body buffer zones, and slope, and a review opinion and evidence snapshot are output. Consistency determination and conflict resolution are performed on the on-site alarm list and the cloud review results. The decision confidence and reason code are calculated, and the decision result set is generated. The decision result set is bound to the spatiotemporal identifier and written to the regulatory interface and log.

Citation Information

Cited By

  • Detailed planning whole-process review method and system, medium and product

    CN122114868A