Cloud-edge integrated natural resource illegal clue real-time identification and positioning method

By combining edge computing devices with cloud databases, a cloud-edge integrated natural resource monitoring method has been adopted to achieve real-time and accurate identification and efficient location of illegal activities related to natural resources. This solves the problems of computing power and bandwidth pressure in existing technologies and improves identification efficiency and accuracy.

CN121505530APending Publication Date: 2026-02-10SURVEYING & MAPPING INST LANDS & RESOURCE DEPT OF GUANGDONG PROVINCE
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511479826.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-16
Publication Date
2026-02-10

AI Technical Summary

Technical Problem

Existing technologies are insufficient for real-time and accurate identification and high-precision visual positioning of illegal activities in natural resource monitoring. Edge computing is ill-suited for computationally intensive tasks, while cloud computing results in excessive bandwidth pressure and identification latency, failing to meet real-time requirements.

Method used

By adopting a cloud-edge integrated approach, edge computing devices are used to identify illegal activities through a lightweight backbone network with visible light and near-infrared dual branches. The identification is then matched with a cloud-based geometric positioning benchmark database. Combined with modal alignment feature fusion and environmental robustness adversarial training, lightweight feature extraction and high-precision positioning are achieved.

Benefits of technology

It improves the efficiency and accuracy of illegal behavior identification, enhances the model's generalization ability and identification robustness in complex environments, and meets the needs of real-time monitoring and efficient positioning.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121505530A_ABST
    Figure CN121505530A_ABST
Patent Text Reader

Abstract

The invention discloses a cloud-edge integrated natural resource illegal clue real-time identification and positioning method, and the method comprises the steps: collecting natural resource data through an edge computing device, recognizing a natural resource illegal behavior through a target detection and segmentation network, and storing the target data of the illegal behavior; and matching the illegal behavior target data with the cloud geometric positioning reference database, and determining the specific coordinate position of the illegal behavior. According to the method, the target detection and segmentation network is constructed based on the visible light and near-infrared double-branch lightweight backbone network to identify the natural resource data, lightweight feature extraction is realized, the calculation amount of a network model is reduced, and the detection and identification efficiency is improved; meanwhile, a visible light branch adopts MobileNetV3-Lite, a near-infrared branch adopts a 5 * 5 depth separable convolution group for feature extraction, and the recognition effect on natural resource data is improved; and a modal alignment feature fusion module is embedded, so that the dislocation problem is avoided, the feature extraction capability of the network is enhanced, and the recognition precision is ensured.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of natural resource law enforcement, and in particular to a cloud-edge integrated method for real-time identification and location of clues to illegal activities in natural resources. Background Technology

[0002] With its significant advantages such as high data timeliness, low deployment cost, wide coverage, and convenient maintenance, monocular surveillance cameras have rapidly replaced traditional manual patrols and satellite remote sensing methods, finding widespread application in the field of natural resource monitoring and becoming an ideal front-end sensing device for building dense monitoring networks. However, natural resource monitoring requires more than just video or image surveillance; it also necessitates real-time and accurate identification of various illegal activities (such as dump trucks, construction machinery, and illegal buildings) and precise acquisition of their actual locations. This demands that the monitoring system possess both real-time, accurate, and intelligent identification capabilities and high-precision visual positioning capabilities, both of which place extremely high demands on the underlying computing power.

[0003] To address the aforementioned needs, edge computing and cloud computing are two relatively mature solutions, each with its own applicable scenarios. Edge computing, by performing preliminary processing at devices close to the data source (such as smart cameras and edge servers), can significantly reduce latency and save bandwidth. However, limited by the computing power of edge nodes, it typically requires the deployment of highly optimized lightweight models and is ill-suited for computationally intensive visual geometric localization tasks. While cloud computing can provide powerful computing capabilities to support model training and high-precision visual geometric localization calculations, transmitting all video streams back to the cloud for processing will result in enormous bandwidth pressure and unacceptable recognition latency, thus failing to meet the real-time requirements of "early detection and early prevention." Summary of the Invention

[0004] To address the shortcomings of existing technologies, this invention provides a cloud-edge integrated method for real-time identification and location of illegal natural resource activities, thereby achieving the integration of edge computing and cloud computing to enable real-time monitoring and efficient location of illegal natural resource activities.

[0005] Therefore, the technical solution adopted by the present invention is as follows: This invention provides a cloud-edge integrated method for real-time identification and location of illegal natural resource clues, the method comprising: Natural resource data is collected using edge computing devices, and illegal activities related to natural resources are identified through a target detection and segmentation network, storing the target data of these illegal activities. The target detection and segmentation network extracts features based on a lightweight backbone network with visible and near-infrared dual branches, and embeds a modal alignment feature fusion module to obtain fused features. A decoupled detection head identifies illegal activities related to natural resources based on these fused features. Specifically, the visible light branch uses MobileNetV3-Lite to extract texture and shape features from the natural resource data, while the near-infrared branch uses 5×5 depth-separable convolutional groups to extract thermal radiation features from the natural resource data. The target data of illegal activities are uploaded to the cloud and matched with the cloud-based geometric positioning reference database to determine the specific coordinates of the illegal activities.

[0006] According to the above scheme, the modal alignment feature fusion module specifically calculates the adaptively fused feature matrix based on the visible light branch features, the near-infrared branch aligned features, the cross-modal cross-correlation matrix, and the learnable weight parameters; wherein, the cross-modal cross-correlation matrix is ​​calculated based on the original near-infrared branch features and the visible light branch features.

[0007] According to the above scheme, the target detection and segmentation network also specifically embeds an SE-HW hybrid attention module optimized based on the SE module to reconstruct the extracted features, and the SE-HW hybrid attention module specifically reconstructs the extracted texture shape features and thermal radiation features according to the height weight, width weight and channel number weight.

[0008] According to the above scheme, when the embedded modal alignment feature fusion module obtains the fused features, it specifically combines the dynamically weighted bidirectional feature pyramid method to perform weighted calculations on the features to be fused, thereby obtaining the fused features.

[0009] According to the above scheme, the target detection and segmentation network is trained using a standard sample dataset and an environment-robust adversarial method. The standard sample dataset is obtained by collecting image data of various types of illegal activities, processing and labeling the data. The illegal activity target data includes the type of natural resource illegal activity, detection time, camera PTZ parameters and video frame data to be located.

[0010] According to the above scheme, the training method using environmental robustness adversarial approach is based on a standard sample dataset. Generative adversarial network is used to generate environmental interference factors that are common in actual observations, and an improved Focal Loss function is used as the loss function for training. Specifically, the improved Focal Loss function improves the training weights of different types of samples in the standard sample dataset, setting the training weights of occluded target samples and small-sized target samples to be higher than the training weights of regular clear target samples.

[0011] According to the above scheme, the geometric positioning reference database is obtained in the following ways: Collect sample data from monocular surveillance cameras on iron towers and on-site image data from drones, and obtain feature points from the sample data and image data; Aerial triangulation is performed on sample data and image data based on feature points to obtain aerial triangulation results, including pose parameters and camera intrinsic parameters of monocular surveillance camera samples and UAV on-site images. The sample data and image data are thinned, and a geometric positioning benchmark database is constructed based on the thinned sample data, image data and aerial triangulation results.

[0012] According to the above scheme, the sample data collected from the monocular surveillance cameras of the tower cameras includes: Set the initial minimum pitch angle and maximum focal length, and make the tower camera rotate one full circle at that pitch angle and focal length to collect the corresponding image data; Decrease the focal length at regular intervals and fix the focal length after each decrease. Increase the pitch angle at regular intervals and rotate the tower camera around the pitch angle and focal length to collect the corresponding image data. Repeat this step until the focal length is reduced to the minimum value and the pitch angle is increased to the maximum value, thus completing the monocular monitoring camera sample data acquisition of the tower camera.

[0013] According to the above scheme, the on-site image data collection by UAV includes: Set the drone's focal length to the minimum focal length, and adjust the pitch angle and its maximum field of view; Obtain the elevation of the tower camera, calculate the range of the drone's pitch angle based on the elevation, set the drone's pitch angle interval, flight path and flight altitude, and capture and collect on-site images according to the set pitch angle interval, flight path and flight altitude to obtain on-site image data of the drone.

[0014] According to the above scheme, matching the target data of illegal activities with the cloud-based geometric positioning reference database specifically involves: The feature points of the video frames to be located in the target data of illegal behavior are extracted by using deep learning network feature extraction algorithms. Based on the feature points, similar images in the cloud geometric positioning benchmark database are located, and a certain number of similar images with the highest similarity are taken as adjacent images. Based on the feature points, the video frame to be located is matched with adjacent images to obtain the successfully matched feature points and their three-dimensional coordinates. Combined with the camera's PTZ parameters, the pose parameters of the video frame to be located are calculated through resection to obtain the specific coordinate location of the illegal act.

[0015] The beneficial effects of this invention are as follows: This invention constructs a target detection and segmentation network based on a lightweight backbone network with visible and near-infrared dual branches to identify natural resource data, achieving lightweight feature extraction, reducing the computational load of the network model, and improving detection and recognition efficiency; at the same time, it uses MobileNetV3-Lite for the visible light branch and 5×5 depthwise separable convolutional groups for the near-infrared branch for feature extraction, improving the recognition effect of natural resource data; and it embeds a modality-aligned feature fusion module, avoiding the problems of misalignment of visible light edge details and thermal radiation areas caused by simple feature addition or splicing and ignoring modal differences, further enhancing the feature extraction capability of the network and ensuring its recognition accuracy.

[0016] Furthermore, this invention calculates the feature matrix of the modality alignment feature fusion module through learnable weight parameters, which can automatically adjust modality dominance. For example, in rainy or foggy scenes where visible light is unfavorable, the near-infrared weight can be increased, thereby improving the versatility and adaptability of the model.

[0017] Furthermore, this invention reconstructs the extracted features by embedding an SE-HW hybrid attention module optimized based on the SE module. While retaining the channel feature enhancement function of the original SE, it also strengthens the spatial dimension processing capability through height weight and width weight, which significantly improves the global perception accuracy of multi-scale targets, such as small targets like construction workers and large targets like buildings under construction.

[0018] Furthermore, this invention trains the target detection and segmentation network using an environmental robustness adversarial method, which can simulate common environmental interference factors in actual observation and improve the model's generalization ability in harsh environments. At the same time, the improved Focal Loss function is used as the loss function for training, making the model more inclined to learn from difficult samples, thereby having a stronger recognition ability when dealing with complex environments or edge targets, and improving the system's recognition robustness and stability under various extreme conditions. Attached Figure Description

[0019] Figure 1 This is a schematic diagram of the method flow for the real-time identification and location of illegal clues related to natural resources using a cloud-edge integrated approach, according to an embodiment of the present invention. Figure 2 This is a schematic diagram of an edge computing device according to an embodiment of the present invention; Figure 3 This is a schematic diagram of the SE-HW module according to an embodiment of the present invention; Figure 4 This is a schematic diagram of a tower camera according to an embodiment of the present invention; Figure 5 This is a schematic diagram of the drone's shooting range according to an embodiment of the present invention; Figure 6This is a schematic diagram of the flight path captured by the drone according to an embodiment of the present invention; Figure 7 This is a schematic diagram of the overall design flow of an embodiment of the present invention; Figures 8(a) and 8(b) are schematic diagrams of real-time monitoring results of illegal acts according to an embodiment of the present invention; Figures 9(a), 9(b), 9(c), and 9(d) are schematic diagrams of feature point matching results for monocular surveillance camera sample data from tower cameras and on-site UAV image data. Detailed Implementation

[0020] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention.

[0021] To address the shortcomings of existing technologies and achieve real-time monitoring and efficient, accurate location of illegal activities related to natural resources, this invention provides a cloud-edge integrated method for real-time identification and location of clues related to illegal activities in natural resources. Figure 1 As shown, the method includes: S1. Collect natural resource data using edge computing devices, identify illegal activities related to natural resources through target detection and segmentation networks, and store target data of illegal activities.

[0022] S2. Upload the target data of the illegal act to the cloud and match it with the cloud geometric positioning reference database to determine the specific coordinates of the illegal act.

[0023] In this embodiment, the edge computing device used is such as Figure 2 As shown.

[0024] Specifically, the target detection and segmentation network employs a lightweight cross-modal feature enhancement network. It extracts features based on a lightweight backbone network with visible and near-infrared dual branches, and embeds a modal alignment feature fusion module to obtain fused features. The decoupled detection head then identifies illegal activities related to natural resources based on these fused features. Specifically, the visible light branch uses MobileNetV3-Lite to replace the traditional detection model backbone for extracting texture and shape features from natural resource data, while the near-infrared branch uses 5×5 depthwise separable convolutional groups to extract thermal radiation features. This depthwise separable convolution reduces the overall number of parameters by more than 80%. By replacing traditional convolution operations with depthwise separable convolution and grouped convolution, the number of model parameters and computational cost are effectively reduced, while retaining crucial semantic integration capabilities. This improved structure not only alleviates the model's burden but also accelerates the forward propagation process, making the model more suitable for deployment on edge computing devices.

[0025] The Modal Alignment Feature Fusion (MAF) module specifically calculates the adaptively fused feature matrix based on the visible light branch features, the near-infrared branch aligned features, the cross-modal cross-correlation matrix, and the learnable weight parameters. The cross-modal cross-correlation matrix is ​​calculated based on the original near-infrared branch features and the visible light branch features.

[0026] Specifically, the module first unifies the dimension of the near-infrared feature channels through 1×1 convolution, and then calculates the cross-modal cross-correlation matrix. Ultimately, learnable weight parameters , (Initial values ​​are 0.7 and 0.3 respectively) Perform adaptive fusion:

[0027] in It is a visible light branching feature. It is a primitive feature of the near-infrared branch. It is a feature after near-infrared branch alignment; These are the parameters of the cross-modal cross-correlation matrix. , These are learnable weight parameters.

[0028] Specifically, the object detection and segmentation network also embeds an SE-HW hybrid attention module optimized based on the SE module to reconstruct the extracted features, and the SE-HW hybrid attention module specifically determines the features based on the height weights. Width weight and channel number weight Extracted features Perform feature reconstruction: This module not only retains the feature enhancement capabilities in the channel dimension of the original SE module, but also adds the ability to process feature information in the height and width dimensions of the feature map, thereby improving the attention to multi-dimensional global features. The overall structure of the SE-HW module is as follows: Figure 3 As shown.

[0029] Specifically, when obtaining the fused features from the embedded modality alignment feature fusion module, a dynamically weighted bidirectional feature pyramid method is used to perform weighted calculations on the features to be fused, resulting in the fused features. The dynamically weighted bidirectional feature pyramid method incorporates learnable hierarchical weights. This mechanism improves the efficiency of cross-level feature interaction while reducing computational load. These are the learnable weight parameters corresponding to the feature channels of each layer in the network. The weights are normalized to positive numbers and summed to 1. This allows the network to automatically learn the importance of features at different levels during fusion without manual configuration. The method establishes a close connection between low-level detail features and high-level semantic features through a top-down and bottom-up bidirectional path. Finally, the fused multi-scale features are fed into a lightweight, decoupled detection head. The classification branch outputs the violation category, the regression branch predicts the target location box, and optionally, the segmentation branch generates pixel-level violation ranges. This method not only significantly reduces computation but also effectively maintains the ability to interact across features at different levels, making it particularly suitable for target detection tasks in complex natural environments.

[0030] Traditional feature pyramid structures suffer from unidirectional feature integration and low information flow efficiency, easily losing information about small targets in complex contexts. The bidirectional feature pyramid used in this embodiment introduces both top-down and bottom-up paths in its structural design, enhancing the smoothness of cross-level feature fusion. Furthermore, by introducing a learnable fusion weight mechanism, the model can automatically allocate fusion ratios based on the importance of features at different scales, thereby improving the effectiveness and flexibility of feature integration. This structure allows for efficient information interaction between feature maps from different scales, particularly establishing a closer connection between low-level detailed features and high-level semantic features. This ensures that both small targets, such as construction workers, and large targets, such as buildings under construction, can be fully perceived and accurately identified.

[0031] Specifically, the target detection and segmentation network is trained using a standard sample dataset and an environment-robust adversarial method. The standard sample dataset is obtained by collecting image data of various types of illegal activities, processing and labeling the data. The target data of illegal activities includes the type of illegal activity in natural resources, detection time, camera PTZ parameters and video frame data to be located. The images of illegal activities include houses under construction, land reclamation, vegetation destruction and construction machinery.

[0032] The training process employs an environmental robustness adversarial approach. Based on a standard sample dataset, a Generative Adversarial Network (GAN) is used to generate simulated environmental disturbances common in real-world observations, including degradation scenarios such as rain, fog, dust storms, and low light. During generation, the generator adds disturbance features consistent with meteorological and physical laws, and image quality constraints (e.g., a peak signal-to-noise ratio (PSNR) of at least 28 dB) ensure the realism and usability of the synthesized samples. A discriminator evaluates the effectiveness of the generated samples, selecting high-quality perturbation images and adding them to the training set to improve the model's generalization ability in harsh environments.

[0033] Furthermore, an improved Focal Loss function is used as the loss function during network training. Specifically, the improved Focal Loss function adjusts the training weights of different types of samples in the standard sample dataset, assigning higher training weights to occluded and small-sized target samples than to regular, clear target samples. In the context of class imbalance, traditional loss functions may underestimate the training value of occluded or small-sized targets; therefore, this method assigns a higher weight (0.8) to these samples, compared to only 0.2 for regular, clear samples. In addition, a focus factor mechanism is used to improve the gradient response to error-prone samples, making the model more inclined to learn from difficult samples, thus possessing stronger recognition capabilities when dealing with complex environments or edge targets.

[0034] The entire training process is divided into three iterative optimization stages, gradually adjusting the diversity of generated samples, weighting strategies, and loss parameters to ensure that the model continuously strengthens its adaptability to adverse environmental factors during convergence. By continuously guiding the model to focus on challenging samples such as low-quality images, small targets, and occluded regions, this strategy significantly improves the system's robustness and stability under various extreme conditions. Specifically, the benchmark database is obtained in the following ways: Collect sample data from monocular surveillance cameras on iron towers and on-site image data from drones, and obtain feature points from the sample data and image data; Aerial triangulation is performed on the sample data and image data based on feature points to obtain the aerial triangulation results, including the pose parameters of the image and the camera intrinsic parameters. The sample data and image data are thinned, and a benchmark database is constructed based on the thinned sample data, image data and aerial triangulation results.

[0035] Specifically, the tower camera used in this embodiment is as follows: Figure 4 As shown, the monocular surveillance camera sample data of the tower camera is displayed at set intervals at multiple focal lengths ( - ) and multiple pitch angles ( - Under these conditions, a 360-degree panoramic view using a tower camera is used to collect camera sample data. For maximum focal length and The minimum focal length specifically includes: Set the initial pitch angle = Starting focal length = Operate the tower camera to slowly rotate 360° at a reasonable speed, ensuring that each frame of video is clear and unblurry, and that the overlap between adjacent images is 80% when acquiring image data; At the same pitch angle Next, decrease the focal length at the set intervals, repeat the previous step, and continue to collect data in a circular motion until the final focal length is reached. Complete data collection; Increase the pitch angle at 10° intervals, repeating the above two steps until the final pitch angle is reached. = Once data acquisition is complete, at this point, at least half of the image captured by the monocular camera on the tower must be of the observed ground features.

[0036] In addition, the on-site image data from the drone is collected using a high-precision aerial survey drone to simulate camera photography, based on the position and angle of the monocular camera on the tower camera. This includes: Set the drone's focal length to the minimum and adjust the pitch angle so that the maximum field of view includes both the nearest and farthest ground features; for example... Figure 5 As shown, the maximum visible range is (Red), the preset radii corresponding to the effective coverage areas are respectively (Yellow inner circle) and (Yellow outer ring) By adjusting the camera angle, ensure that both the farthest and nearest objects are clearly visible in the image at the same time; The elevation of the tower's location is obtained by generating terrain data from satellite imagery. This elevation, plus the camera's mounting height, is used as the tower-camera elevation. The lower elevations of local areas on the concentric circles of the central optical ray of the camera are respectively and (Lower elevation is defined as expanding the ring by 10m), and elevation sampling statistics are performed on the scanned annular area; then, based on trigonometric functions, the elevation from the camera to the concentric circles is estimated. and The range of pitch angles, i.e. and ,in:

[0037]

[0038] Based on the range of overhead shooting angles and Set the drone camera's overhead shooting angle at 5° intervals (covering...) and Subsequently, corresponding to flight paths with radii of 20m and 50m, flight altitudes were set to -20m, 0m, 20m, and 40m respectively; finally, based on the aforementioned overhead angle, flight path, and flight altitude, data was collected along the flight path from the perspective of the simulated tower camera with a 90% overlap, such as... Figure 6 As shown.

[0039] Preferably, the monocular monitoring camera sample data of the tower camera and the on-site image data of the UAV are automatically matched for connection points and subjected to high-precision geometric adjustment to obtain high-precision reference data, including oriented images, digital surface models and orthophotos.

[0040] Specifically, in this embodiment, when performing aerial triangulation on the drone's on-site image data simulating a camera, feature points from similar images are used as observations. These similar images are located based on the feature points of the drone's on-site image data, and image position parameters measured by high-precision RTK are used as constraints. Global SFM is then used to calculate the image pose and camera parameters, and the object coordinates of all connection points are calculated. Subsequently, using the connection points of the drone image set simulating a camera as control data, aerial triangulation of the camera sample data is performed. Since the camera only rotates and does not displace, incremental SFM is used to calculate the camera pose image by image. During incremental SFM calculation, the camera position remains almost unchanged, and the 3D points matched with the drone image set can be used as constraints for local adjustment, ensuring the accuracy of the aerial triangulation results for the camera images.

[0041] Finally, based on the actual observation perspective in real-world application scenarios, the drone image samples and camera image samples, which resemble cameras, were thinned with a 60% overlap. A benchmark database was established by retrieving information such as the aerial triangulation results, global feature vectors, camera intrinsic parameters, and the mapping relationship between focal length and scaling factor for all images.

[0042] Specifically, matching the target data of illegal activities with the cloud-based geometric positioning benchmark database involves: The feature points of the video frames to be located in the target data of illegal behavior are extracted by using deep learning network feature extraction algorithms. Based on the feature points, similar images in the cloud geometric positioning benchmark database are located, and a certain number of similar images with the highest similarity are taken as adjacent images. The video frame to be located is matched with adjacent images using SuperGlue features to obtain the successfully matched feature points and their three-dimensional coordinates. Combined with the camera's PTZ parameters, the pose parameters of the video frame to be located are calculated through resection to obtain the specific coordinate location of the illegal act.

[0043] Preferably, in this embodiment, the SuperPoint deep learning network feature extraction algorithm is used to extract feature points from the video frame to be located. Then, a NetVLAD-based image retrieval method is used to obtain similar images of the video frame to be located. The ten most similar images are selected as adjacent images, and SuperGlue feature matching is performed between the video frame to be located and the adjacent images to obtain successfully matched feature points and their corresponding 3D coordinates. Additionally, based on the scaling factor of the video frame to be located, its actual focal length value is retrieved. Then, combined with camera intrinsic data, and using the successfully matched feature points and corresponding 3D points, the pose parameters of the video frame to be located are calculated through resection to complete the localization.

[0044] The overall design flowchart of this embodiment is as follows: Figure 7 As shown in the figure; the results of the cloud-edge integrated real-time identification and location method for illegal clues of natural resources in this embodiment are shown in the figure; Figure 8(a) and Figure 8(b) represent the real-time monitoring results of illegal behavior; Figure 9(a), Figure 9(b), Figure 9(c) and Figure 9(d) represent the feature point matching results of four images of monocular monitoring camera sample data of iron tower camera and on-site image data of UAV respectively.

[0045] This invention employs a lightweight backbone network based on visible and near-infrared dual branches to construct a target detection and segmentation network for natural resource data identification. This achieves lightweight feature extraction, reduces the computational load of the network model, and improves detection and recognition efficiency. Furthermore, MobileNetV3-Lite is used for feature extraction in the visible light branch, and 5×5 depthwise separable convolutional groups are used in the near-infrared branch, enhancing the recognition effect on natural resource data. Additionally, an embedded modality-aligned feature fusion module avoids the problems of misalignment between visible light edge details and thermal radiation areas caused by simple feature addition or splicing and neglecting modal differences, further enhancing the network's feature extraction capabilities and ensuring its recognition accuracy.

[0046] Furthermore, in this embodiment of the invention, the feature matrix calculation of the modality alignment feature fusion module is performed through learnable weight parameters, which can automatically adjust modality dominance. For example, in visible light-disadvantaged scenarios such as rain and fog, the near-infrared weight can be increased, thereby improving the versatility and adaptability of the model.

[0047] Furthermore, in this embodiment of the invention, the extracted features are reconstructed by embedding an SE-HW hybrid attention module optimized based on the SE module. While retaining the channel feature enhancement function of the original SE, the spatial dimension processing capability is also enhanced by height weight and width weight, which significantly improves the global perception accuracy of multi-scale targets, such as small targets of construction workers and large targets of buildings under construction.

[0048] Furthermore, this embodiment of the invention trains the target detection and segmentation network using an environmental robustness adversarial method, which can simulate common environmental interference factors in actual observation and improve the model's generalization ability in harsh environments. At the same time, the improved Focal Loss function is used as the loss function for training, making the model more inclined to learn from difficult samples, thereby having a stronger recognition ability when dealing with complex environments or edge targets, and improving the system's recognition robustness and stability under various extreme conditions.

[0049] It should be noted that, depending on the implementation needs, the various steps / components described in this application can be broken down into more steps / components, or two or more steps / components or parts of the operation of steps / components can be combined into new steps / components to achieve the purpose of this invention.

[0050] The order of the steps in the above embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.

[0051] It should be understood that those skilled in the art can make improvements or modifications based on the above description, and all such improvements and modifications should fall within the protection scope of the appended claims.

Claims

1. A method for real-time identification and location of illegal natural resource clues integrating cloud and edge computing, characterized in that, The method includes: Natural resource data is collected using edge computing devices, and illegal activities related to natural resources are identified through a target detection and segmentation network, storing the target data of these illegal activities. The target detection and segmentation network extracts features based on a lightweight backbone network with visible and near-infrared dual branches, and embeds a modal alignment feature fusion module to obtain fused features. A decoupled detection head identifies illegal activities related to natural resources based on these fused features. Specifically, the visible light branch uses MobileNetV3-Lite to extract texture and shape features from the natural resource data, while the near-infrared branch uses 5×5 depth-separable convolutional groups to extract thermal radiation features from the natural resource data. The target data of illegal activities are uploaded to the cloud and matched with the cloud-based geometric positioning reference database to determine the specific coordinates of the illegal activities.

2. The method for real-time identification and location of illegal natural resource clues integrating cloud and edge computing as described in claim 1, characterized in that, The modal alignment feature fusion module specifically calculates the adaptively fused feature matrix based on the visible light branch features, the near-infrared branch aligned features, the cross-modal cross-correlation matrix, and the learnable weight parameters; wherein, the cross-modal cross-correlation matrix is ​​calculated based on the original near-infrared branch features and the visible light branch features.

3. The method for real-time identification and location of illegal natural resource clues integrating cloud and edge computing as described in claim 2, characterized in that, The target detection and segmentation network also specifically embeds an SE-HW hybrid attention module optimized based on the SE module to reconstruct the extracted features. Specifically, the SE-HW hybrid attention module reconstructs the extracted texture shape features and thermal radiation features based on height weight, width weight, and channel number weight.

4. The method for real-time identification and location of illegal natural resource clues integrating cloud and edge computing as described in claim 3, characterized in that, When the embedded modal alignment feature fusion module obtains the fused features, it specifically combines the dynamically weighted bidirectional feature pyramid method to perform weighted calculations on the features to be fused, thus obtaining the fused features.

5. The method for real-time identification and location of illegal natural resource clues integrating cloud and edge computing as described in claim 3, characterized in that, The object detection and segmentation network is trained using a standard sample dataset and an environment-robust adversarial method. The standard sample dataset is obtained by collecting image data of various types of illegal activities, and then processing and labeling the data. The target data for illegal activities includes the type of illegal activity related to natural resources, the detection time, the camera's PTZ parameters, and the video frame data to be located.

6. The method for real-time identification and location of illegal natural resource clues integrating cloud and edge computing as described in claim 5, characterized in that, The training method employs an environment-robust adversarial approach. Specifically, it uses a standard sample dataset as a basis, employs a generative adversarial network to generate environmental interference factors that are common in actual observations, and uses an improved Focal Loss function as the loss function for training. Specifically, the improved Focal Loss function improves the training weights of different types of samples in the standard sample dataset, setting the training weights of occluded target samples and small-sized target samples to be higher than the training weights of regular clear target samples.

7. The method for real-time identification and location of illegal natural resource clues integrating cloud and edge computing as described in claim 1, characterized in that, The geometric positioning reference database is obtained in the following ways: Collect sample data from monocular surveillance cameras on iron towers and on-site image data from drones, and obtain feature points from the sample data and image data; Aerial triangulation is performed on sample data and image data based on feature points to obtain aerial triangulation results, including pose parameters and camera intrinsic parameters of monocular surveillance camera samples and UAV on-site images. The sample data and image data are thinned, and a geometric positioning reference database is constructed based on the thinned sample data, image data and aerial triangulation results.

8. The method for real-time identification and location of illegal natural resource clues integrating cloud and edge computing as described in claim 7, characterized in that, The sample data collected from the monocular surveillance cameras on the tower cameras includes: Set the initial minimum pitch angle and maximum focal length, and make the tower camera rotate one full circle at that pitch angle and focal length to collect the corresponding image data; Decrease the focal length at regular intervals and fix the focal length after each decrease. Increase the pitch angle at regular intervals and rotate the tower camera around the pitch angle and focal length to collect the corresponding image data. Repeat this step until the focal length is reduced to the minimum value and the pitch angle is increased to the maximum value, thus completing the monocular monitoring camera sample data acquisition of the tower camera.

9. The cloud-edge integrated real-time identification and location method for illegal clues of natural resources as described in claim 7, characterized in that, The drone-based on-site image data collection includes: Set the drone's focal length to the minimum focal length, and adjust the pitch angle and its maximum field of view; Obtain the elevation of the tower camera, calculate the range of the drone's pitch angle based on the elevation, set the drone's pitch angle interval, flight path and flight altitude, and capture and collect on-site images according to the set pitch angle interval, flight path and flight altitude to obtain on-site image data of the drone.

10. The method for real-time identification and location of illegal natural resource clues integrating cloud and edge computing as described in claim 5, characterized in that, Matching the target data of illegal activities with the cloud-based geometric positioning benchmark database specifically involves: The feature points of the video frames to be located in the target data of illegal behavior are extracted by using deep learning network feature extraction algorithms. Based on the feature points, similar images in the cloud geometric positioning benchmark database are located, and a certain number of similar images with the highest similarity are taken as adjacent images. Based on the feature points, the video frame to be located is matched with adjacent images to obtain the successfully matched feature points and their three-dimensional coordinates. Combined with the camera's PTZ parameters, the pose parameters of the video frame to be located are calculated through resection to obtain the specific coordinate location of the illegal act.