A low-altitude remote sensing small target rapid detection method and related equipment

CN122551208APending Publication Date: 2026-08-11GUANGDONG KENUO SURVEYING ENG CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-04-14
Publication Date
2026-08-11

AI Technical Summary

Technical Problem

[0002]当前低空遥感技术在工业检测领域的应用日益广泛,但小目标检测仍面临多重技术瓶颈:一方面,低空场景下小目标尺度极小、特征微弱,且受光照变化、云雾遮挡、相似物干扰等复杂环境因素影响,传统检测方法易出现漏检、误检;另一方面,大范围巡检产生海量影像数据,对检测系统的实时性与数据处理效率提出极高要求,而现有技术存在数据采集适应性差、特征提取不充分、多源数据融合不足、小样本学习泛化能力弱等问题

Benefits of technology

[0015]本发明的实施例至少包括以下有益效果:本发明提供一种低空遥感小目标快速检测方法和相关设备,该方案通过对采集影像进行图像增强以及标准化,得到标准化影像,有效抑制背景噪声、光照不均等干扰;通过轻量化学生模型,对标准化影像进行视觉检测,生成第一疑似缺陷数据,并将第一疑似缺陷数据与定位信息进行关联,实现实时、快速的初步缺陷定位与提取,显著降低计算开销与延迟,为后续精细化处理筛选出关键候选区域;通过多尺度空频特征融合网络,对第一疑似缺陷数据进行特征优化,得到第二疑似缺陷数据,强化对小目标及细微缺陷的特征表征能力,从而提升检测精度;第二疑似缺陷数据包括视觉置信度;通过教师模型,对第二疑似缺陷数据进行二次检测,得到第三疑似缺陷数据,能够有效排除因纹理相似、阴影等导致的视觉伪缺陷,提升判断的语义合理性与准确性;第三疑似缺陷数据包括语义验证标记;根据语义验证标记以及视觉置信度,对第三疑似缺陷数据进行筛选,得到有效缺陷数据,实现高置信度缺陷的精准筛选;通过大语言模型以及领域知识库,对有效缺陷数据进行语义验证,输出目标检测结果,并将目标检测结果与定位信息进行关联,得到结构化检测记录,该方案够提高小目标检测的精度和效率。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122551208A_ABST
    Figure CN122551208A_ABST
Patent Text Reader

Abstract

This invention discloses a rapid detection method and related equipment for small targets in low-altitude remote sensing, comprising: image enhancement and standardization of acquired images to obtain standardized images; visual inspection of the standardized images using a lightweight student model to generate first suspected defect data, which is then associated with location information; optimization of the first suspected defect data using a multi-scale spatial-frequency feature fusion network to obtain second suspected defect data; secondary detection of the second suspected defect data using a teacher model to obtain third suspected defect data; filtering of the third suspected defect data based on semantic verification tags and visual confidence to obtain valid defect data; semantic verification of the valid defect data using a large language model and domain knowledge base, outputting target detection results, and associating them with location information to obtain structured detection records. This invention can improve the accuracy and efficiency of small target detection and can be widely applied in the field of computer vision technology.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of computer vision technology, and in particular to a method and system for rapid detection of small targets in low-altitude remote sensing. Background Technology

[0002] Currently, low-altitude remote sensing technology is increasingly widely used in industrial inspection, but small target detection still faces multiple technical bottlenecks: On the one hand, small targets in low-altitude scenarios are extremely small in scale and have weak features, and are affected by complex environmental factors such as changes in illumination, cloud and fog obstruction, and interference from similar objects, making traditional detection methods prone to missed detections and false detections; on the other hand, large-scale inspections generate massive amounts of image data, placing extremely high demands on the real-time performance and data processing efficiency of the detection system, while existing technologies suffer from problems such as poor data acquisition adaptability, insufficient feature extraction, inadequate fusion of multi-source data, and weak generalization ability for small sample learning.

[0003] Meanwhile, existing technical solutions for small-sample defect detection mostly focus on sample generation in the image preprocessing stage. They only guide adversarial generative networks to expand multi-class samples through image category features, lacking in-depth guidance from knowledge in the industrial inspection field. The generated samples have poor adaptability to real low-altitude remote sensing defect scenarios and do not form a linkage with subsequent detection and inference stages, failing to fundamentally solve the false detection problem of small target detection. Single-terminal or cloud computing architectures are difficult to balance detection accuracy and real-time performance, resulting in limited engineering applicability. Existing detection technologies mostly focus on optimizing a single stage, lacking a systematic design of the entire process from data acquisition, feature enhancement, semantic understanding to collaborative computing, and cannot efficiently solve the core problem of low-altitude remote sensing small target detection. Summary of the Invention

[0004] In view of this, the main objective of the embodiments of the present invention is to provide a method and system for rapid detection of small targets in low-altitude remote sensing, in order to solve at least one of the problems of the prior art. The present invention can improve the accuracy and efficiency of small target detection.

[0005] To achieve the above objectives, one aspect of the present invention provides a method for rapid detection of small targets in low-altitude remote sensing, the method comprising: The acquired images are enhanced and standardized to obtain standardized images; A lightweight student model is used to perform visual inspection on the standardized image to generate first suspected defect data, and the first suspected defect data is associated with the location information. The first suspected defect data is optimized using a multi-scale spatial-frequency feature fusion network to obtain the second suspected defect data; the second suspected defect data includes visual confidence. The second suspected defect data is subjected to a second detection using a teacher model to obtain a third suspected defect data; the third suspected defect data includes semantic verification tags. Based on the semantic verification markers and the visual confidence level, the third suspected defect data is filtered to obtain valid defect data; Using a large language model and a domain knowledge base, the valid defect data is semantically verified, the target detection results are output, and the target detection results are associated with the location information to obtain a structured detection record.

[0006] In some embodiments, the process of image enhancement and standardization of the acquired images to obtain standardized images includes the following steps: An enhanced image is obtained by strengthening the edge distinction between small targets and the background in the acquired image through an adaptive histogram equalization algorithm. The enhanced image is scaled and normalized to obtain the standardized image.

[0007] In some embodiments, the step of optimizing the features of the first suspected defect data through a multi-scale space-frequency feature fusion network to obtain the second suspected defect data includes the following steps: The first suspected defect data is subjected to wavelet transform by the wavelet decomposition and feature splicing module to obtain frequency domain sub-bands; The frequency domain sub-band is spliced ​​with the first suspected defect data, and channel fusion and dimensionality reduction are performed to obtain the first enhanced feature. The first enhanced feature is decomposed by the feature channel space-frequency decomposition and fusion module to obtain the initial spatial domain feature and the initial frequency domain feature. A non-uniform interval feature selection strategy is used to reorganize the spatial context information of the initial spatial features to obtain the target spatial features. The initial frequency domain features are subjected to wavelet transform to obtain the target frequency domain features; The target spatial domain features and the target frequency domain features are fused to obtain the second enhanced feature; By using the cross-self-attention Swin Transformer module, the shallow and deep features of the second enhanced feature are strongly correlated to obtain the second suspected defect data.

[0008] In some embodiments, the step of performing a second detection on the second suspected defect data using a teacher model to obtain the third suspected defect data includes the following steps: The second suspected defect data and the preset domain knowledge prompts are input into the teacher model; The second suspected defect data is feature-encoded using a visual encoder to obtain visual features; Semantic features are obtained by feature encoding the domain knowledge prompts using a semantic encoder. The visual features are matched with the semantic features using a cross-modal attention mechanism, and the semantic verification tag is output. The semantic verification tag is associated with the second suspected defect data to generate the third suspected defect data.

[0009] In some embodiments, the method further includes the following steps: Based on the full amount of valid data, the multi-scale space-frequency feature fusion network is trained to generate optimized network parameters; The teacher model is updated based on the optimized network parameters; Based on the updated teacher model, lightweight parameters are generated through knowledge distillation techniques. The lightweight student model is updated based on the lightweight parameters. The full set of valid data includes real defect samples and synthetic defect samples.

[0010] In some embodiments, the method further includes the following steps: When it is detected that the number of samples in the full amount of valid data is insufficient, the adversarial generative network is invoked to generate the synthetic defective samples. Based on the structural similarity between the synthesized defect sample and the real defect sample, obtain the structural similarity loss; Obtain the synthetic feature embedding vector of the synthetic defect sample; Obtain the true feature embedding vector of the real defect sample; The domain consistency loss is obtained based on the cosine distance between the synthesized feature embedding vector and the real feature embedding vector. The adversarial generative network is trained based on the adversarial loss, the structural similarity loss, and the domain consistency loss.

[0011] To achieve the above objectives, another aspect of the present invention provides a rapid detection device for small targets in low-altitude remote sensing, the device comprising: The preprocessing module is used to enhance and standardize the acquired images to obtain standardized images; The lightweight visual inspection module is used to perform visual inspection on the standardized image using a lightweight student model, generate first suspected defect data, and associate the first suspected defect data with the positioning information. A multi-scale feature optimization module is used to optimize the features of the first suspected defect data through a multi-scale space-frequency feature fusion network to obtain the second suspected defect data; the second suspected defect data includes visual confidence. The cross-modal verification module is used to perform secondary detection on the second suspected defect data through the teacher model to obtain the third suspected defect data; the third suspected defect data includes semantic verification tags; The filtering module is used to filter the third suspected defect data based on the semantic verification marker and the visual confidence level to obtain valid defect data; The semantic verification module is used to perform semantic verification on the valid defect data through a large language model and a domain knowledge base, output the target detection result, and associate the target detection result with the location information to obtain a structured detection record.

[0012] To achieve the above objectives, another aspect of the present invention provides an electronic device, the electronic device including a memory and a processor, the memory storing a computer program, and the processor executing the computer program to implement the method described above.

[0013] To achieve the above objectives, another aspect of the present invention provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the methods described above.

[0014] To achieve the above objectives, another aspect of the present invention provides a computer program product or computer program that includes computer instructions stored in a computer-readable storage medium. A processor of a computer device can read the computer instructions from the computer-readable storage medium and execute the computer instructions to cause the computer device to perform the aforementioned method.

[0015] The embodiments of the present invention include at least the following beneficial effects: The present invention provides a method and related equipment for rapid detection of small targets in low-altitude remote sensing. This scheme obtains standardized images by enhancing and standardizing the acquired images, effectively suppressing interference such as background noise and uneven illumination; visual detection of the standardized images is performed using a lightweight student model to generate first suspected defect data, and the first suspected defect data is associated with the location information to achieve real-time and rapid preliminary defect location and extraction, significantly reducing computational overhead and latency, and filtering out key candidate regions for subsequent refined processing; feature optimization of the first suspected defect data is performed using a multi-scale spatial-frequency feature fusion network to obtain second suspected defect data, enhancing the feature representation capability of small targets and subtle defects. This improves detection accuracy. The second suspected defect data includes visual confidence. Using a teacher model, the second suspected defect data undergoes secondary detection to obtain the third suspected defect data. This effectively eliminates visual false defects caused by texture similarity, shadows, etc., improving the semantic rationality and accuracy of the judgment. The third suspected defect data includes semantic verification tags. Based on the semantic verification tags and visual confidence, the third suspected defect data is filtered to obtain valid defect data, achieving accurate screening of high-confidence defects. Through a large language model and domain knowledge base, semantic verification is performed on the valid defect data, outputting target detection results. These results are then associated with positioning information to obtain structured detection records. This scheme improves the accuracy and efficiency of small target detection. Attached Figure Description

[0016] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0017] Figure 1 This is a flowchart of the rapid detection method for small targets in low-altitude remote sensing provided in an embodiment of the present invention; Figure 2 This is a schematic diagram of the overall architecture of the low-altitude remote sensing small target rapid detection system provided in an embodiment of the present invention; Figure 3 This is a flowchart of the low-altitude remote sensing small target rapid detection system provided in this embodiment of the invention; Figure 4 This is a network architecture diagram of SF-SCST provided in an embodiment of the present invention; Figure 5 This is a schematic diagram of the cross-self-attention Swin Transformer module provided in an embodiment of the present invention; Figure 6This is a schematic diagram of the hardware structure of the electronic device provided in an embodiment of the present invention. Detailed Implementation

[0018] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative of the invention and are not intended to limit the invention. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with those of this invention; they are merely examples of apparatuses and methods consistent with some aspects of the embodiments of this invention as detailed in the appended claims.

[0019] It should be noted that although functional modules are divided in the system diagram and a logical order is shown in the flowchart, in some cases, the steps shown or described may be performed in a different order than the module division in the system or the order in the flowchart. The terms "first / S100" and "second / S200" in the specification, claims, and the foregoing drawings may be used herein to describe various concepts, but unless specifically stated otherwise, these concepts are not limited by these terms. These terms are used only to distinguish one concept from another. For example, first information may also be referred to as second information without departing from the scope of the embodiments of the invention, and similarly, second information may also be referred to as first information. Depending on the context, the words "if" or "when" as used herein may be interpreted as "when," "in response to a determination," or "in the event of a determination."

[0020] The terms “at least one,” “multiple,” “each,” “any,” etc., used in this invention, “at least one” includes one, two, or more than two; “multiple” includes two or more than two; “each” refers to each of the corresponding multiple; and “any” refers to any one of the multiple.

[0021] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains. The terminology used herein is for the purpose of describing embodiments of the invention only and is not intended to limit the invention.

[0022] With the rapid development of UAV platforms, sensor technology, and artificial intelligence algorithms, low-altitude remote sensing has become an indispensable core technology in fields such as water conservancy project inspection, photovoltaic power station operation and maintenance, power line inspection, and transportation facility monitoring. In key sectors like photovoltaics, water conservancy, power, and transportation, the small targets that low-altitude remote sensing needs to detect are characterized by their tiny scale, scattered distribution, and weak features. Currently, low-altitude remote sensing small target detection technology mainly revolves around data acquisition, feature extraction, and target recognition. However, in terms of scene-adaptive acquisition, it cannot dynamically adjust according to terrain complexity, target distribution density, and real-time environmental interference (light, fog, wind disturbance), resulting in problems such as blurred details of small targets and severe background interference in the acquired images, affecting subsequent detection accuracy. Regarding feature extraction and fusion, existing technologies are mostly limited to single-spatial-domain feature extraction, failing to effectively fuse frequency domain detail information and lacking a collaborative fusion mechanism for multi-source data (visible light, infrared, etc.), making it difficult to comprehensively capture the weak features of small targets. This leads to a high false negative rate in scenarios with extremely small target scales and complex environmental interference. In terms of few-shot learning and semantic understanding, existing data augmentation methods lack domain knowledge guidance, resulting in poor sample quality. Models rely solely on visual feature reasoning, lacking semantic knowledge support, leading to poor robustness and high false detection rates in challenging scenarios such as blurriness and occlusion. Regarding practicality and engineering applicability, existing terminal computing power cannot support complex models, and cloud processing suffers from transmission delays, making it difficult to balance detection accuracy and real-time requirements, thus failing to meet the engineering application requirements of large-scale inspections. Furthermore, existing high-precision detection models are structurally complex and difficult to run efficiently on UAV edge devices; lightweight models suffer from insufficient feature extraction capabilities, failing to balance detection speed and accuracy.

[0023] In view of this, this invention provides a method and related equipment for rapid detection of small targets in low-altitude remote sensing. This scheme combines an improved image enhancement algorithm to enhance and standardize the acquired images, improving the distinction between small targets and the background, obtaining high-quality acquired data, and laying the foundation for subsequent detection. A lightweight student model is used for preliminary visual detection, and a designed multi-scale spatial-frequency feature fusion network is used to fully integrate spatial semantic information and frequency-domain detail features, comprehensively capturing weak features of small targets and reducing the false negative rate in complex scenes. Based on knowledge distillation technology, semantic transfer is achieved from a multimodal large model (teacher model) to a pure visual lightweight model (student model), thereby filtering and retaining effective defect data. A large language model is introduced to construct a cross-modal semantic reasoning framework. Through visual-language feature alignment and fusion, domain knowledge priors are injected to perform semantic reasoning on defect features, thereby correcting misjudgments in blurred / occluded scenes and outputting the final detection result.

[0024] The low-altitude remote sensing small target rapid detection method provided in this invention relates to the interdisciplinary fields of UAV remote sensing, computer vision, and artificial intelligence. It is applicable to scenarios such as photovoltaic power plant operation and maintenance, water conservancy project inspection, power line inspection, and traffic facility monitoring, enabling automated, high-precision, and high-efficiency detection of minute defects smaller than 32×32 pixels (such as microcracks in photovoltaic panels, cracks in water conservancy dams, and damaged power cables). The low-altitude remote sensing small target rapid detection method provided in this invention can be deployed in a lightweight three-layer collaborative architecture of "cloud-based coordinated training - edge-based precise verification - edge-based real-time detection." It constructs a complete detection process of "adaptive flight path planning and scene-adaptive data acquisition, edge-based preprocessing and preliminary detection, edge-based secondary precise verification, cloud-based semantic reasoning and model iteration, and data storage and result output," achieving fully automated and intelligent processing from data acquisition to result output, thereby improving the accuracy, efficiency, and robustness of small target detection in low-altitude scenarios. Optionally, as... Figure 2 As shown, in the forward detection process, the edge layer acts as the data entry point to complete data collection and preliminary detection, outputting suspected defect data to the edge layer; the edge layer acts as an intermediate hub to complete secondary precise verification, filtering valid defect data and uploading it to the cloud layer; the cloud layer acts as the global brain to complete final semantic verification and result archiving, realizing a closed-loop detection process. In the reverse optimization process, after the cloud layer completes iterative model training based on the full dataset, it distributes the optimized model parameters and domain knowledge to the edge layer; after the edge layer completes local model updates, it distills and generates lightweight parameters adapted to the edge layer, which are then simultaneously distributed to the edge layer; the edge layer completes iterative iteration of the local inference model, realizing continuous self-iteration of "detection-optimization-re-detection".

[0025] For example, such as Figure 3As shown, the edge layer is equipped with a drone platform and supporting hardware, serving as the core for data acquisition and front-end inference, enabling rapid response in acquisition, preprocessing, and preliminary detection. The edge-side data acquisition unit includes a visible light / infrared dual-light sensor (≥20 million pixels, supporting simultaneous shooting) and a GNSS positioning module (supporting GPS / BeiDou dual-mode, static positioning accuracy ≤1m). During acquisition, it simultaneously records image data, latitude and longitude coordinates, shooting altitude, attitude angle, and other metadata, embedding them into the image file via EXIF ​​information. The edge-side flight path planning unit incorporates a scene complexity assessment submodule and a parameter dynamic adjustment submodule, analyzing terrain, target distribution, and environmental data in real time, automatically generating and adjusting flight parameters (altitude, heading overlap, etc.), and supporting real-time reception of edge feedback commands via 4G / 5G. The edge-side preliminary processing unit deploys a basic preprocessing module (improved adaptive histogram equalization algorithm) and a lightweight student model optimized by knowledge distillation, rapidly enhancing, standardizing, and detecting suspected defects in the acquired images. Detection results (including image fragments, coordinate information, and confidence levels) are temporarily stored in a local cache. In the edge data transmission unit, it supports near real-time wireless transmission of suspected defect data via 4G / 5G or batch offline export via USB interface, compatible with mainstream drone data management software (such as DJI Terra). For the edge layer, acting as an intermediate hub between the edge and cloud, it balances real-time performance and detection accuracy, reducing cloud transmission pressure. In the edge-side precision verification unit, a teacher model (a simplified version of a multimodal large model) and a simplified version of the multi-scale space-frequency feature fusion (Space-Frequency Separated Cross-Attention Swin Transformer, SF-SCST) network are deployed. It receives suspected defect data uploaded from the edge, performs secondary precision detection and feature optimization, and filters out valid defect data with a confidence level ≥ 0.7. In the edge-side data interaction unit, valid defect data is uploaded to the cloud, and model parameter update instructions and domain knowledge supplementary information from the cloud are received and synchronized to the edge to optimize front-end detection performance. For the cloud layer, acting as the "brain" of the system architecture, it is responsible for model training, knowledge storage, and global data processing. In the cloud-based model training unit, a complete multi-scale spatial-frequency feature fusion network (SF-SCST), a large language model (LLM), and a domain knowledge-guided adversarial generative network are deployed. Iterative model training, few-shot generation, and cross-modal semantic reasoning are performed based on the full dataset. In the cloud-based knowledge storage unit, a domain knowledge base (containing defect terminology, morphological descriptions, and location rules for scenarios such as photovoltaics and water conservancy), a model parameter database (storing the weights of each trained version of the model), and a detection result database (structured storage of defect types, coordinates, confidence levels, and other information) are constructed.In the decision output unit at the cloud layer, the final detection results are sent to the edge and end sides, and simultaneously synchronized to the GIS platform for visualization, supporting the generation of inspection reports and defect distribution heat maps.

[0026] Figure 1 This is an optional flowchart of a rapid detection method for small targets in low-altitude remote sensing provided by an embodiment of the present invention. Figure 1 The method may include, but is not limited to, steps S100 to S600: Step S100: Image enhancement and standardization are performed on the acquired images to obtain standardized images; Step S200: Visual inspection of standardized images is performed using a lightweight student model to generate first suspected defect data, and the first suspected defect data is associated with the location information; Step S300: Through a multi-scale spatial-frequency feature fusion network, feature optimization is performed on the first suspected defect data to obtain the second suspected defect data; the second suspected defect data includes visual confidence. Step S400: Using the teacher model, perform a second detection on the second suspected defect data to obtain the third suspected defect data; the third suspected defect data includes semantic verification tags; Step S500: Based on semantic verification tags and visual confidence, the third suspected defect data is filtered to obtain valid defect data; Step S600: Semantic verification of valid defect data is performed using a large language model and domain knowledge base, target detection results are output, and the target detection results are associated with the location information to obtain structured detection records.

[0027] In step S100 of some embodiments, the scene is perceived, corresponding images are acquired, and the acquired images are enhanced and standardized to obtain standardized images. For example, taking a UAV as an example, the UAV loads terrain data and historical target distribution information of the inspection area, while simultaneously acquiring environmental parameters such as current illumination, cloud cover, and wind speed in real time. Based on a three-dimensional assessment of "terrain complexity, target density, and environmental interference," parameters such as flight altitude and heading overlap are automatically generated or adjusted (e.g., 30-50m altitude for photovoltaic panel inspection). Visible light / infrared images are simultaneously acquired using dual optical sensors, and information such as shooting altitude and attitude angle are embedded into the image EXIF ​​field. GPS / BeiDou dual-mode positioning is used to bind latitude and longitude coordinates to the acquired images, laying the foundation for subsequent spatial positioning. After acquiring the images, improved adaptive histogram equalization (CLAHE) is used for image enhancement, and the enhanced images are standardized to lay the foundation for subsequent detection.

[0028] In some embodiments, step S100 may include, but is not limited to, steps S110 to S120: Step S110: The edge discrimination between small targets and background in the acquired image is enhanced by an adaptive histogram equalization algorithm to obtain an enhanced image. Step S120: The enhanced image is scaled and normalized to obtain a standardized image.

[0029] Before step S110 in some embodiments, the system architecture dynamically optimizes the edge parameters based on the task scenario and real-time feedback. Taking the UAV edge as an example, the UAV can collect data through adaptive flight path planning (as shown in Table 1). First, the scenario complexity is classified and quantified, establishing a multi-dimensional evaluation index system including terrain complexity, target distribution density, and environmental interference intensity. Terrain complexity is calculated using the local variance of the digital elevation model (DEM). Target distribution density is modeled based on the statistical distribution of historical detection results; environmental interference intensity comprehensively considers illumination uniformity, cloud and fog coverage ratio, and wind disturbance level. The flight path parameter generation module automatically calculates and recommends the optimal flight altitude, speed, heading overlap rate, and lateral overlap rate based on the evaluation results. Among them, flight altitude is negatively correlated with target distribution density and terrain complexity, while overlap rate is positively correlated with environmental interference intensity and task accuracy requirements. The real-time dynamic adjustment module continuously monitors environmental parameters such as illumination intensity, wind speed, and wind direction during flight.

[0030] Table 1. Decision Table of Key Parameters for Adaptive Route Planning

[0031] In step S110 of some embodiments, a contrast-limited adaptive histogram equalization algorithm is used to divide the acquired image into local windows of 16×16 pixels, with a cropping limit set to 0.03, thereby enhancing the edge distinction between small targets and the background. The calculation formula is as follows: ; In the formula, Indicates an enhanced image; Represents the pixels of the acquired image gray levels within the local window The cumulative distribution function; Indicates the total number of pixels within the window; Indicates the number of gray levels; This represents the cropping limit. The cropping limit value was experimentally determined within the range of 0.01 to 0.05 through grid search, and 0.03 was ultimately selected as the optimal value. Experiments show that too small a cropping limit leads to insufficient local contrast enhancement and decreased edge discrimination; too large a cropping limit easily introduces noise, affecting the integrity of the target's shape. At the selected value, the algorithm can effectively enhance the grayscale difference between small targets and the background.

[0032] In step S120 of some embodiments, the enhanced image is scaled to 640×640 pixels (while maintaining the aspect ratio) and the pixel values ​​are normalized to the [0,1] range to obtain a normalized image.

[0033] In step S200 of some embodiments, a lightweight student model optimized by distillation (with SF-SCST network channels cut to 1 / 2 of the original) is invoked to perform rapid inference on the standardized image and generate the first suspected defect data. This may include, but is not limited to, generating defect bounding boxes, defect types, and defect confidence scores. The first suspected defect data is then associated with the corresponding GNSS coordinates, temporarily stored locally on the edge, and synchronously uploaded to the edge device.

[0034] In step S300 of some embodiments, the edge end of the system architecture transmits suspected defect image fragments, GNSS coordinates, and defect confidence information uploaded by the receiving end via 4G / 5G wireless transmission, and uses a simplified multi-scale space-frequency feature fusion network to perform feature optimization on these first suspected defect data. The multi-scale space-frequency feature fusion network includes a wavelet decomposition and feature stitching module (WDFC), a feature channel space-frequency decomposition and fusion module (SFDF), and a cross-self-attention SwinTransformer module (CS-Swin). Optionally, as... Figure 4 As shown, a simplified multi-scale spatial-frequency feature fusion network is used to extract spatial / frequency features from the image, enhance the representation of small target details, and take standardized image as input. The image is processed by three-level modules: WDFC, SFDF, and CS-Swin. The high-frequency details of small targets are locked in sequence, spatial-frequency dual-domain features are fused, and the deep and shallow feature correlations are established. Finally, the structured visual dimension detection results such as defect coordinates, visual confidence, and category probability are output, which is the second suspected defect data.

[0035] In some embodiments, step S300 may include, but is not limited to, steps S310 to S370: Step S310: The first suspected defect data is subjected to wavelet transform through the wavelet decomposition and feature splicing module to obtain the frequency domain subband. Step S320: The frequency domain subband and the first suspected defect data are spliced ​​together and channel fusion and dimensionality reduction are performed to obtain the first enhanced feature; Step S330: The first enhanced feature is decomposed through the feature channel space-frequency decomposition and fusion module to obtain the initial spatial domain feature and the initial frequency domain feature. Step S340: Using a non-uniform interval feature selection strategy, the spatial context information of the initial spatial features is reorganized to obtain the target spatial features. Step S350: Perform wavelet transform on the initial frequency domain features to obtain the target frequency domain features; Step S360: The target spatial domain features and the target frequency domain features are fused to obtain the second enhanced feature; Step S370: Through the cross-self-attention Swin Transformer module, the shallow features and deep features of the second enhanced feature are strongly correlated to obtain the second suspected defect data.

[0036] In steps S310 to S320 of some embodiments, the wavelet decomposition and feature stitching module in the multi-scale space-frequency feature fusion network is used to perform a two-dimensional discrete wavelet transform on the feature map of the input first suspected defect data. Decomposed into low-frequency subbands (Approximate information) and three high-frequency sub-bands (Horizontal details) (Vertical details) (Diagonal details). The calculation formula is: To maintain spatial dimensionality consistency, bilinear interpolation upsampling was performed on each frequency domain sub-band to the original feature map size. These four frequency domain sub-bands were then compared with the original spatial domain feature map. The feature maps of the first suspected defective data are concatenated along the channel dimension, and then fused and reduced in dimensionality using a 1×1 convolutional layer to obtain the first enhanced feature. The calculation formula is as follows:

[0037] In the formula, Indicates a splicing operation; This represents a 1×1 convolution operation.

[0038] By employing simple two-dimensional discrete wavelet transform calculations, high-frequency information such as edges and textures can be effectively captured. Rich high-frequency details are preserved even in shallow network layers, making it suitable for small object detection tasks. Wavelet decomposition and feature concatenation modules enable the network to process spatial semantic information and frequency detail information in parallel. Compared to the baseline model mAP@50, after introducing the WDFC module, the proposed model's mAP@50 improves by 1.4% and 1.9% respectively on the public datasets UAV-DA and VisDrone (as shown in Table 2). mAP@50 is a core metric used in object detection to evaluate the overall performance of a model.

[0039] Table 2. Performance improvements of the WDFC module on public datasets

[0040] In steps S330 to S360 of some embodiments, the feature channel space-frequency decomposition and fusion module in the multi-scale space-frequency feature fusion network is used to divide the input first enhanced feature into initial spatial domain features and initial frequency domain features in the channel dimension. In the spatial domain branch, a non-uniform interval feature selection strategy is used to reorganize the spatial context information of the initial spatial domain features, thereby capturing the overall structure of the target and obtaining the target spatial domain features. In the frequency domain branch, a two-dimensional discrete wavelet transform is applied again to the initial frequency domain features to enhance the extraction of high-frequency detail features, thereby obtaining the target frequency domain features. Finally, the target spatial domain features and the target frequency domain features are fused to obtain the second enhanced feature. The calculation formula is as follows: ; In the formula, Indicates the second enhancement feature; Indicates the initial spatial characteristics; Indicates the initial frequency domain characteristics; Represents a two-dimensional discrete wavelet transform; , This indicates the corresponding processing function; , This represents the learnable fusion weights. The fusion weights are initialized to equal values ​​and dynamically updated during training via gradient backpropagation to adaptively balance the contributions of spatial and frequency domain features. Spatial features focus on the macroscopic shape and spatial context of the target, while frequency domain features focus on local high-frequency information such as edges and textures. The two complement each other effectively at the feature level, jointly enhancing the representation ability of small targets.

[0041] In step S370 of some embodiments, the cross-self-attention Swin Transformer module (e.g., in the multi-scale space-frequency feature fusion network) is utilized. Figure 5As shown, a strong correlation is established between shallow detail features and deep semantic features. The Swin Transformer module with cross-attention first aligns the channel numbers of the shallow and deep features of the second enhancement feature using a 1×1 convolution, ensuring their semantic comparability. Shallow features mainly contain local detail information such as edges and corners, while deep features encode more abstract target categories and structural semantics. The two have a natural hierarchical correlation and complementarity in detection tasks. The calculation of cross-attention can be simplified as follows: ; In the formula, This indicates a query derived from deep features. Keys represent features derived from shallow layers. This represents the value derived from shallow features. Represents the Softmax function; Indicates the scaling factor; Indicates the transpose operation; This indicates cross attention; In step S400 of some embodiments, a simplified version of the multimodal large model (teacher model) is invoked, and the suspected defects are subjected to secondary semantic verification in combination with domain knowledge. The second suspected defect data output by SF-SCST and the inspection domain knowledge prompt words are used as dual inputs. The visual and semantic features are encoded by dual encoders. The defect professional definition is matched through cross-modal attention, and the third suspected defect data including semantic matching degree, category judgment result and false detection exclusion mark is finally output.

[0042] In some embodiments, step S400 may include, but is not limited to, steps S410 to S450: Step S410: Input the second suspected defect data and the preset domain knowledge prompts into the teacher model; Step S420: The second suspected defect data is feature-encoded by a visual encoder to obtain visual features; Step S430: Encode the domain knowledge prompts using a semantic encoder to obtain semantic features; Step S440: Through a cross-modal attention mechanism, visual features are matched with semantic features to output semantic verification tags; Step S450: Associate the semantic verification tag with the second suspected defect data to generate the third suspected defect data.

[0043] In step S410 of some embodiments, the second suspected defect data and domain knowledge prompts from a preset domain knowledge base are used as joint inputs and fed into the teacher model, completing the context switching from single-modal visual detection to multi-modal joint reasoning. The teacher model is a simplified version of a large multi-modal model.

[0044] In steps S420 to S430 of some embodiments, the teacher model uses its built-in visual encoder and semantic encoder to encode features of the second suspected defect data and domain knowledge prompts, respectively, to obtain visual features and semantic features. Optionally, the built-in visual encoder of the teacher model is usually a pre-trained convolutional neural network or VisionTransformer, and the built-in semantic encoder is usually a language model based on the Transformer architecture.

[0045] In step S440 of some embodiments, the teacher model matches the encoded visual features with semantic features through its cross-modal attention mechanism to output a semantic verification label. This semantic verification label may include a semantic matching degree, a category determination result, and a false positive exclusion label. Optionally, the cross-modal attention mechanism uses visual features as queries and semantic features as keys and values ​​to calculate the attention distribution of visual content in the semantic concept space. This mechanism evaluates the degree of matching between visual evidence and professional definitions, ultimately outputting a semantic dimension verification label.

[0046] In step S450 of some embodiments, the semantic verification tag representing the semantic level verification result is structurally associated and encapsulated with the second suspected defect data (including coordinates, visual confidence, etc.) representing the visual level detection result, generating a new, more comprehensive data object, namely the third suspected defect data. This completes the upgrade from visual perception result to visual-semantic joint decision result.

[0047] For example, visual-linguistic feature alignment and fusion are performed using a visual encoder, a language encoder, and a cross-modal fusion module within the teacher model. During operation, the visual encoder extracts image features from the second suspected defect data. Meanwhile, the language encoder processes text information related to the scene and generates text features. The text information is domain-specific prompt text, constructed based on a professional domain knowledge base for the task scenario, including standard terminology, typical descriptive sentence structures, and contextual association rules. For example, in photovoltaic panel inspection, the prompt text includes professional terms such as "hot spot," "microcrack," and "stain obstruction," along with descriptions of their common locations and shapes. The text prompt length is typically controlled between 16 and 32 words, and semantic density is adjusted by the proportion of noun entities and the use of logical conjunctions to balance information content with noise interference. The cross-modal fusion module employs a Transformer-based attention mechanism to achieve alignment and deep fusion of visual and linguistic features. The calculation formula is as follows: ; In the formula, Indicates cross-modal fusion operation; The fused features contain both visual details and semantic priors and contextual knowledge, thus enabling more accurate identification of small targets that are difficult to judge visually.

[0048] In step S500 of some embodiments, the results of SF-SCST are combined with the results of the teacher model (including labeled samples containing visual dimension detection results and semantic dimension verification results), false detection samples with a confidence level lower than 0.7 are removed, valid defect data are retained, and the valid defect data are uploaded to the cloud.

[0049] In step S600 of some embodiments, in the system architecture, the cloud receives valid defect data uploaded by the edge device, which includes structured information such as image features, coordinates, and type. The cloud calls a Large Language Model (LLM) and combines professional terms, defect morphology descriptions, and location rules from the domain knowledge base to perform semantic reasoning on the features of the valid defect data, correcting misjudgments in blurred / occluded scenes, and outputting the final detection result. The final detection result (outputting defect bounding boxes, defect types, confidence scores, and other suspected defect result data) is then bound to the GNSS coordinates corresponding to the acquired image data to generate a structured detection record with spatial attributes. The structured detection record is stored in the cloud database, supporting retrieval by time, region, and defect type.

[0050] In some embodiments, a method for rapid detection of small targets in low-altitude remote sensing further includes: training a multi-scale spatial-frequency feature fusion network based on full-volume valid data to generate optimized network parameters; updating the teacher model based on the optimized network parameters; generating lightweight parameters by distillation using knowledge distillation techniques based on the updated teacher model; and updating the lightweight student model based on the lightweight parameters; wherein the full-volume valid data includes real defect samples and synthetic defect samples.

[0051] For example, in the cloud, a complete multi-scale spatial-frequency feature fusion network is iteratively trained using all available data (including real and synthetic defect samples) to update the SF-SCST model weights and generate optimized network parameters. The cloud then distributes the optimized network parameters to the edge device to update the teacher model. The cloud also stores the iterated SF-SCST model weights in a parameter library and adds new defect description rules to the domain knowledge base. The edge device uses the updated teacher model and, through knowledge distillation, generates the latest lightweight parameters for a fixed lightweight student model architecture. The edge device distributes these latest lightweight parameters to the device, which receives and loads them to update its local student model. The device uses the updated model for inference, and the resulting data is fed back to the cloud to drive the next round of model training and optimization, forming a continuous iterative closed loop.

[0052] Alternatively, efficient inference can be achieved through a cross-modal knowledge distillation framework, which guides a lightweight, purely visual detection model (student model) to learn from a powerful, multimodal large model (teacher model). The student model only receives images and learns by minimizing the KL divergence between the student model's output and the teacher's "soft labels." ), and minimize the mean squared error between the intermediate feature maps of the student model and the corresponding intermediate feature maps of the teacher model ( The learning process is conducted using the following formula: ; In the formula, This represents the predicted distribution of the teacher model; This represents the predicted distribution of the student model; Indicates the balance coefficient; , Indicates intermediate layer features; This represents the total loss from knowledge distillation.

[0053] In some embodiments, a method for rapid detection of small targets in low-altitude remote sensing further includes: when the number of samples in the full set of valid data is insufficient, invoking an adversarial generative network to generate synthetic defect samples; obtaining a structural similarity loss based on the structural similarity between the synthetic defect samples and the real defect samples; obtaining the synthetic feature embedding vector of the synthetic defect samples; obtaining the real feature embedding vector of the real defect samples; obtaining a domain consistency loss based on the cosine distance between the synthetic feature embedding vector and the real feature embedding vector; and training the adversarial generative network based on the adversarial loss, the structural similarity loss, and the domain consistency loss.

[0054] For example, when the system architecture iterates models in the cloud, if the number of samples for a certain type of defect is insufficient (e.g., <100), highly realistic simulated samples are generated based on domain knowledge embedding vectors to alleviate the class imbalance problem. Optionally, adversarial generative networks and domain knowledge can be used to guide the completion of small samples. Furthermore, traditional adversarial loss... Based on this, structural similarity loss is introduced. Domain consistency loss Then, the total loss is calculated, and the adversarial generative network is trained. The total loss function is then: ; In the formula, This represents a real defect sample; Indicates a synthetic defect sample; Used to constrain the structural similarity between synthetic defect sample images and real defect sample images; Used to constrain the features of synthetic defect sample images through a pre-trained domain classifier. Features of real defect sample images The distribution is consistent in the feature space; and To balance the weights, the weights for each loss term are determined using a grid search based on its contribution to the validation set, and are ultimately set as follows: , , The domain classifier employs a three-layer fully connected network, pre-trained on a labeled dataset consisting of real defect samples and normal samples without defects. The domain consistency loss is specifically calculated as the cosine distance between the features of the synthetic defect sample image and the features of the real defect sample image in the feature space of the last hidden layer of the classifier.

[0055] This invention also provides a rapid detection device for small targets in low-altitude remote sensing, which can implement the above-mentioned rapid detection method for small targets in low-altitude remote sensing. The device includes: The preprocessing module is used to enhance and standardize the acquired images to obtain standardized images; The lightweight visual inspection module is used to perform visual inspection on standardized images using a lightweight student model, generate first suspected defect data, and associate the first suspected defect data with the location information. The multi-scale feature optimization module is used to optimize the features of the first suspected defect data through a multi-scale spatial-frequency feature fusion network to obtain the second suspected defect data; the second suspected defect data includes visual confidence. The cross-modal verification module is used to perform secondary detection on the second suspected defect data through the teacher model to obtain the third suspected defect data; the third suspected defect data includes semantic verification tags; The filtering module is used to filter the third suspected defect data based on semantic verification tags and visual confidence to obtain valid defect data; The semantic verification module is used to perform semantic verification on valid defect data through a large language model and domain knowledge base, output target detection results, and associate the target detection results with the location information to obtain structured detection records.

[0056] It is understood that the content of the above method embodiments is applicable to the present device embodiments. The specific functions implemented by the present device embodiments are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.

[0057] This invention also provides an electronic device, which includes a processor and a memory. The memory stores a computer program, and the processor executes the computer program to implement the above-described method. This electronic device can be any smart terminal, including a tablet computer, an in-vehicle computer, or similar device.

[0058] It is understood that the content of the above method embodiments is applicable to this device embodiment. The specific functions implemented by this device embodiment are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.

[0059] refer to Figure 6 , Figure 6 The hardware structure of an electronic device according to another embodiment is illustrated. The electronic device includes: The processor 701 can be implemented using a general-purpose CPU (Central Processing Unit), microprocessor, application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of the present invention. The memory 702 can be implemented as a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM). The memory 702 can store the operating system and other application programs. When the technical solutions provided in the embodiments of this specification are implemented through software or firmware, the relevant program code is stored in the memory 702 and is called and executed by the processor 701. The input / output interface 703 is used to implement information input and output; The communication interface 704 is used to enable communication and interaction between this device and other devices. Communication can be achieved through wired means (such as USB, Ethernet cable, etc.) or wireless means (such as mobile network, WIFI, Bluetooth, etc.). Bus 705 transmits information between various components of the device (e.g., processor 701, memory 702, input / output interface 703, and communication interface 704); The processor 701, memory 702, input / output interface 703, and communication interface 704 are connected to each other within the device via bus 705.

[0060] This invention also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the above-described method.

[0061] It is understood that the content of the above method embodiments is applicable to this storage medium embodiment. The specific functions implemented in this storage medium embodiment are the same as those in the above method embodiments, and the beneficial effects achieved are also the same as those achieved in the above method embodiments.

[0062] This invention also provides a computer program product or computer program that includes computer instructions stored in a computer-readable storage medium. A processor of a computer device can read the computer instructions from the computer-readable storage medium and execute the computer instructions to cause the computer device to perform the aforementioned method.

[0063] In summary, the low-altitude remote sensing method and related equipment for rapid detection of small targets according to embodiments of the present invention have the following advantages: 1. This invention dynamically optimizes parameters such as flight altitude and overlap rate by quantifying terrain complexity, target distribution density, and environmental interference intensity. Simultaneously, it combines an improved CLAHE algorithm to enhance the edge differentiation between small targets and the background. Compared to a fixed-flight acquisition mode, this method specifically improves the clarity of small target details and suppresses background interference, providing higher-quality input data for subsequent detection and reducing the risk of missed detections from the source.

[0064] 2. This invention employs a multi-scale spatial-frequency feature fusion network comprising three modules (WDFC, SFDF, and CS-Swin) to simultaneously capture spatial semantic information and frequency-domain detailed features. It also supports the collaborative fusion of multi-source data, including visible light and infrared, overcoming the limitations of existing technologies that focus only on single spatial features or single modal data. Through a cross-attention mechanism with strong correlation between deep and shallow features, it can more comprehensively capture the subtle features of small targets, significantly improving detection performance in scenarios with extremely small target scales and complex environmental interference.

[0065] 3. This invention introduces domain knowledge guidance into few-shot learning, generating highly realistic and targeted defect samples through adversarial generative networks, effectively alleviating the class imbalance problem. Simultaneously, it integrates a large language model to construct a cross-modal semantic understanding framework, injecting domain knowledge priors into visual reasoning. Compared to simple data augmentation and pure visual feature matching, it can more accurately identify targets in challenging scenarios such as blurriness and occlusion, significantly reducing false detection rates and improving detection robustness.

[0066] 4. This invention employs a three-layer collaborative architecture: real-time inference on the device side, precise verification at the edge, and model training in the cloud. It utilizes knowledge distillation technology to achieve semantic transfer from a large multimodal model to a lightweight model. Deploying a lightweight model on the device side meets real-time detection requirements, while secondary verification at the edge improves accuracy. The cloud coordinates and optimizes model performance, solving the problem of insufficient computing power on a single terminal and avoiding transmission delays associated with centralized cloud processing. This achieves a perfect balance between detection accuracy and real-time performance, making it more suitable for large-scale inspection engineering applications.

[0067] 5. The data acquisition module of this invention supports multi-sensor expansion, and the detection model can be quickly adapted to different inspection scenarios such as photovoltaics, water conservancy, power, and transportation by adjusting the domain knowledge base and model parameters. The "cloud-edge-device" architecture is compatible with mainstream drone hardware and GIS platforms, and can be integrated into existing workflows without system reconstruction. Compared with existing technologies that are highly targeted but have weak scalability, this invention has stronger universality, lower maintenance costs, and can better meet the diverse detection needs of different industries.

[0068] In some alternative embodiments, the functions / operations mentioned in the block diagrams may not occur in the order shown in the operation diagrams. For example, depending on the functions / operations involved, two consecutively shown blocks may actually be executed substantially simultaneously, or the blocks may sometimes be executed in reverse order. Furthermore, the embodiments presented and described in the flowcharts of this invention are provided by way of example to provide a more comprehensive understanding of the technology. The disclosed methods are not limited to the operations and logic flows presented herein. Alternative embodiments are contemplated in which the order of various operations is altered and sub-operations described as part of a larger operation are executed independently.

[0069] Furthermore, although the invention has been described in the context of functional modules, it should be understood that, unless otherwise stated, one or more of the described functions and / or features may be integrated into a single physical device and / or software module, or one or more functions and / or features may be implemented in a separate physical device or software module. It is also understood that a detailed discussion of the actual implementation of each module is unnecessary for understanding the invention. Rather, given the properties, functions, and internal relationships of the various functional modules in the apparatus disclosed herein, the actual implementation of the module will be understood within the scope of conventional skill of an engineer. Therefore, those skilled in the art can implement the invention as set forth in the claims using ordinary techniques without excessive experimentation. It is also understood that the specific concepts disclosed are merely illustrative and not intended to limit the scope of the invention, which is determined by the full scope of the appended claims and their equivalents.

[0070] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this invention, essentially, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0071] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a sequenced list of executable instructions for implementing logical functions, and can be embodied in any computer-readable medium for use by, or in conjunction with, an instruction execution system, apparatus, or device (such as a computer-based system, a processor-including system, or other system that can fetch and execute instructions from, an instruction execution system, apparatus, or device). For the purposes of this specification, "computer-readable medium" can be any means that can contain, store, communicate, propagate, or transmit programs for use by, or in conjunction with, an instruction execution system, apparatus, or device.

[0072] More specific examples of computer-readable media (a non-exhaustive list) include: electrical connections (electronic devices) having one or more wires, portable computer disk drives (magnetic devices), random access memory (RAM), read-only memory (ROM), erasable and editable read-only memory (EPROM or flash memory), fiber optic devices, and portable optical disc read-only memory (CDROM). Furthermore, computer-readable media can even be paper or other suitable media on which the program can be printed, because the program can be obtained electronically, for example, by optically scanning the paper or other medium, followed by editing, interpreting, or otherwise processing as necessary, and then stored in computer memory.

[0073] It should be understood that various parts of the present invention can be implemented in hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented in software or firmware stored in memory and executed by a suitable instruction execution system. For example, if implemented in hardware, as in another embodiment, it can be implemented using any one or a combination of the following techniques known in the art: discrete logic circuits having logic gates for implementing logical functions on data signals, application-specific integrated circuits (ASICs) having suitable combinational logic gates, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), etc.

[0074] In the description of this specification, references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of the invention. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples.

[0075] Although embodiments of the invention have been shown and described, those skilled in the art will understand that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the claims and their equivalents.

[0076] The above is a detailed description of the preferred embodiments of the present invention. However, the present invention is not limited to the embodiments described. Those skilled in the art can make various equivalent modifications or substitutions without departing from the spirit of the present invention. All such equivalent modifications or substitutions are included within the scope defined by the claims of the present invention.

Claims

1. A method for rapid detection of small targets in low-altitude remote sensing, characterized in that, Includes the following steps: The acquired images are enhanced and standardized to obtain standardized images; A lightweight student model is used to perform visual inspection on the standardized image to generate first suspected defect data, and the first suspected defect data is associated with the location information. The first suspected defect data is optimized using a multi-scale spatial-frequency feature fusion network to obtain the second suspected defect data; the second suspected defect data includes visual confidence. The teacher model is used to perform a second detection on the second suspected defect data to obtain the third suspected defect data; The third suspected defect data includes semantic verification tags; Based on the semantic verification markers and the visual confidence level, the third suspected defect data is filtered to obtain valid defect data; Using a large language model and a domain knowledge base, the valid defect data is semantically verified, the target detection results are output, and the target detection results are associated with the location information to obtain a structured detection record.

2. The method according to claim 1, characterized in that, The acquired images are then enhanced and standardized to obtain standardized images. Includes the following steps: An enhanced image is obtained by strengthening the edge distinction between small targets and the background in the acquired image through an adaptive histogram equalization algorithm. The enhanced image is scaled and normalized to obtain the standardized image.

3. The method according to claim 1, characterized in that, The step of optimizing the first suspected defect data using a multi-scale space-frequency feature fusion network to obtain the second suspected defect data includes the following steps: The first suspected defect data is subjected to wavelet transform by the wavelet decomposition and feature splicing module to obtain frequency domain sub-bands; The frequency domain sub-band is spliced ​​with the first suspected defect data, and channel fusion and dimensionality reduction are performed to obtain the first enhanced feature. The first enhanced feature is decomposed by the feature channel space-frequency decomposition and fusion module to obtain the initial spatial domain feature and the initial frequency domain feature. A non-uniform interval feature selection strategy is used to reorganize the spatial context information of the initial spatial features to obtain the target spatial features. The initial frequency domain features are subjected to wavelet transform to obtain the target frequency domain features; The target spatial domain features and the target frequency domain features are fused to obtain the second enhanced feature; By using the cross-self-attention Swin Transformer module, the shallow and deep features of the second enhanced feature are strongly correlated to obtain the second suspected defect data.

4. The method according to claim 1, characterized in that, The process of using a teacher model to perform a second detection on the second suspected defect data to obtain the third suspected defect data includes the following steps: The second suspected defect data and the preset domain knowledge prompts are input into the teacher model; The second suspected defect data is feature-encoded using a visual encoder to obtain visual features; Semantic features are obtained by feature encoding the domain knowledge prompts using a semantic encoder. The visual features are matched with the semantic features using a cross-modal attention mechanism, and the semantic verification tag is output. The semantic verification tag is associated with the second suspected defect data to generate the third suspected defect data.

5. The method according to claim 1, characterized in that, The method further includes the following steps: Based on the full amount of valid data, the multi-scale space-frequency feature fusion network is trained to generate optimized network parameters; The teacher model is updated based on the optimized network parameters; Based on the updated teacher model, lightweight parameters are generated through knowledge distillation techniques. The lightweight student model is updated based on the lightweight parameters. The full set of valid data includes real defect samples and synthetic defect samples.

6. The method according to claim 5, characterized in that, The method further includes the following steps: When it is detected that the number of samples in the full amount of valid data is insufficient, the adversarial generative network is invoked to generate the synthetic defective samples. Based on the structural similarity between the synthesized defect sample and the real defect sample, obtain the structural similarity loss; Obtain the synthetic feature embedding vector of the synthetic defect sample; Obtain the true feature embedding vector of the real defect sample; The domain consistency loss is obtained based on the cosine distance between the synthesized feature embedding vector and the real feature embedding vector. The adversarial generative network is trained based on the adversarial loss, the structural similarity loss, and the domain consistency loss.

7. A rapid detection device for small targets in low-altitude remote sensing, characterized in that, include: The preprocessing module is used to enhance and standardize the acquired images to obtain standardized images; The lightweight visual inspection module is used to perform visual inspection on the standardized image using a lightweight student model, generate first suspected defect data, and associate the first suspected defect data with the positioning information. A multi-scale feature optimization module is used to optimize the features of the first suspected defect data through a multi-scale space-frequency feature fusion network to obtain the second suspected defect data; the second suspected defect data includes visual confidence. The cross-modal verification module is used to perform a second detection on the second suspected defect data through the teacher model to obtain the third suspected defect data; The third suspected defect data includes semantic verification tags; The filtering module is used to filter the third suspected defect data based on the semantic verification marker and the visual confidence level to obtain valid defect data; The semantic verification module is used to perform semantic verification on the valid defect data through a large language model and a domain knowledge base, output the target detection result, and associate the target detection result with the location information to obtain a structured detection record.

8. An electronic device, characterized in that, Including the processor and memory; The memory is used to store programs; The processor executes the program to implement the method as described in any one of claims 1 to 6.

9. A computer-readable storage medium, characterized in that, The storage medium stores a program that is executed by a processor to implement the method as described in any one of claims 1 to 6.

10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the method as described in any one of claims 1 to 6.