Man-machine collaborative screening method for interpretable features

By employing a human-machine collaborative screening method with interpretable features in UAV remote sensing images, combined with deep learning models and multi-dimensional screening, the problems of missed detection and false detection in black box screening of UAV images are solved, and fast and accurate target identification and verification are achieved.

CN120853177BActive Publication Date: 2025-11-25TIANJIN ZERO ONE INTELLIGENT TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511357756.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-09-23
Publication Date
2025-11-25
Estimated Expiration
2045-09-23

AI Technical Summary

Technical Problem

Existing technologies struggle to quickly and accurately identify black boxes from low-altitude remote sensing images from drones, resulting in missed detections and false detections, and manual verification is inefficient.

Method used

A human-machine collaborative screening method with interpretable features is adopted. Parallel detection is performed using a pre-trained YOLOv5-small deep learning model. Multi-dimensional screening is performed by combining size and color features. A visual table is constructed for manual confirmation to achieve target verification.

Benefits of technology

It reduced the false negative rate and the false positive rate, improved the detection speed and accuracy, reduced the labor intensity of manual verification, and achieved rapid and reliable target screening.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120853177B_ABST
    Figure CN120853177B_ABST
Patent Text Reader

Abstract

The application provides a man-machine collaborative screening method with interpretable features, comprising collecting remote sensing images of a search area; pre-processing the remote sensing images to obtain a cropped unit; performing parallel detection on the cropped unit through a pre-trained detection model, and outputting a detection result; performing non-maximum suppression on a first detection box of the detection result to obtain a second detection box; performing table quick screening on the second detection box to obtain a third detection box; and performing original image fine screening on the third detection box to label a real search result. The application has the beneficial effects of fusing three screening dimensions, forming progressive logic, and ensuring that the target verification result is reliable; the screening logic is based on physical characteristics, is easy to adjust, and can be popularized to other unmanned aerial vehicle micro-target search tasks; the size and color characteristics are combined to batch exclude obvious false detections and reduce the number of artificial verification targets.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application belongs to the technical field of artificial intelligence image recognition processing, in particular to a man-machine collaborative screening method with interpretable features. BACKGROUND

[0002] In civil aviation accident investigation, the black box (flight recorder) is the core basis for restoring the accident process and analyzing the accident cause. Its rapid positioning and recovery is crucial to the efficiency of accident investigation. Using unmanned aerial vehicles to collect low-altitude images of large-scale accident sites for rapid search of black boxes is also a crucial task. Currently, using unmanned aerial vehicles to collect low-altitude remote sensing images of accident sites has become a mainstream technical means for searching for black boxes in large-scale and complex environments, but this technical route still faces many key problems that need to be solved in practical application. First, the black box only occupies tens to hundreds of pixels in the aerial photograph of the unmanned aerial vehicle, and the visual features are blurred. Existing target recognition technologies cannot capture effective features and cannot accurately locate the target. Second, the terrain and lighting of the accident site are complex, and there are a large number of interference objects similar in color and shape to the black box. Existing technologies are difficult to exclude interference and affect detection accuracy. Third, although existing detection algorithms optimize performance, when processing massive data, the black box features are easily misjudged as similar objects, resulting in a large number of invalid false detection results. Finally, existing technologies rely on manual verification of detection results frame by frame, which is slow and labor-intensive, and cannot meet the urgent needs of accident investigation.

[0003] Therefore, how to design a system that can assist humans in quickly and accurately screening out real targets from massive detection results is a technical problem that needs to be solved in the current field of black box unmanned aerial vehicle search. SUMMARY

[0004] To solve the above technical problems, the present application provides a man-machine collaborative screening method with interpretable features, which is especially suitable for quickly searching for scattered black boxes in accidents.

[0005] The technical solution adopted by the present application is to provide a man-machine collaborative screening method with interpretable features, comprising the following steps:

[0006] Collecting remote sensing images of the search area;

[0007] Pretreating the remote sensing images to obtain a cropped unit;

[0008] Parallelly detecting the cropped unit through a pre-trained detection model and outputting a detection result;

[0009] Sorting all first detection boxes in the detection result from high to low according to confidence;

[0010] Calculate the intersection over union of adjacent first bounding boxes after sorting, and remove the first bounding boxes with an intersection over union greater than a preset overlap threshold to screen out second bounding boxes;

[0011] Convert the local pixel coordinates of the second bounding box to global pixel coordinates in the remote sensing image;

[0012] Extract interpretable features of the second bounding box, including size features and color features;

[0013] Construct a visual table carrier, set a screening threshold based on the prior features of the search target, including a size threshold and a color threshold;

[0014] Compare the interpretable features with the corresponding screening threshold to screen out third bounding boxes;

[0015] Based on the global pixel coordinates, spatially align the third bounding box with the corresponding remote sensing image;

[0016] Extract the scene features of the corresponding area of the third bounding box in the remote sensing image;

[0017] Based on the scene features, verify the third bounding box for the search target, and mark the real search target.

[0018] Further, the preprocessing of the remote sensing image to obtain the cropped unit includes the following steps:

[0019] According to the maximum size of the search target, set the overlap rate between the cropped units;

[0020] According to the overlap rate, block and crop the remote sensing image to obtain a plurality of cropped units;

[0021] Zero-fill the cropped units whose pixel number after cropping does not reach the preset standard;

[0022] Synchronously record the coordinate position information of the cropped units in the remote sensing image.

[0023] Further, the size features include the width pixel value, height pixel value and area pixel value of the second bounding box, and the color features include the white pixel ratio in the Lab color space and the red pixel ratio in the HSV color space of the image area within the second bounding box.

[0024] Further, the white pixel ratio in the Lab color space is calculated according to the following steps:

[0025] Calculate the Euclidean distance between the total pixels in the second bounding box and the standard white pixels;

[0026] The number of pixels whose Euclidean distance is less than a preset white threshold is taken as the number of white pixels;

[0027] The proportion of white pixels is obtained by comparing the number of white pixels with the total number of pixels in the second detection frame.

[0028] Furthermore, the percentage of red pixels in the HSV color space is calculated using the following steps:

[0029] Pixels with hue H within the red hue range in the second detection box are extracted as red pixel intervals;

[0030] Pixels with a saturation S greater than a preset red saturation threshold are selected from the red pixels and designated as the red pixels.

[0031] The number of red pixels obtained after filtering is counted;

[0032] The proportion of red pixels is obtained by comparing the number of red pixels with the total number of pixels in the second detection frame.

[0033] Furthermore, the pre-trained detection model includes a pre-trained YOLOv5-small deep learning small object detection model.

[0034] The advantages and positive effects of this invention are as follows: By adopting the above-mentioned technical solution and integrating three screening dimensions to form a progressive logic, the false negative rate and false positive rate are reduced, ensuring the reliability of target verification results; the screening logic is based on physical features, is easy to adjust, and can be extended to other UAV small target search tasks; block pruning reduces the amount of single model input data, adapts to parallel computing, and improves detection processing speed; the overlap rate is set based on the maximum target size to avoid edge target defects and reduce the risk of false negatives; the combination of size and color features can eliminate obvious false positives in batches, reducing the number of targets to be manually verified; color difference is quantified by Euclidean distance to accurately identify white pixels and reduce color misjudgment; color is combined with saturation screening to eliminate non-red and low-saturation interference and extract red paint. Attached Figure Description

[0035] Figure 1 This is a flowchart illustrating a human-machine collaborative screening method for interpretable features according to an embodiment of the present invention.

[0036] Figure 2 This is a schematic diagram of a drone-captured search scene according to an embodiment of the present invention;

[0037] Figure 3 This is a schematic diagram of the width and height distribution of the target detection results according to an embodiment of the present invention;

[0038] Figure 4FIG. 1 is a schematic diagram of a red and white proportion of a search target detection result according to an embodiment of the present application. DETAILED DESCRIPTION

[0039] The present disclosure will be described more fully hereinafter with reference to the accompanying drawings, in which example embodiments of the present disclosure are shown. The following description of example embodiments is made with reference to the accompanying drawings. The following description of example embodiments is made with reference to the accompanying drawings. Obviously, the described embodiments are only a part of embodiments of the present disclosure, rather than all the embodiments. Based on the embodiments in the present disclosure, all other embodiments obtained by those of ordinary skill in the art without creative work fall within the scope of the present disclosure.

[0040] As shown in Figure 1 The present application provides a human-computer collaborative screening method for interpretable features, comprising the following steps:

[0041] S01, collecting remote sensing images of a search area;

[0042] Batch inputting high-resolution aerial images of the target search area collected by the unmanned aerial vehicle into the system, and the remote sensing images need to cover the entire target search area to ensure that the area where the black box may exist is included in the collection range.

[0043] S02, pre-processing the remote sensing images to obtain a cropped unit;

[0044] Adaptive overlapping cropping To solve the problem of computational burden of directly processing high-resolution large images and the problem of incomplete edge targets, an image cropping strategy with an overlap rate is adopted. According to the maximum size of the black box (such as 60 pixels), the overlap rate between the cropped blocks (i.e. the cropped unit) is set. The cropped blocks with insufficient size after cropping are zero-filled, and the position information of each block in the original image is recorded for subsequent result relocation.

[0045] S03, performing parallel detection on the cropped unit through the pre-trained detection model, and outputting the detection result;

[0046] The system calls the internally integrated pre-trained YOLOv5-small deep learning small target detection model (prior art) to perform parallel detection on all cropped units, and outputs the first detection frame containing the suspected target position coordinates, contour size and confidence. The position coordinates of the first detection frame are based on the local pixel coordinates of the corresponding cropped unit.

[0047] S04, sorting all first detection frames in the detection result in descending order of confidence;

[0048] S05, calculating the intersection over union of adjacent first detection frames after sorting, and removing the first detection frame with an intersection over union greater than a preset overlap threshold, and screening out a second detection frame;

[0049] The intersection over union (IoU) of the first detection box after sorting is calculated, low-confidence initial detection boxes with an intersection over union greater than a preset overlap threshold are removed, and high-confidence first detection boxes for the same suspected target are retained as second detection boxes.

[0050] When the image is overlapped and cropped, the adjacent cropped blocks may both contain the same target, causing the target to be detected multiple times, resulting in multiple overlapping detection boxes, which causes duplication of results and data redundancy. The NMS algorithm is applied for preliminary screening: the confidence of the detection box is sorted, the intersection over union of adjacent boxes is calculated, low-confidence boxes with high overlap with high-confidence boxes are removed, and the optimal detection box for the same target is retained to remove obvious redundant results.

[0051] S06, convert the local pixel coordinates of the second detection box to global pixel coordinates in the remote sensing image;

[0052] S07, extract the interpretable features of the second detection box, which include size features and color features;

[0053] The size features are the width pixel value, height pixel value, and area pixel value of the second detection box.

[0054] S08, construct a visual table carrier, set a screening threshold based on the prior features of the search target, which includes a size threshold and a color threshold;

[0055] Each row of the visual table carrier uniquely corresponds to a second detection box, and each column respectively carries the interpretable feature data and thumbnail information of the optimal detection box, and displays its thumbnail and extracted size, color, etc. properties, and the thumbnail information is the thumbnail of the image area corresponding to the second detection box. Set a screening threshold based on the prior features of the search target. The operator can sort any attribute column to quickly locate suspicious results, and perform batch confirmation or exclusion operations through multi-selection, full-selection, etc., to realize the rapid cycle screening of "setting threshold, filtering, sorting, batch operation".

[0056] S09, compare the interpretable features with the corresponding screening threshold to select third detection boxes;

[0057] According to the size threshold, the second detection boxes with width, height, or area exceeding the corresponding range in the visual table carrier are batch excluded.

[0058] S10, based on the global pixel coordinates, spatially align the third detection boxes with the corresponding remote sensing images, so that each third detection box accurately matches its corresponding actual shooting area in the remote sensing image;

[0059] S11, extract the scene features of the corresponding area of the third detection box in the remote sensing image;

[0060] The scene features include the relative position relationship between the search area and the surrounding ground objects, image texture features, and shadow shape features, the surrounding ground objects include grassland, snowfield, ruins, and road, the image texture features include whether there is a surface texture unique to the black box, and the shadow shape features match the physical shape of the black box.

[0061] S12, based on the scene features, verifying the third detection frame to search for the target, and marking the real search target.

[0062] The operator performs auditing in the original map interface, and through zooming in, zooming out, panning, and the like, can observe the context information of the target in the real scene, such as the relationship between the target and the surrounding ground objects, texture details, shadows, and the like. These information is the key to distinguish high-similarity false detections from real black boxes. A list of all results is displayed on one side of the interface, and the original map is displayed in the other side or middle area. By clicking any result in the list, the original map view will automatically focus and highlight the result area, facilitating the operator to make detailed comparison and final confirmation.

[0063] By using the above method, three screening dimensions are fused to form a progressive logic, the missing detection rate and the misjudgment rate are reduced, and the target verification result is ensured to be reliable; the screening logic is based on physical features, is easy to adjust, and can be popularized to other unmanned aerial vehicle micro target search tasks; the basic work is automatically processed, the human complex scene judgment is reserved, the advantages of man and machine are complementary, and the burden of the operator is reduced; the detection result scale is compressed through multi-stage automatic processing, the range of manual intervention is reduced, the labor intensity is reduced, and the low efficiency problem of pure manual verification is solved.

[0064] In order to solve the problems of heavy high-resolution image processing burden, edge target easy to be broken, detection frame difficult to return to the original map, and inconsistent model input size, an implementation manner is provided in the embodiment.

[0065] In an embodiment, the remote sensing image is preprocessed to obtain a clipping unit, including the following steps:

[0066] The overlap rate between the clipping units is set according to the maximum size of the search target;

[0067] Based on the physical properties of the search target, the parameters of the remote sensing image acquisition device, and the acquisition height, the maximum size of the target is calculated combined with the imaging principle, and the overlap rate between the clipping units is set according to the maximum size of the search target, for example, the search target is a black box, the maximum pixel size of the black box in the image is calculated combined with the imaging principle according to the physical size of the black box, the unmanned aerial vehicle shooting height, and the camera parameters, so as to set the overlap rate between the clipping units, and avoid the edge target from being broken due to clipping.

[0068] The remote sensing image is divided and clipped according to the overlap rate to obtain a plurality of clipping units;

[0069] The zero value filling is performed on the cropped unit whose pixel number after cropping does not reach the preset standard, so as to ensure that the pixel numbers of all the cropped units are consistent.

[0070] The coordinate position information of the cropped unit in the remote sensing image is recorded, and the coordinate position information includes the global pixel coordinates of the upper left corner of the cropped unit in the remote sensing image and the pixel size of the cropped unit.

[0071] By using the above method, the block cropping reduces the data amount of single model input, adapts to parallel computing, and improves the detection processing speed; the maximum target size is used to set the overlap rate, so as to avoid the edge target defect and reduce the risk of missed detection; the coordinate position information is recorded, so as to provide a basis for the detection frame coordinate conversion and the original picture back labeling, and ensure the continuity of the position information; the zero value filling makes the sizes of the cropped units consistent, so as to avoid the precision fluctuation caused by uneven input of the model.

[0072] In order to solve the problems that the features of the traditional model cannot be interpreted, the false detection cannot be excluded by the features, and the precision of single feature screening is insufficient, an implementation manner is provided in the embodiment.

[0073] In an embodiment, the size feature includes a width pixel value, a height pixel value and an area pixel value of the second detection frame, and the color feature includes a white pixel proportion in the Lab color space and a red pixel proportion in the HSV color space of an image region in the second detection frame.

[0074] By using the above method, the physical properties are used, the adjustment is easy to understand, no professional model knowledge is needed, and the use threshold is reduced; the size and color features are combined, the obvious false detection is excluded in batches, and the number of targets verified by manual verification is reduced; the multi-dimensional screening avoids the single feature from being disturbed by the environment, and the reliability of the screening result is improved.

[0075] In order to solve the problems that the white pixel is easily misjudged under complex illumination, and the precision of the traditional brightness threshold is insufficient, an implementation manner is provided in the embodiment.

[0076] In an embodiment, the white pixel proportion in the Lab color space is calculated according to the following steps:

[0077] The Euclidean distance between the total pixels in the second detection frame and the standard white pixel is calculated.

[0078] The number of pixels with a Euclidean distance less than a preset white threshold is counted as the number of white pixels.

[0079] The number of white pixels is compared with the total number of pixels in the second detection frame, and the proportion of white pixels is obtained.

[0080] By using the above method, the color difference is quantified by the Euclidean distance, the white pixels are accurately identified, and the color misjudgment is reduced; the Lab space isolates the influence of light, ensuring the stability of the white proportion calculation under complex light; and the accurate data provides a basis for threshold screening, excluding interference objects with insufficient white proportion.

[0081] To solve the problem of red color being easily interfered with in a complex background and the misjudgment of red color extraction in the traditional RGB space, an implementation manner is provided in the embodiment.

[0082] In an embodiment, the proportion of red color pixels in the HSV color space is calculated according to the following steps:

[0083] Extracting the pixels with hue H in the red hue range in the second detection frame as the red pixel interval;

[0084] Screening the pixels with saturation S greater than the preset red saturation threshold from the red pixels as the red pixels;

[0085] Counting the number of red pixels obtained after screening;

[0086] Calculating the proportion of red pixels by comparing the number of red pixels with the total number of pixels in the second detection frame.

[0087] By using the above method, the hue is combined with the saturation screening to exclude non-red and low-saturation interference and extract red paint; the HSV space is independent of the luminance channel to avoid misjudgment of red features caused by light and enhance the robustness; the low-saturation “pseudo-red” is removed to reduce the false detection caused by color similarity and reduce the verification pressure of the original image.

[0088] In an embodiment, the pre-trained detection model includes a pre-trained YOLOv5-small deep learning small target detection model.

[0089] To facilitate the human-computer collaborative screening method of the interpretable features provided by the present disclosure, the present disclosure further provides a human-computer collaborative screening system for interpretable features, which includes a data preprocessing module, a primary screening module and a secondary screening module:

[0090] The data preprocessing module is used to batch receive remote sensing images of a target search area, and pre-process the remote sensing images to obtain a clipping unit;

[0091] The primary screening module is used to call a pre-trained small target deep learning detection model to detect the clipping unit to output an initial detection frame, remove redundancy from the first detection frame to obtain a second detection frame, and convert the coordinates of the second detection frame;

[0092] The secondary screening module includes a table-based quick screening unit and an original image-based fine screening unit. The table-based quick screening unit is used to construct a visual table carrier, extract and fill features, and set thresholds to filter and obtain the third detection box in the table. The original image-based fine screening unit is used to achieve spatial alignment between the detection box and the remote sensing image, extract scene features, and complete target verification.

[0093] The following description, in conjunction with a preferred embodiment, illustrates the content involved in the above embodiments.

[0094] like Figures 2-4 As shown, a black box is deployed in a simulated accident site covering one square kilometer. A drone is conducting aerial photography from a height of 60 meters. Figure 2 As shown), where The altitude of the drone, The width of the aerial image (i.e., the remote sensing image). For the height of aerial images, To detect the width of the target image region (i.e., the detection box). To detect the height of the target image region, The unit is meters. , , and The unit is pixels, and a total of 1371 high-resolution images were acquired. First, the data was input and preprocessed, and the 1371 images were imported into the system of this invention in batches. The system sets the cropping block size to 1024x1024 pixels and the overlap rate to 60 pixels, cropping the original image to generate tens of thousands of smaller images, and recording the coordinate information of each smaller image.

[0095] Then, an initial screening is performed. The system calls its internally integrated YOLOv5-small object detection model to perform parallel detection on tens of thousands of small images, initially outputting several thousand suspected object boxes. Subsequently, the system automatically executes the NMS algorithm to merge duplicate boxes, obtaining approximately 1,500 preliminary screening results.

[0096] Finally, a second round of screening is conducted. In the table-based quick screening stage (the first step of the second screening), feature extraction is performed: the system automatically calculates the size (width, height, area) and color (red percentage, white percentage) features for these 1500 results. Figures 3-4As shown, the blue circle in the figure, i.e., 0, represents a false detection target, and the orange square, i.e., 1, represents a correct detection target. The "high" and "width" represent the height and width of the detection frame in pixels. The "red proportion" and "white proportion" represent the red proportion and white proportion in the detection frame. A result table is generated. Rule application: The operator sets the screening rules according to prior knowledge in the table interface: size rule: area range [400, 6400] pixels. Color rule: red proportion > 0.05, white proportion > 0.01. Batch exclusion: Click the "Apply Screening" button, and the table is refreshed instantly. About 1400 results that do not meet the rules are hidden. About 100 remaining results are displayed in the table. Quick review: The operator sorts the 100 results by "area" from large to small. By quickly browsing the thumbnails, it is found that most of them are stones or signs with bright colors but inconsistent shapes. Using the multi-selection function, the operator excludes about 90 obvious false detections in 1 minute. Original image fine screening stage (second screening second step): Original image alignment: There are only about 10 highly suspicious results left in the table. The operator switches to the "original image fine screening" interface. Context judgment: The interface loads the original large image containing the 10 results. The operator clicks the results in the list one by one, and the view automatically jumps and zooms in. For result A: In the original image, it is a red warning sign in the shape of a rectangle, which is inconsistent with the cylinder of the black box. Combined with the context that it is located on the roadside, it is determined to be a false detection. For result B: In the original image, it is an object partially covered by mud, but through magnified observation, the exposed part presents standard red and white stripes, and the surrounding environment is consistent with the scene of being thrown off. The operator finally determines it to be a real black box in combination with the texture and shadow in the image. Result confirmation: The operator marks the real black box result as "confirmed" and the others as "excluded".

[0097] Result and efficiency analysis: Time consumption statistics: The entire secondary screening process, including table quick screening and original image fine screening, takes the operator only about 5 minutes. Comparison: If a purely manual method is used, the operator needs to search for each of the 1371 images, which takes an average of about 33 hours. Conclusion: In this simulation task, using the system, the real black box is successfully locked from the massive detection results, and the total time consumption is about 0.67 hours (including detection and screening), which is about 50 times more efficient than the 33 hours of pure manual search, verifying the outstandingness and practicality of the method.

[0098] Based on the embodiments of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium and a computer program product.

[0099] An electronic device includes at least one processor; and a memory communicatively connected with the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method for human-computer collaborative screening of interpretable features provided by the present disclosure.

[0100] Electronic device is intended to refer to various forms of digital computers, such as laptops, desktops, tablets, personal digital assistants, servers, blade servers, mainframes, and other appropriate computers. Electronic device can also refer to various forms of mobile devices, such as personal digital assistants, cellular telephones, smartphones, wearable devices, and other similar computing devices. The components shown in the figures, their connections and relationships, and their functions, are meant to be examples only, and are not intended to limit the implementations of the present disclosure described and / or claimed in this document.

[0101] A non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to enable a computer to perform the method for human-computer collaborative screening of interpretable features provided by the present disclosure.

[0102] Various implementations in the present disclosure can be implemented in digital electronic circuitry, integrated circuitry, field programmable gate arrays (FPGA), application specific integrated circuits (ASIC), application specific standard products (ASSP), system on a chip systems (SOC), complex programmable logic devices (CPLD), computer hardware, firmware, software, and / or combinations thereof. These various implementations can include implementation in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which can be special or general purpose, coupled to receive data and instructions from, and to transmit data and instructions to, a storage system, at least one input device, and at least one output device.

[0103] A computer program product includes computer programs / instructions, and the computer programs / instructions perform the method for human-computer collaborative screening of interpretable features provided by the present disclosure when executed by a processor.

[0104] Program code for carrying out methods of the present disclosure can be written in any combination of one or more programming languages. The program code can be provided to a processor or controller of a general purpose computer, special purpose computer, or other programmable data processing apparatus to produce a machine, such that the program code, when executed by the processor or controller, produces the functions / operations specified in the flowcharts and / or the block diagrams. The program code can be executed entirely on a machine, partially on a machine, partially on a machine as a standalone software package, or entirely on a remote machine or server.

[0105] In the context of the present disclosure, a machine-readable medium can be a tangible medium that contains or stores a program for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of the machine-readable storage medium will include one or more lines of electrical connections, portable computer disks, hard disk drives, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or Flash memory), optical fibers, portable compact disc read-only memories (CD-ROMs), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0106] The above detailed description of the embodiments of the present application has been presented for the purposes of illustration and description. It is not intended to be exhaustive or to limit the application to the precise form disclosed. Many modifications and variations will be apparent to practitioners skilled in this art. Embodiments were chosen and described in order to best explain the principles of the application and its best mode, and to enable others skilled in the art to best utilize the application.

Claims

1. A human-machine collaborative screening method for interpretable features, characterized in that, Includes the following steps: Collect remote sensing images of the search area; The remote sensing image is preprocessed to obtain a cropping unit; The cropping unit is subjected to parallel detection using a pre-trained detection model, and the detection results are output. All first detection boxes in the detection results are sorted from high to low confidence level; Calculate the intersection-union ratio (CUIR) of adjacent first detection boxes after sorting, remove the first detection boxes whose CUIR is greater than a preset overlap threshold, and select the second detection boxes. The local pixel coordinates of the second detection box are converted into global pixel coordinates in the remote sensing image; Extract interpretable features from the second detection box, the interpretable features including size features and color features; A visual table is constructed, and a filtering threshold is set based on the prior features of the target being searched. The filtering threshold includes a size threshold and a color threshold. The interpretable features are compared with the corresponding filtering thresholds to filter out the third detection box; Based on the global pixel coordinates, the third detection box is spatially aligned with the corresponding remote sensing image; Extract scene features of the corresponding region of the third detection box in the remote sensing image; Based on the scene features, the third detection box is used to verify the search target and mark the real search target.

2. The human-machine collaborative screening method for interpretable features according to claim 1, characterized in that, Preprocessing the remote sensing image to obtain the cropping unit includes the following steps: The overlap rate between the cropping units is set according to the maximum size of the target being searched. The remote sensing image is divided into blocks and cropped according to the overlap rate to obtain a plurality of cropping units; Zero-value filling is performed on cropping units whose number of pixels after cropping does not meet the preset standard; The coordinate position information of the cropping unit in the remote sensing image is recorded synchronously.

3. The human-machine collaborative screening method for interpretable features according to claim 1, characterized in that: The size features include the width pixel value, height pixel value, and area pixel value of the second detection box, and the color features include the proportion of white pixels in the Lab color space and the proportion of red pixels in the HSV color space of the image area within the second detection box.

4. The human-machine collaborative screening method for interpretable features according to claim 3, characterized in that, The percentage of white pixels in the Lab color space is calculated using the following steps: Calculate the Euclidean distance between the total number of pixels within the second detection frame and the standard white pixel; The number of pixels whose Euclidean distance is less than a preset white threshold is taken as the number of white pixels; The proportion of white pixels is obtained by comparing the number of white pixels with the total number of pixels in the second detection frame.

5. The human-machine collaborative screening method for interpretable features according to claim 3, characterized in that, The percentage of red pixels in the HSV color space is calculated using the following steps: Pixels with hue H within the red hue range in the second detection box are extracted as red pixel intervals; Pixels with a saturation S greater than a preset red saturation threshold are selected from the red pixels and designated as the red pixels. The number of red pixels obtained after filtering is counted; The proportion of red pixels is obtained by comparing the number of red pixels with the total number of pixels in the second detection frame.

6. The human-machine collaborative screening method for interpretable features according to claim 1, characterized in that: The pre-trained detection model includes a pre-trained YOLOv5-small deep learning small object detection model.

Citation Information

Patent Citations

  • Target detection method based on remote sensing scene classification

    CN112800982A

  • Weak and small target identification and positioning method for airborne wide-view-field reconnaissance load

    CN118674912A