A method, apparatus, device, and medium for target detection in images

By cyclically optimizing the remote sensing image model and filtering the output results of the model with physical rules, the problem of low detection accuracy of small objects in remote sensing images is solved, and the detection accuracy and reliability are significantly improved.

CN119625291BActive Publication Date: 2025-05-27BEIJING HUITIAN ZHUOTE TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510169639.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-17
Publication Date
2025-05-27
Estimated Expiration
2045-02-17

AI Technical Summary

Technical Problem

The prior art has problems such as low detection accuracy, difficulty in model reliance on high-quality labeling data, and easy missed or missed in complex backgrounds and low-quality images in remote sensing images.

Method used

By performing model optimization operations in a loop, using the physical rules of the target object to standardize the model output results, gradually optimize the model and improve the detection effect. The specific steps include obtaining the sample data set, training the preparatory model, performing model optimization operations, determining whether the output results meet the standardization requirements, and using physical rules to filter the results, and finally combining the filtered results for model training.

Benefits of technology

The detection accuracy and reliability of the model are improved, and the detection effect of targets in the image is enhanced, especially in complex backgrounds and low-quality images.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119625291B_ABST
    Figure CN119625291B_ABST
Patent Text Reader

Abstract

The embodiments of this specification disclose a method for target detection in images, including: training a preliminary model using a sample data set; repeatedly performing model optimization operations; the first model optimization operation includes: inputting a sample image into the preliminary model, determining whether the output result meets the standardization requirements, screening the results that meet the standardization requirements using the physical rules matched by the target object, merging the screened results with the existing training data into new training data, and training a new model; starting from the second model optimization operation, each model optimization operation includes: inputting a sample image into the new model, determining whether the output result meets the standardization requirements, screening the results that meet the standardization requirements using the physical rules matched by the target object, merging the screened results with the existing training data into new training data, and training a new model; using the latest model to perform target object detection operations on the image to be detected to obtain target object detection results.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of image processing technology, and particularly to a method, device, equipment and medium for target detection in images. Background Art

[0002] With the rapid development of deep learning and remote sensing technology, remote sensing images are increasingly widely used in geographic information acquisition and target positioning. Optical remote sensing images can provide rich information for the recognition and positioning of ground object targets. However, traditional image processing methods have limitations in terms of automation and positioning accuracy.

[0003] Currently, the tasks of target detection and instance segmentation in remote sensing images mainly rely on deep learning models such as YOLO. These deep learning models can detect and segment targets in scenes. For example, to improve the positioning accuracy and processing efficiency of remote sensing images, the deep learning model YOLOv11 is combined with optical remote sensing images, and the efficient target detection and instance segmentation capabilities of YOLOv11 are utilized to achieve precise positioning of targets. This combination can not only automatically identify targets in images, but also significantly improve the automation and accuracy of positioning, meeting the requirements of practical applications.

[0004] However, in the face of small targets in remote sensing images (such as street lights in remote sensing images), the existing technologies have certain challenges: (1) The models rely on a large amount of high-quality labeled data, but the labeling cost of small targets in the remote sensing field is high and it is difficult to obtain data; (2) For small targets (such as street lights), the existing models are prone to missed detection or false detection in complex backgrounds or low-quality images; (3) The existing technologies usually only rely on pixel-level convolutional features and ignore the physical relationship between the target and the environment (such as the lamp post and the shadow); (4) The inference results of the models are usually directly used as outputs, lacking further screening and optimization based on physical laws or other prior knowledge.

[0005] The deficiencies of the existing technologies in small target detection and instance segmentation mainly focus on the weak ability to utilize the physical characteristics of targets and the lack of targeted optimization solutions suitable for remote sensing images. Especially in the street light detection task, the geometric physical laws such as the directionality and angle of the lamp post and its shadow are not fully utilized, resulting in low inference accuracy, and there is an urgent need for technical solutions to improve the detection effect. Summary of the Invention

[0006] The embodiments of this specification provide a method, device, equipment and medium for target detection in images, so as to solve the technical problem of how to improve the detection effect of targets in images.

[0007] To solve the above technical problem, the embodiments of this specification provide the following technical solutions:

[0008] An embodiment of this specification provides a method for target detection in images. The method includes:

[0009] Obtain a sample data set and train a preliminary model using the sample data set;

[0010] Loop through and perform model optimization operations;

[0011] Among them, the first model optimization operation includes: obtaining a sample image, inputting the sample image into the preliminary model, and obtaining the result output by the preliminary model; determining whether the result meets the standardization requirements, screening the results that meet the standardization requirements using the physical rules matched by the target object, merging the screened results with the existing training data into new training data, and training the preliminary model using the new training data to obtain an optimized new model;

[0012] Starting from the second model optimization operation, each model optimization operation includes: obtaining a sample image, inputting the sample image into the new model, and obtaining the result output by the new model; determining whether the result meets the standardization requirements, screening the results that meet the standardization requirements using the physical rules matched by the target object, merging the screened results with the existing training data into new training data, and training the new model using the new training data to obtain an optimized new model;

[0013] After reaching the end condition of the model optimization operation, use the latest model to perform target object detection operations on the image to be detected to obtain target object detection results.

[0014] Optionally, obtaining the sample data set includes:

[0015] Obtain the target area image, label the target object in the target area image to form shape feature data of the target object;

[0016] Generate small slices of the labeled data, adjust the fixed line width and generate a labeled file;

[0017] Convert the labeled file into a sample data set in a specific format.

[0018] Optionally, determining whether the result meets the standardization requirements includes:

[0019] Extract each polygon from the result and determine whether each polygon meets the standardization conditions.

[0020] Optionally, determining whether each polygon meets the standardization conditions includes:

[0021] For any polygon, perform a standardization judgment operation on this polygon in a loop;

[0022] Among them, the first standardized judgment operation includes:

[0023] Extracting the center line of the polygon according to the pixel coordinates of the polygon, and determining whether the center line meets a preset condition;

[0024] If yes, the center line is saved and the standardization judgment operation is no longer performed on the polygon;

[0025] If not, the next standardization judgment operation is performed;

[0026] The next standardization judgment operation includes:

[0027] Simplify the polygon, extract the center line according to the pixel coordinates of the simplified polygon, and determine whether the newly extracted center line meets the preset conditions;

[0028] If yes, the center line is saved and the standardization judgment operation is no longer performed;

[0029] If not, the next standardization judgment operation is performed;

[0030] If a center line that meets the preset conditions is obtained within a preset number of standardization judgment operations, the polygon meets the standardization conditions; if a center line that meets the preset conditions is not obtained within a preset number of standardization judgment operations, the polygon does not meet the standardization conditions.

[0031] Optionally, if the number of vertices of the center line is a first preset value, the center line satisfies the preset condition.

[0032] Optionally, using the physical rules matched by the target object to filter the results that meet the standardization requirements includes:

[0033] Determine the screening factors based on the physical rules matched by the target object;

[0034] The results that meet the standardization requirements are grouped according to the screening factors, and the groups whose number of results does not meet the preset number are eliminated.

[0035] Optionally, the screening factors include the size and / or direction of the angle between the target object and its own shadow.

[0036] The present invention provides a device for detecting an object in an image, the device comprising:

[0037] An initial training module, used to obtain a sample data set and train a preliminary model using the sample data set;

[0038] A loop optimization module, used to execute model optimization operations in a loop;

[0039] Among them, the first model optimization operation includes: obtaining a sample image, inputting the sample image into the preliminary model, and obtaining the result output by the preliminary model; determining whether the result meets the standardization requirements, screening the result that meets the standardization requirements by using the physical rules matched by the target object, combining the screened result with the existing training data into new training data, and training the preliminary model with the new training data to obtain an optimized new model;

[0040] Starting from the second model optimization operation, each model optimization operation includes: obtaining a sample image, inputting the sample image into the new model, and obtaining the result output by the new model; determining whether the result meets the standardization requirements, screening the result that meets the standardization requirements by using the physical rules matched by the target object, combining the screened result with the existing training data into new training data, and training the new model with the new training data to obtain an optimized new model;

[0041] The object detection module is used to perform an object detection operation on the image to be detected by using the latest model after reaching the end condition of the model optimization operation, and obtain an object detection result.

[0042] An embodiment of this specification provides an object detection device in an image, including:

[0043] At least one processor;

[0044] And,

[0045] A memory communicatively connected to the at least one processor;

[0046] Among them,

[0047] The memory stores instructions executable by the at least one processor, and when the instructions are executed by the at least one processor, the at least one processor can execute the above object detection method in an image.

[0048] An embodiment of this specification provides a computer-readable storage medium, and the computer-readable storage medium stores computer-executable instructions, and when the computer-executable instructions are executed by a processor, the above object detection method in an image is implemented.

[0049] The above at least one technical solution adopted by the embodiment of this specification can achieve the following beneficial effects:

[0050] By performing model optimization operations cyclically, the model can be gradually optimized, and the detection effect of the model can be improved.

[0051] In each model optimization operation, the output results of the model are subject to standardized judgment, and the physical rules matched by the target object are used to screen the results that meet the standardized requirements, accurately identifying and eliminating the possible error results or non-compliant results in the model optimization process, effectively improving the model optimization efficiency, significantly enhancing the detection accuracy and reliability of the optimized model, and thus improving the detection effect of the target in the image. Description of the Drawings

[0052] In order to more clearly illustrate the technical solutions in the embodiments of this specification or the prior art, the following will briefly describe the drawings required for use in the description of the embodiments of this specification or the prior art. Obviously, only the drawings required for use in some embodiments of this application are described below. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0053] Figure 1 It is a schematic flowchart of the method for detecting a target in an image provided by the first embodiment of this specification;

[0054] Figure 2 It is a schematic technical route diagram of the model optimization example in the first embodiment of this specification;

[0055] Figure 3 It is a schematic diagram of the three-point broken line annotation mode in the first embodiment of this specification;

[0056] Figure 4 It is a schematic flowchart of the center line extraction process in the first embodiment of this specification;

[0057] Figure 5 It is a schematic diagram of the polygon pixel coordinates in the first embodiment of this specification;

[0058] Figure 6 It is a schematic diagram of the center line in the first embodiment of this specification;

[0059] Figure 7 It is a schematic diagram of the reconstructed polygon in the first embodiment of this specification;

[0060] Figure 8 It is a schematic structural diagram of the device for detecting a target in an image provided by the second embodiment of this specification. Detailed Implementation Modes

[0061] To enable those skilled in the art to better understand the technical solutions in this specification, the technical solutions in the embodiments of this specification will be clearly and completely described below in conjunction with the accompanying drawings. Obviously, the embodiments involved in the specific implementation manners are only a part of the embodiments of this application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the specific implementation manners without creative efforts shall fall within the protection scope of this application.

[0062] The first embodiment of this specification (hereinafter referred to as "Embodiment 1") provides a method for target detection in an image. The execution subject of Embodiment 1 includes but is not limited to a terminal, a server, an operating system, or an application program. That is, the execution subject can be various and can be set, used, or transformed according to needs.

[0063] As Figure 1 shown, the method for target detection in an image provided by Embodiment 1 includes:

[0064] S101: Obtain a sample data set and train a preliminary model using the sample data set;

[0065] In Embodiment 1, a sample data set can be obtained and a preliminary model can be trained using the sample data set.

[0066] Embodiment 1 does not limit the specific method for obtaining the sample data set. For example, obtaining the sample data set may include:

[0067] Obtain an image of the target area, label the target objects in the image of the target area to form shape feature data of the target objects;

[0068] Generate small slices of the labeled data, adjust the fixed line width and generate a labeled file;

[0069] Convert the labeled file into a sample data set in a specific format.

[0070] The following gives a specific example of obtaining the sample data set (Embodiment 1 is not limited to the following example):

[0071] 1. Obtain the required image of the target area. For example, use the Bigemap platform and the 1:5000 standard map sheet to download the high-resolution remote sensing image with a resolution of Level_19 within the airport area (that is, the target area can be the airport area), and the image projection is unified to the WGS84 / EPSG: 4326 coordinate system.

[0072] 2. Label the target object in the image. For example, if the target object is a street lamp, you can use the LabelMe tool to label the target in the image. The labeled target is the airport street lamp and its shadow. The labeling mode uses the three-point polyline mode, that is, labeling the top of the lamp pole, the root of the lamp pole and the end point of the shadow to form accurate shape feature data, such as Figure 3 shown.

[0073] 3. After the annotation is completed, the image can be converted into a sample data set in a specific format (for example, YOLO format). Specifically, the annotation data (for example, the target object is a street lamp, the annotation data is a broken line formed by the three top lines of the street lamp top, the street lamp base and the top of the shadow) can be sliced, the fixed line width can be adjusted (for example, the target object is a street lamp, after determining the fixed line width, a polygonal pixel coordinate of the street lamp and its shadow can be obtained) and the corresponding JSON file can be generated as the annotation file, and then the annotation file can be converted into a sample data set in YOLO format. The sample data set includes information such as bounding boxes and category labels.

[0074] In the first embodiment, after obtaining the sample data set, the sample data set can be used to train a preliminary model. The following takes the YOLOV11-seg (You Only Look Once version 11-Segmentation) model (this model is an extension of the YOLO series model and can be used for instance segmentation tasks) as an example to illustrate how to train the preliminary model.

[0075] (1) Environment construction: Complete the construction of the training framework in a GPU-supported environment, including upgrading the GPU driver, installing and configuring CUDA and cuDNN; using Conda to create a virtual environment and install PyTorch; and deploying the YOLO training framework.

[0076] (2) Dataset division: Divide the sample data set into training set, validation set, and test set in proportion to ensure reasonable data distribution. The division ratio of the training set, validation set, and test set can be set as needed, such as 8:1:1 or 7:2:1. In particular, the division ratio can be determined according to the size of the sample data set to ensure reasonable data distribution and improve the model training effect.

[0077] (3) Model training: Use the YOLOV11-seg model for training, load the formatted training set for segmentation task training, use the validation set for verification, and use the test set for testing. After reaching the predetermined training goal, save the training weight file best.pt to obtain the trained model, that is, the preliminary model, for use in the subsequent model inference stage.

[0078] S103: Repeatedly perform model optimization operations;

[0079] In the first embodiment, the model optimization operations can be performed repeatedly. Among them, the first model optimization operation may include: obtaining a sample image, inputting the sample image into the preliminary model, and obtaining the result output by the preliminary model; determining whether each result (i.e., the model output result) meets the standardization requirements, screening the results that meet the standardization requirements using the physical rules matched by the target object, merging the screened results with the existing training data into new training data, and training the preliminary model with the new training data to obtain an optimized new model;

[0080] Starting from the second model optimization operation, each model optimization operation may include: obtaining a sample image, inputting the sample image into the new model, and obtaining the result output by the new model; determining whether each result (i.e., the model output result) meets the standardization requirements, screening the results that meet the standardization requirements using the physical rules matched by the target object, merging the screened results with the existing training data into new training data, and training the new model with the new training data to obtain an optimized new model.

[0081] In each model optimization operation of the first embodiment, determining whether the model output result meets the standardization requirements may include: extracting each polygon from each model output result and determining whether each polygon meets the standardization conditions.

[0082] Among them, determining whether each polygon meets the standardization conditions may include:

[0083] S1031: For any polygon, repeatedly perform standardization judgment operations on the polygon;

[0084] Among them, the first standardization judgment operation may include:

[0085] Extracting the center line of the polygon according to the pixel coordinates of the polygon and determining whether the center line meets the preset conditions;

[0086] If so, save the center line and no longer perform the standardization judgment operation on the polygon;

[0087] If not, perform the next standardization judgment operation;

[0088] The next standardization judgment operation may include:

[0089] Simplifying the polygon, extracting the center line according to the pixel coordinates of the simplified polygon, and determining whether the newly extracted center line meets the preset conditions;

[0090] If so, save the center line and no longer perform the standardization judgment operation;

[0091] If not, perform the next standardization judgment operation.

[0092] S1033: If a center line that meets the preset conditions is obtained within the preset number of standardization judgment operations, the polygon meets the standardization conditions; if a center line that meets the preset conditions is not obtained within the preset number of standardization judgment operations, the polygon does not meet the standardization conditions.

[0093] The above preset conditions can be set as needed. For example, if the number of vertices of the center line is a preset value, it is determined that the center line meets the standardization conditions.

[0094] For any model output result, if the polygon extracted from the model output result meets the standardization conditions, the model output result meets the standardization requirements.

[0095] In each model optimization operation of the first embodiment, screening the results that meet the standardization requirements by using the physical rules matched by the target object may include: determining screening factors according to the physical rules matched by the target object; grouping the results that meet the standardization requirements according to the screening factors, and removing the groups whose included result quantity does not meet the preset quantity.

[0096] The specific content of the above screening factors is not limited. For example, the screening factors include the size and / or direction of the angle between the target object and its own shadow.

[0097] The following combines Figure 2 , and gives specific examples of each model optimization operation (the first embodiment is not limited to the following examples):

[0098] The first model optimization operation

[0099] In this example, the first model optimization operation can be carried out according to the following content:

[0100] 1. Automatically download sample images. For example, generate the border longitude and latitude coordinates of the airport range using the airport geographical longitude and latitude coordinate data, use a script to automatically download Google images with a resolution of Level_19, and convert the image projection to the WGS84 / EPSG: 4326 coordinate system, and use the obtained image as the sample image.

[0101] 2. After obtaining the sample image, preprocessing operations can be performed on the sample image. For example, for ultra-large area images, the large image is cropped into small tiles suitable for YOLO input, or the Slicing Aided Hyper Inference (SAHI) tool is combined to achieve efficient inference.

[0102] 3. Input the sample image (or the preprocessed image) into the preliminary model, and let the model perform inference and output the corresponding results. For example, input the image into the aforementioned best.pt model for inference, that is, perform instance segmentation on the image and output the corresponding results. In this example, if the role of the model is to detect the target object, the output results may include a txt file of the detection label (the content of the txt file may include the label type, the normalized polygon coordinates, etc.) and the corresponding png file.

[0103] 4. Determine whether the results output by the preliminary model meet the standardization requirements, including extracting polygons separately from each of the output results (specifically, extracting the normalized polygon coordinates from the txt file), and determining whether each of the extracted polygons meets the standardization conditions.

[0104] Reference Figure 4 , for any polygon, determining whether the polygon meets the standardization conditions may include:

[0105] Perform the standardization judgment operation on this polygon in a loop.

[0106] The first standardization judgment operation is as follows:

[0107] (1) Coordinate conversion: For any polygon, convert its normalized coordinates to pixel coordinates, as shown in, for example Figure 5 .

[0108] (2) Center line coordinate extraction: Extract the center line according to the polygon pixel coordinates of this polygon, as shown in, for example Figure 6 .

[0109] (3) Determine whether the center line meets the preset conditions, specifically, determine whether the number of vertices of the center line is the first preset value. Depending on the different target objects, the preset conditions or the first preset value may be different. For example, if the target object is a street lamp and the three-point broken line mode is used for the aforementioned annotation of the street lamp, the preset condition can be that the number of vertices is 3, and the first preset value is 3. That is, if the number of vertices of the center line is 3, the center line meets the preset conditions; if the number of vertices of the center line is not 3, the center line does not meet the preset conditions. When the target object changes, the annotation method or the number of annotated vertices may be different, and the preset conditions or the first preset value may be different. Generally, the preset conditions or the first preset value correspond to the annotation method or the number of annotated vertices used for the target object.

[0110] (4) If the center line meets the preset conditions, save the center line and no longer perform the standardization judgment operation on this polygon.

[0111] If the center line does not meet the preset conditions, perform the next standardization judgment operation.

[0112] The next standardization judgment operation is as follows:

[0113] (1)Simplify the polygon (for example, use the RDP algorithm or other appropriate methods for simplification). Specifically, simplify the pixel coordinates of the polygon, then extract the center line according to the pixel coordinates of the simplified polygon, and determine whether the newly extracted center line meets the above preset conditions.

[0114] (2)If the newly extracted center line meets the preset conditions, save the center line and no longer perform the standardization judgment operation on this polygon.

[0115] If the newly extracted center line does not meet the preset conditions, continue to perform the next standardization judgment operation, including continuing to simplify the polygon, extracting the center line, and determining whether the newly extracted center line meets the preset conditions.

[0116] Of course, the standardization judgment operation is not carried out infinitely. It can be set that the standardization judgment operation is carried out at most n times. If a center line that meets the preset conditions is obtained in a certain standardization judgment operation within n times of standardization judgment operations, it can be determined that this polygon meets the standardization conditions. If no center line that meets the preset conditions is obtained after n times of standardization judgment operations, it is determined that this polygon does not meet the preset conditions, and this polygon can be discarded.

[0117] For each polygon extracted from the output results of each model, it can be judged whether it meets the standardization conditions as above. Since each polygon is the result of the preliminary model detecting the target object, the polygon that meets the standardization conditions is the polygon that matches the shape special features of the target object. In the above example, the target object is a street lamp. If the center line of the polygon has three vertices, it means that the shape features of the polygon match the annotation pattern and shape features of the street lamp. Therefore, the polygon meets the shape feature standard that the target object should have and meets the standardization conditions. The target object may be included in the polygon (whether the target object is really included in the polygon is determined by screening through the following physical rules).

[0118] For any result output by the preliminary model, if the polygon extracted from this result meets the standardization conditions, it can be considered that this result meets the standardization requirements.

[0119] 5. For the model output results that meet the standardization requirements, the physical rules matched by the target object can be used to screen the model output results that meet the standardization requirements. Among them, the physical rules can be selected as needed, including various geometric characteristic rules. For example, when the imaging time is unified, in the same image, the size and direction of the included angle formed by all target objects and their shadows are the same. Then the physical rule matched by the target object can be the included angle consistency rule, and the included angle consistency rule includes that the size and direction of the included angle are the same.

[0120] Taking the angle consistency rule as an example, the following illustrates how to use the physical rules matched by the target object to screen the results that meet the standardization requirements.

[0121] According to the physical rules matched by the target object, the screening factors are determined. If the physical rule is the angle consistency rule, the screening factors can include the magnitude and / or direction of the angle between the target object and its own shadow. If the physical rule is other rules, the screening factors can be determined accordingly.

[0122] Group the model output results that meet the standardization requirements according to the screening factors, and eliminate the groups whose included result quantities do not meet or are less than the preset quantity. The specific process can include: through the aforementioned standardization judgment process, the polygon centerlines of each model output result that meets the standardization requirements have been obtained. Then, the "polygon centerlines of each model output result that meets the standardization requirements" (hereinafter referred to as the centerlines to be grouped) obtained can be screened according to the angle and direction respectively.

[0123] Angle screening

[0124] The angle between the two broken lines formed by the three vertices of each centerline to be grouped can be used as the angle of the centerline to be grouped, so that the angle of each centerline to be grouped can be determined.

[0125] The non-overlapping first groups can be determined, and each first group covers a certain angle range. For example, calculate the average value A of the angles of the above-mentioned polygon centerlines, and determine the span value a (variable), the minimum value A - pa (the minimum value is generally not greater than the minimum value of the angles of the centerlines to be grouped), and the maximum value A + qa (the maximum value is generally not greater than the maximum value of the directions of the centerlines to be grouped). Then, the first groups are respectively [A - pa, A - (p - 2)a), [(p - 2)a, (p - 4)a), ……, [A - a, A + a), [A + a, A + 2a), ……, [A + (q - 4)a, A + (q - 2)a), [A + (q - 2)a, A + qa). According to the angle of each centerline to be grouped, the centerlines to be grouped are assigned to the corresponding first group, so as to group the centerlines to be grouped according to the angle.

[0126] Determine the number of centerlines actually included in each first group, and then eliminate the first groups whose included centerline quantities are less than the first preset quantity, and retain the other first groups, so as to achieve angle screening. Take the centerlines retained after angle screening as the centerlines to be grouped and perform the following direction screening.

[0127] Direction screening

[0128] The direction of the angle bisector between two broken lines formed by three vertices of each center line to be grouped can be used as the direction of the center line to be grouped. A unified coordinate system is established for each center line to be grouped, and the direction of each center line to be grouped corresponds to a vector. Values can be assigned to the directions of the determined center lines to be grouped. For example, the slope of the direction of the center line to be grouped can be used as the value of the direction of the center line to be grouped, or the angle value of the direction of the center line to be grouped in the above coordinate system can be used as the value of the direction of the center line to be grouped.

[0129] Each non-overlapping second group can be determined, and each second group covers a certain value range. For example, calculate the average value B of the directions of the center lines of the polygons to be grouped, and determine the span value b (which can vary), the minimum value B - xb (the minimum value is generally not greater than the minimum value of the directions of the center lines to be grouped), and the maximum value B + yb (the maximum value is generally not greater than the maximum value of the directions of the center lines to be grouped). Then the second groups are respectively [B - xb, B - (x - 2)b), [(x - 2)b, (x - 4)b), ……, [B - b, B + b), [B + b, B + 2b), ……, [B + (y - 4)b, B + (y - 2)b), [B + (y - 2)b, B + yb). According to the values of the directions of each center line to be grouped, the center lines to be grouped are assigned to the corresponding second groups, thereby grouping the center lines to be grouped according to their directions.

[0130] Determine the number of center lines actually included in each second group, and then eliminate the second groups whose included number of center lines is less than the second preset number, and retain the other second groups, thereby achieving direction screening.

[0131] The above angle screening and direction screening both belong to physical rule screening. The first preset number and the second preset number both belong to preset numbers, and the first preset number and the second preset number are variable. In actual situations, there is no absolute order between angle screening and direction screening. If direction screening is performed first, the center lines retained after direction screening are used as the center lines to be grouped for angle screening.

[0132] After physical rule screening, the model output results to which the finally retained center lines belong or correspond are the retained model output results. Correspondingly, the model output results to which the center lines eliminated during the physical rule screening belong or correspond are the eliminated model output results.

[0133] After physical rule screening, for any retained model output result, it can be considered that the target object exists in the model output result, and the target object is included in the aforementioned polygon extracted from the model output result.

[0134] It can be seen that through the above-mentioned standardized judgment and physical rule screening, it is possible to confirm whether the model output result contains the target object and also confirm the position of the target object (the position of the center line of the polygon of the model output result is the position of the target object).

[0135] 6. The model output results retained after screening can be merged with the existing training data to form new training data, which can be carried out in the following specific ways:

[0136] (1) Perform small graph generation and polygon regularization operations on the retained model output results, including: for any retained model output result, generate a graph block of a specific size centered on a specific position of the target object in the model output result, and reconstruct the polygon according to the center line of the target object. Taking the target object as a street lamp as an example, the above-mentioned specific position can be the root of the street lamp, that is, the second vertex of the center line of the street lamp, and the specific size is, for example, 320×320. The reconstructed polygon is, for example Figure 7 as shown

[0137] (2) Perform geographic coordinate conversion operations, including: using the coordinate relationship between the original large image (i.e., the sample image) and the inference small image (i.e., the above-mentioned graph block of a specific size), convert the detected target position into the actual geographic coordinates. The target position here can be determined as needed. Taking the target object as a street lamp as an example, the target position can be the pixel coordinates of the root of the street lamp, so as to convert the pixel coordinates of the root of the street lamp into the longitude and latitude coordinates of the root of the street lamp.

[0138] (3) Generate a JSON file, including: create a corresponding standard JSON file, and write information such as the label category obtained above, the normalized polygon coordinates in the retained model output result, and the target geographic longitude and latitude (i.e., the above-mentioned actual geographic coordinates) into the JSON file.

[0139] (4) Merge the newly generated JSON file with the existing training data to form new training data, that is, a new training set.

[0140] Continue to perform iterative training on the preliminary model with the new training data and optimize the preliminary model to obtain an optimized new model.

[0141] The next model optimization operation

[0142] Starting from the second model optimization operation, each model optimization operation can be carried out according to the following content:

[0143] 1. Obtain the sample image, which can refer to the method of obtaining the sample image in the first model optimization operation, but the sample images obtained each time are not necessarily the same.

[0144] 2. When needed, the sample image preprocessing can be referred to the first model optimization operation.

[0145] 3. Input the sample image (or the preprocessed image) into the new model obtained from the previous model optimization operation, and obtain the result output by the new model.

[0146] 4. Determine whether the result output by the new model meets the standardization requirements. The content for determining whether the result output by the model meets the standardization requirements in the first model optimization operation can be referred to for execution.

[0147] 5. For the results that meet the standardization requirements, use the physical rules matched by the target object to screen the results that meet the standardization requirements. The content for physical rule screening in the first model optimization operation can be referred to for execution.

[0148] 6. Combine the retained model output results after screening with the existing training data (the existing training data is the training data existing before the current model optimization operation, including the training data in S101 and the retained model output results of each previous model optimization operation before the current time) to form new training data. The content for combining to obtain new training data in the first model optimization operation can be referred to for execution.

[0149] Iteratively train the new model obtained from the previous model optimization operation with the new training data to continuously obtain an optimized new model for the next model optimization operation.

[0150] S105: After reaching the end condition of the model optimization operation, use the latest model to perform a target object detection operation on the image to be detected, and obtain the target object detection result.

[0151] In the first embodiment, the model optimization operation is not carried out without limit. The end condition of the model optimization operation can be preset. When the end condition of the model optimization operation is reached, the model optimization operation will no longer be carried out. The end condition of the model optimization operation can be set as needed. For example, the number of times of the model optimization operation reaching a preset number can be set as the end condition, or the target object detection effect (including but not limited to the correct rate) of the latest obtained model reaching a preset effect can be set as the end condition.

[0152] After reaching the end condition of the model optimization operation, the latest model (i.e., the optimized new model obtained from the last model optimization operation) can be used to perform a target object detection operation on the image to be detected, and obtain the target object detection result.

[0153] The content of the image to be detected and the acquisition method are not limited in the first embodiment.

[0154] The first embodiment can achieve the following beneficial effects:

[0155] By performing model optimization operations through cycles, the model can be gradually optimized, achieving self-improvement of the model and enhancing the detection effect of the model on the target object.

[0156] In each model optimization operation, a standardized judgment is made on the output result of the model, and the physical rules matched by the target object are used to screen the results that meet the standardized requirements, accurately identifying and eliminating possible error results or non-compliant results during the model optimization process, effectively improving the model optimization efficiency, significantly enhancing the detection accuracy and reliability of the optimized model for the target object, and thus improving the detection effect and efficiency of using the optimized model to detect the image to be detected.

[0157] Specifically, the physical laws such as the size and direction of the angle formed by the target object and its shadow in the remote sensing image can be utilized, and a screening method for the model output result based on the principle of geometric angle consistency can be adopted. By calculating the angle size and direction between the target object and its shadow, possible error results or non-compliant results during the model inference process can be accurately identified and eliminated, significantly improving the screening accuracy and reliability of the model output result, and further effectively improving the model optimization effect.

[0158] In each model optimization operation, the retained model output result after screening is merged with the existing training data to form new training data for model training. Through this data merging and cyclic training optimization strategy, the scale of the training data can be gradually increased, and the detection ability of the model can be optimized.

[0159] During the model optimization operation process, preprocessing including multiple aspects is performed on the sample image, making the preprocessed image more adaptable to the model optimization operation, realizing the combination of model optimization and image processing, and thus improving the model optimization effect and efficiency.

[0160] During the model optimization operation process, based on the obtained model output result, combined with geometric processing techniques such as polygon coordinate extraction, centerline simplification, and regularized polygon construction, the centerline can be extracted from the instance segmentation result, simplified and regularized, and finally a polygon area that meets the actual application requirements can be reconstructed for subsequent data merging. These geometric processing techniques can accurately describe the physical characteristics of the target, providing high-quality data support for the standardized judgment operation, physical rule screening, and data merging, thereby improving the model optimization effect and efficiency. Moreover, these geometric processing techniques can specifically reconstruct a polygon area that meets the actual application requirements for small-area or small-volume target objects, ensuring that the optimized model can specifically detect small-area or small-volume target objects, improving the detection effect and detection efficiency for small-area or small-volume target objects.

[0161] In the first embodiment, in each model optimization operation, after obtaining the model output result, a series of post-processing techniques are adopted to further process the model output result. These post-processing techniques include: (1) screening the model output result through physical rules; (2) extracting the polygons and centerlines of each region (i.e., each model output result) after instance segmentation and converting them into standardized data for output; (3) eliminating the error detection results that do not conform to the actual physical laws or the results that do not meet the requirements to improve the reliability of the inference result; (4) integrating and converting the retained model output results after screening into a standard format suitable for model training, and merging them with the existing training data to achieve cyclic training. Through these post-processing techniques, the model optimization effect and efficiency can be effectively improved, and further, the effect and efficiency of using the optimized model to detect the image to be detected can be improved.

[0162] As Figure 8 shown, the second embodiment of this specification provides an object detection device in an image corresponding to the method described in the first embodiment. The device includes:

[0163] An initial training module 202, configured to obtain a sample data set and train a preliminary model using the sample data set;

[0164] A cyclic optimization module 204, configured to perform model optimization operations cyclically;

[0165] Among them, the first model optimization operation includes: obtaining a sample image, inputting the sample image into the preliminary model, and obtaining the result output by the preliminary model; determining whether the result meets the standardization requirements, screening the result that meets the standardization requirements using the physical rules matched by the target object, merging the screened result with the existing training data into new training data, and training the preliminary model using the new training data to obtain an optimized new model;

[0166] Starting from the second model optimization operation, each model optimization operation includes: obtaining a sample image, inputting the sample image into the new model, and obtaining the result output by the new model; determining whether the result meets the standardization requirements, screening the result that meets the standardization requirements using the physical rules matched by the target object, merging the screened result with the existing training data into new training data, and training the new model using the new training data to obtain an optimized new model;

[0167] An object detection module 206, configured to perform an object detection operation on the image to be detected using the latest model after reaching the end condition of the model optimization operation, and obtain an object detection result.

[0168] Optionally, obtaining the sample data set includes:

[0169] Obtain the image of the target area, label the target object in the image of the target area to form the shape feature data of the target object;

[0170] Generate small slices of the labeled data, adjust the fixed line width and generate a labeled file;

[0171] Convert the labeled file into a sample data set in a specific format.

[0172] Optionally, determining whether the result meets the standardization requirements includes:

[0173] Extract each polygon from the result and determine whether each polygon meets the standardization conditions.

[0174] Optionally, determining whether each polygon meets the standardization conditions includes:

[0175] For any polygon, perform a standardization judgment operation on this polygon in a loop;

[0176] Among them, the first standardization judgment operation includes:

[0177] Extract the center line of the polygon according to the pixel coordinates of the polygon, and determine whether the center line meets the preset conditions;

[0178] If so, save the center line and no longer perform the standardization judgment operation on this polygon;

[0179] If not, perform the next standardization judgment operation;

[0180] The next standardization judgment operation includes:

[0181] Simplify the polygon, extract the center line according to the pixel coordinates of the simplified polygon, and determine whether the newly extracted center line meets the preset conditions;

[0182] If so, save the center line and no longer perform the standardization judgment operation;

[0183] If not, perform the next standardization judgment operation;

[0184] If a center line that meets the preset conditions is obtained within the preset number of standardization judgment operations, then this polygon meets the standardization conditions; if a center line that meets the preset conditions is not obtained within the preset number of standardization judgment operations, then this polygon does not meet the standardization conditions.

[0185] Optionally, if the number of vertices of the center line is the first preset value, then the center line meets the preset conditions.

[0186] Optionally, using the physical rules matched by the target object to screen the results that meet the standardization requirements includes:

[0187] Determine the screening factors according to the physical rules matched by the target object;

[0188] Group the results that meet the standardization requirements according to the screening factors, and eliminate the groups whose number of included results does not meet the preset number.

[0189] Optionally, the screening factors include the magnitude and / or direction of the angle between the target object and its own shadow.

[0190] The third embodiment of this specification provides an object detection device in an image, including:

[0191] At least one processor;

[0192] And,

[0193] A memory communicatively connected to the at least one processor;

[0194] Wherein,

[0195] The memory stores instructions executable by the at least one processor, and when the instructions are executed by the at least one processor, the at least one processor can execute the object detection method in the image described in the first embodiment.

[0196] The fourth embodiment of this specification provides a computer-readable storage medium, which stores computer-executable instructions, and when the computer-executable instructions are executed by a processor, the object detection method in the image described in the first embodiment is implemented.

[0197] The above embodiments can achieve the same technical effects, and the embodiments can be used in combination.

[0198] The above is only for the embodiments of this specification and is not used to limit this application. For those skilled in the art, this application can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of this application shall be included within the scope of the claims of this application.

Claims

1. A method for detecting an object in an image, characterized in that: The method comprises: Obtaining a sample data set, and using the sample data set to train a preliminary model; Execute model optimization operations in a loop; The first model optimization operation includes: obtaining a sample image, inputting the sample image into the preparatory model, and obtaining a result output by the preparatory model; judging whether the result meets the standardization requirements, screening the results that meet the standardization requirements using the physical rules matched by the target object, merging the screened results with the existing training data into new training data, and training the preparatory model using the new training data to obtain an optimized new model; Starting from the second model optimization operation, each model optimization operation includes: obtaining a sample image, inputting the sample image into a new model, and obtaining a result output by the new model; judging whether the result meets the standardization requirements, screening the results that meet the standardization requirements using the physical rules matched by the target object, merging the screened results with the existing training data into new training data, and training the new model using the new training data to obtain an optimized new model; After the end condition of the model optimization operation is reached, the target object detection operation is performed on the image to be detected using the latest model to obtain the target object detection result; Determining whether the result meets the standardization requirements includes: Extracting each polygon from the result, and determining whether each polygon satisfies a standardization condition; Judging whether each polygon meets the standardized conditions includes: For any polygon, perform the standardization judgment operation on the polygon in a loop; Among them, the first standardized judgment operation includes: Extracting the center line of the polygon according to the pixel coordinates of the polygon, and determining whether the center line meets a preset condition; If yes, the center line is saved and the standardization judgment operation is no longer performed on the polygon; If not, the next standardization judgment operation is performed; The next standardization judgment operation includes: Simplify the polygon, extract the center line according to the pixel coordinates of the simplified polygon, and determine whether the newly extracted center line meets the preset conditions; If yes, the center line is saved and the standardization judgment operation is no longer performed; If not, the next standardization judgment operation is performed; If a center line that meets the preset conditions is obtained within a preset number of standardization judgment operations, the polygon meets the standardization conditions; if a center line that meets the preset conditions is not obtained within a preset number of standardization judgment operations, the polygon does not meet the standardization conditions.

2. The method according to claim 1, characterized in that Get the sample dataset including: Acquire a target area image, and mark the target object in the target area image to form shape feature data of the target object; Generate small slices of the annotation data, adjust the fixed line width and generate the annotation file; The annotation file is converted into a sample data set in a specific format.

3. The method according to claim 1, characterized in that If the number of vertices of the center line is a first preset value, the center line satisfies the preset condition.

4. The method according to any one of claims 1 to 3, characterized in that Using the physical rules matched by the target object to screen the results that meet the standardization requirements includes: Determine the screening factors based on the physical rules matched by the target object; The results that meet the standardization requirements are grouped according to the screening factors, and the groups whose number of results does not meet the preset number are eliminated.

5. The method according to claim 4, characterized in that The screening factors include the size and / or direction of the angle between the target object and its own shadow.

6. A device for detecting an object in an image, characterized in that: The device comprises: An initial training module, used to obtain a sample data set and train a preliminary model using the sample data set; A loop optimization module, used to execute model optimization operations in a loop; The first model optimization operation includes: obtaining a sample image, inputting the sample image into the preparatory model, and obtaining a result output by the preparatory model; judging whether the result meets the standardization requirements, screening the results that meet the standardization requirements using the physical rules matched by the target object, merging the screened results with the existing training data into new training data, and training the preparatory model using the new training data to obtain an optimized new model; Starting from the second model optimization operation, each model optimization operation includes: obtaining a sample image, inputting the sample image into a new model, and obtaining a result output by the new model; judging whether the result meets the standardization requirements, screening the results that meet the standardization requirements using the physical rules matched by the target object, merging the screened results with the existing training data into new training data, and training the new model using the new training data to obtain an optimized new model; An object detection module is used to perform a target object detection operation on the image to be detected using the latest model after the end condition of the model optimization operation is met to obtain a target object detection result; Wherein, judging whether the result meets the standardization requirements includes: Extracting each polygon from the result, and determining whether each polygon satisfies a standardization condition; Judging whether each polygon meets the standardized conditions includes: For any polygon, perform the standardization judgment operation on the polygon in a loop; Among them, the first standardized judgment operation includes: Extracting the center line of the polygon according to the pixel coordinates of the polygon, and determining whether the center line meets a preset condition; If yes, the center line is saved and the standardization judgment operation is no longer performed on the polygon; If not, the next standardization judgment operation is performed; The next standardization judgment operation includes: Simplify the polygon, extract the center line according to the pixel coordinates of the simplified polygon, and determine whether the newly extracted center line meets the preset conditions; If yes, the center line is saved and the standardization judgment operation is no longer performed; If not, the next standardization judgment operation is performed; If a center line that meets the preset conditions is obtained within a preset number of standardization judgment operations, the polygon meets the standardization conditions; if a center line that meets the preset conditions is not obtained within a preset number of standardization judgment operations, the polygon does not meet the standardization conditions.

7. An object detection device in an image, characterized in that: include: at least one processor; as well as, a memory communicatively coupled to the at least one processor; in, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor, so that the at least one processor can execute the method for detecting objects in images according to any one of claims 1 to 5.

8. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer-executable instructions, and when the computer-executable instructions are executed by a processor, the method for detecting an object in an image according to any one of claims 1 to 5 is implemented.

Citation Information

Patent Citations

  • Method and device for determining image recognition model

    CN114580517A