Method and device for updating image recognition result correction model and storage medium

The image recognition result correction model corrects the preliminary identification results of the drone detection system, and uses user feedback to train and update the model, solving the problems of high false alarm rate and high false alarm rate in complex environments, achieving higher recognition accuracy and adaptability.

CN120032276APending Publication Date: 2025-05-23AUTEL INTELLIGENT AUTOMOBILE CORP LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510095345.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-21
Publication Date
2025-05-23

AI Technical Summary

Technical Problem

The existing drone detection system has high false alarm rates and high false alarm rates in complex environments, making it difficult to effectively identify and monitor drones.

Method used

The preliminary recognition results are corrected through the image recognition result correction model, and labeled and trained in combination with user feedback, and dynamically updated model parameters to improve recognition accuracy.

Benefits of technology

It improves the adaptability, flexibility and accuracy of drone detection, reduces the false alarm and missed alarm rates, and enhances the corrective ability of the model in different scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120032276A_ABST
    Figure CN120032276A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of unmanned aerial vehicles, and discloses an updating method and device of an image recognition result correction model and a storage medium, the method comprises the steps that a detection image and a preliminary recognition result of the detection image are acquired, and the preliminary recognition result comprises a target in the detection image and a prediction category of the target; correcting the prediction category through an image recognition result correction model to obtain a corrected recognition result; displaying the detection image and the corrected recognition result; obtaining a confirmation result indicating whether the prediction category in the corrected recognition result input by the user is correct or not; setting a label for the target in the detection image according to the confirmation result; and training an image recognition result correction model by taking the detection image after setting and labeling as training data, and updating parameters of the image recognition result correction model. In this way, the accuracy of the displayed image recognition result is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present application relate to the field of drone technology, and specifically to a method and device for updating an image recognition result correction model, a target recognition and display system, and a storage medium. Background Art

[0002] With the widespread use of drones and their rapid development in civil, commercial and military fields, related security issues have become increasingly prominent. The illegal use of drones may pose threats to air safety, personal privacy and public safety. Therefore, drone monitoring technology has become an important part of the modern security system.

[0003] Current drone detection systems mainly rely on single sensors, such as radar or visual sensors to detect drones, but these single detection methods have limited effectiveness in complex environments and have problems such as high false alarm rates and high missed alarm rates. Summary of the invention

[0004] In view of the above problems, the embodiments of the present application provide a method, device, target recognition and display system and storage medium for updating an image recognition result correction model, which are used to solve the problems of high false alarm rate and high missed alarm rate when detecting drones in the prior art.

[0005] According to one aspect of an embodiment of the present application, a method for updating an image recognition result correction model is provided, the method comprising: obtaining a detection image and a preliminary recognition result of the detection image, the preliminary recognition result comprising a target in the detection image and a predicted category of the target; correcting the predicted category through an image recognition result correction model to obtain a corrected recognition result; displaying the detection image and the corrected recognition result; obtaining a confirmation result input by a user, wherein the confirmation result is used to indicate whether the predicted category in the corrected recognition result is correct; setting a label for the target in the detection image according to the confirmation result; training the image recognition result correction model using the detection image with the label as training data, and updating the parameters of the image recognition result correction model.

[0006] In an optional manner, the annotation includes a first annotation and a second annotation, and setting the annotation for the target in the detection image according to the confirmation result includes: if the confirmation result is that the predicted category in the corrected recognition result is correct, setting the first annotation for the target in the detection image; if the confirmation result is that the predicted category in the corrected recognition result is wrong, setting the second annotation for the target in the detection image.

[0007] In an optional manner, the detected image includes multiple targets, and the preliminary recognition result also includes the confidence corresponding to the predicted category. The method also includes: if the confirmation result input by the user for some of the multiple targets is not obtained, setting a pseudo-label for the partial target in the detected image according to the confidence corresponding to the predicted category of the partial target; using the detection image after setting the annotation and the pseudo-label as training data, training the image recognition result correction model, and updating the parameters of the image recognition result correction model.

[0008] In an optional manner, the pseudo-identifier includes a first pseudo-identifier and a second pseudo-identifier, and the pseudo-identifier is set for the partial target in the detected image according to the confidence corresponding to the predicted category of the partial target, including: for each target in the partial target: if the confidence corresponding to the predicted category of the target in the preliminary recognition result is greater than or equal to a first confidence threshold, then the first pseudo-identifier is set for the target in the detected image; if the confidence corresponding to the predicted category of the target in the preliminary recognition result is less than or equal to a second confidence threshold, then the second pseudo-identifier is set for the target in the detected image, wherein the second confidence threshold is less than the first confidence threshold.

[0009] In an optional manner, the preliminary recognition result also includes a target box, which is used to represent the position information of the target in the detected image. Before displaying the detected image and the corrected recognition result, the method also includes: adjusting the target box through the image recognition result correction model so that the adjusted target box correctly represents the position information of the target in the detected image.

[0010] In an optional manner, the preliminary recognition result also includes a target box and a confidence level corresponding to the predicted category, and the target box is used to indicate position information of the target in the detected image; and displaying the detected image and the corrected recognition result includes: displaying the detected image and displaying the target box and the predicted category according to the confidence level.

[0011] In an optional manner, the preliminary recognition result also includes a confidence level corresponding to the predicted category, and the correcting the predicted category by using an image recognition result correction model includes: the image recognition result correction model correcting the predicted category according to the confidence level.

[0012] According to another aspect of an embodiment of the present application, a device for updating an image recognition result correction model is provided, comprising a memory, a processor, and a computer program stored on the memory, wherein the processor executes the computer program to implement the method for updating the image recognition result correction model as described above.

[0013] According to another aspect of an embodiment of the present application, a target recognition and display system is provided, including a radar and a visual device and an updating device for the image recognition result correction model as described above, wherein the radar and the visual device are used to detect a control area to obtain a detection image, recognize the detection image to obtain a preliminary recognition result, and transmit the detection image and the recognition result to the updating device for the image recognition result correction model, wherein the preliminary recognition result includes a target in the detection image and a predicted category of the target.

[0014] According to another aspect of an embodiment of the present application, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the method for updating the image recognition result correction model as described above is implemented.

[0015] In the embodiment of the present application, after obtaining the preliminary recognition result of the detection image, the predicted category in the preliminary recognition result is corrected by the image recognition result correction model, and then the detection image is annotated according to the user's feedback information (i.e., the confirmation result of whether the predicted category in the corrected recognition result input by the user is correct or not). Since the annotation is set based on the confirmation result input by the user, the quality of the annotated data is relatively high. Then, the image recognition result correction model is trained and updated using the annotated detection image, so that the model can be dynamically adjusted to improve the accuracy of the model's correction of the preliminary recognition result, thereby improving the adaptability, flexibility and accuracy of drone detection. In addition, by effectively utilizing the data fed back by users, the model can be continuously optimized to enhance the model's ability to correct the preliminary recognition results of the detection images obtained in different scenarios, thereby reducing the false positives and false negatives in the displayed recognition results.

[0016] The above description is only an overview of the technical solution of the embodiment of the present application. In order to more clearly understand the technical means of the embodiment of the present application, it can be implemented in accordance with the contents of the specification. In order to make the above and other purposes, features and advantages of the embodiment of the present application more obvious and easy to understand, the specific implementation methods of the present application are listed below. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] The accompanying drawings are only used to illustrate the embodiments and are not to be considered as limiting the present application. In addition, the same reference symbols are used to represent the same components throughout the accompanying drawings. In the accompanying drawings:

[0018] Figure 1 A schematic diagram of the structure of a target recognition and display system provided in an embodiment of the present application is shown;

[0019] Figure 2 A schematic diagram showing the structure of a device for updating an image recognition result correction model provided by an embodiment of the present application is shown;

[0020] Figure 3 A flowchart of a method for updating an image recognition result correction model provided by an embodiment of the present application is shown;

[0021] Figure 4 A schematic diagram of an interface for displaying detected images and predicted categories provided in an embodiment of the present application is shown. DETAILED DESCRIPTION

[0022] The exemplary embodiments of the present application will be described in more detail below with reference to the accompanying drawings. Although the exemplary embodiments of the present application are shown in the drawings, it should be understood that the present application can be implemented in various forms and should not be limited by the embodiments set forth herein.

[0023] Radar sensors have the ability to detect all-weather and long-distance detection, and can detect information such as the distance, speed and direction of drones, especially at night or in bad weather. However, radar detection is easily affected by object reflection and multipath interference in complex environments, resulting in a high false alarm rate. Visual sensors can provide rich images and image data to assist in identifying the appearance characteristics of the target, but the accuracy of visual recognition is poor in poor lighting conditions or when obstacles block it. In addition, due to the complex and changeable appearance and movement patterns of drones, a single radar or visual detection system cannot meet the needs of all scenarios, especially in densely populated urban areas or complex terrain.

[0024] In order to solve the above problems, radar and visual sensor fusion technology has gradually attracted attention. Radar and visual fusion technology can effectively combine the long-range detection capability of radar and the high-resolution image recognition capability of vision to improve the target detection accuracy and reliability of the system. However, although radar and visual fusion technology can improve the performance of drone detection, the existing system still faces the problems of poor environmental adaptability and insufficient model generalization ability. Especially when dealing with complex environmental changes and new types of drones, the flexibility and adaptability of traditional algorithms are limited, resulting in the continued existence of false alarms and missed alarms.

[0025] The inventors of the present application have discovered that after using radar and vision fusion technology to detect a drone and obtain a detection image and preliminary recognition results, if the preliminary recognition results are corrected using an image recognition result correction model and then the corrected recognition results are displayed, the accuracy of the displayed recognition results can be improved, thereby reducing the false alarm rate and the missed alarm rate.

[0026] However, as mentioned above, the appearance and movement patterns of drones are complex and changeable. If the parameters of the image recognition result correction model remain unchanged, then when encountering new models of drones or abnormal flight conditions, the image recognition result correction model's ability to correct the initial recognition results may decline.

[0027] Based on this, the present application proposes a method for updating an image recognition result correction model, which obtains a detection image and a preliminary recognition result of the detection image, corrects the preliminary recognition result through the image recognition result correction model and displays the detection image and the corrected recognition result, then obtains a confirmation result of whether the corrected recognition result input by the user is correct, annotates the detection image according to the confirmation result and uses the annotated image as training data, uses the training data to train the image recognition result correction model and updates the parameters of the model, thereby improving the correction ability of the model in different scenarios and reducing the false alarm rate and the missed alarm rate.

[0028] Figure 1 FIG. 1 shows a schematic diagram of the structure of the target recognition and display system provided by the embodiment of the present application. Figure 1 As shown, the target recognition and display system 10 includes a radar and vision device 100 and an updating device 200 for correcting a model of image recognition results.

[0029] The radar and vision device 100 includes a radar sensor and a vision sensor. The radar and vision device 100 detects the control area through the radar sensor to determine the distance, speed and direction of the target, and shoots the control area through the vision sensor to obtain a detection image, and then identifies the content in the detection image to obtain a preliminary recognition result. The preliminary recognition result includes the target in the detection image and the predicted category of the target, wherein the predicted category of the target is the category to which the radar and vision device 100 identifies and predicts the target. Then the radar and vision device 100 transmits the detection image and the preliminary recognition result to the update device 200 of the image recognition result correction model.

[0030] The updating device 200 of the image recognition result correction model is used to execute the updating method of the image recognition result correction model provided in the embodiment of the present application. The updating device 200 of the image recognition result correction model can be a device including one or more processors, such as a touch-screen mobile phone, a smart phone, a tablet computer, a portable electronic device or other electronic device with a display screen. The processor may be a central processing unit CPU, or an application-specific integrated circuit (ASIC), or one or more integrated circuits configured to implement the embodiment of the present invention, which is not limited here. The one or more processors included in the updating device 200 of the image recognition result correction model can be processors of the same type, such as one or more CPUs; or they can be processors of different types, such as one or more CPUs and one or more ASICs, which are not limited here.

[0031] Since the radar and vision device 100 continuously provides the detection images and preliminary recognition results to the updating device 200 of the image recognition result correction model. These preliminary recognition results will help the updating device 200 of the image recognition result correction model to further optimize the image recognition result correction model. For example, if there are many false positives in the preliminary recognition results provided by the radar and vision device 100, the updating device 200 of the image recognition result correction model can further fine-tune the optimization model according to the user's feedback on the false positives, so that the image recognition result correction model can correct the false positives.

[0032] Figure 2 The schematic diagram of the structure of the updating device of the image recognition result correction model provided by the embodiment of the present application is shown. The specific embodiment of the present application does not limit the specific implementation of the updating device of the image recognition result correction model.

[0033] like Figure 2 As shown, the updating device 200 for the image recognition result correction model includes: a processor (processor) 202 and a memory (memory) 204.

[0034] The memory 204 is used to store a computer program 206. The memory 204 may include a high-speed RAM memory, or may also include a non-volatile memory, such as at least one disk memory. The computer program 206 may include computer executable instructions.

[0035] The processor 202 is used to execute the computer program 206 to implement the method for updating the image recognition result correction model provided in the embodiment of the present application.

[0036] The processor 202 may be a central processing unit (CPU), or an application-specific integrated circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of the present application. The one or more processors included in the image recognition result correction model updating device 200 may be processors of the same type, such as one or more CPUs; or may be processors of different types, such as one or more CPUs and one or more ASICs.

[0037] Figure 3 FIG. 2 is a flowchart of a method for updating an image recognition result correction model provided by an embodiment of the present application, and the method is executed by an updating device 200 for updating an image recognition result correction model. Figure 3 As shown, the method comprises the following steps:

[0038] Step 110: Obtain a detection image and a preliminary recognition result of the detection image.

[0039] Among them, in this step, the updating device 200 of the image recognition result correction model obtains the detection image sent by the radar and vision device 100 and the preliminary recognition result of the detection image.

[0040] The preliminary recognition result includes the target in the detection image and the predicted category of the target. It should be noted that the target in the embodiment of the present application refers to a target that can move, such as a drone, a bird, etc., not just a drone. The predicted category refers to the category of the target obtained by identifying and predicting the target in the detection image based on the content of the detection image. The categories of the target include drones, birds, etc.

[0041] Step 120: Correct the predicted category using the image recognition result correction model to obtain a corrected recognition result.

[0042] The image recognition result correction model updating device 200 is equipped with a pre-trained image recognition result correction model, which is used to correct the predicted category in the preliminary recognition result. The image recognition result correction model can be a known model such as a semi-supervised learning model and an incremental learning model.

[0043] Step 130: Display the detected image and the corrected recognition result.

[0044] Figure 4 FIG. 1 shows a schematic diagram of an interface for displaying detected images and predicted categories provided by an embodiment of the present application. Figure 4As shown, the detection image includes three targets, namely the first target 101, the second target 102 and the third target 103. And the predicted category of each target is displayed next to each target. In the figure, the predicted category of the first target 101 is a drone, the predicted category of the second target 102 is a drone, and the predicted category of the third target 103 is a bird. Among them, the predicted category shown in the figure is the predicted category in the corrected recognition result.

[0045] Step 140: Obtain the confirmation result input by the user, wherein the confirmation result is used to indicate whether the predicted category in the corrected recognition result is correct or not.

[0046] like Figure 4 As shown, Figure 4 The image recognition result correction model includes three targets. For the convenience of user operation, the user can select a target by clicking on the screen, and after entering the confirmation result of whether the predicted category of the target is correct, the updating device 200 of the image recognition result correction model can obtain the confirmation result of the predicted category of the target entered by the user. For example, the user can confirm the correct drone target by clicking on a target in the detection image displayed on the screen or by selecting a target in the detection image displayed on the screen, and mark it as a real drone target. The user can also delete the falsely reported target and add the mark of the missed drone target.

[0047] In order to accurately obtain the confirmation result input by the user, for example, after the user clicks on the first target 101 and enters the confirmation result, in order to avoid mistaking the confirmation result for the confirmation result of the second target 102, in the embodiment of the present application, it is assumed that the screen coordinates clicked by the user on the screen are Ps (xs, ys), and the coordinates corresponding to the image content displayed by the screen coordinates on the screen in the detection image are the detection image coordinates Pv (xv, yv), the width and height of the screen are Ws and Hs respectively, and the width and height of the detection image are Wv and Hv respectively. The screen aspect ratio As is determined by the formula As = Ws / Hs, and the detection image aspect ratio Av is determined by the formula Av = Wv / Hv. When the height of the screen matches the height of the detection image, if As>Av, it means that there are black borders on the left and right ends of the screen when the detection image is displayed (i.e., there is no detection image content), then the scaling factor scalefactor is determined by the formula scalefactor=Ws / Wv, the width of the detection image Wv′ after scaling is Wv′=Wv×scalefactor, and the horizontal offset offsetx (i.e., the width of the black border) is: offsetx=(Ws-Wv′) / 2. If As≤Av, it means that there are black borders on the upper and lower ends of the screen when the detection image is displayed (i.e., there is no detection image content), then the scaling factor scalefactor is determined by the formula scalefactor=Hs / Hv, the height of the detection image Hv′ after scaling is Hv′=Hv×scalefactor, and the vertical offset offsety (i.e., the height of the black border) is offsety=(Hs-Hv′) / 2. The horizontal coordinate xv in the detection image coordinate is determined by the formula xv=(xs-offsetx) / scalefactor, and the vertical coordinate yv in the detection image coordinate is determined by the formula yv=(ys-offsety) / scalefactor.

[0048] For example, the screen width Ws = 1920 pixels, and the screen height Hs = 1080 pixels. The detection image width Wv = 1280 pixels, and the detection image height Hv = 960 pixels. The screen aspect ratio As is As = Ws / Hs = 1920 / 1080 = 1.777. The detection image aspect ratio Av is Av = Wv / Hv = 1280 / 960 = 1.333. Since As>Av, there are black edges at both ends of the screen in the horizontal direction (left and right direction). The scaling factor scalefactor (filling the screen width horizontally) is scalefactor = Ws / Wv = 1920 / 1280 = 1.5. The height Hv′ of the scaled detection image is Hv′ = Hv×scalefactor = 960×1.5 = 1440 pixels. Since the height Hv′ = 1440 pixels of the scaled detection image exceeds the screen height Hs = 1080 pixels, this means that the detection image will have black edges in the vertical direction. The height of the black border offsety is offsety = (Hv′-Hs) / 2 = (1440-1080) / 2 = 180 pixels, so there are 180 pixels of black border above and below the detected image. Assume that the screen position clicked by the user is Ps(1000,500), which needs to be converted to the position Pv(xv,yv) in the detected image. First, calculate the horizontal coordinate conversion: xv = xs / scalefactor = 1000 / 1.5 = 666.67; secondly, calculate the vertical coordinate conversion: in the vertical direction, the screen click coordinates need to subtract the offset and then scale, so yv = (ys+offsety) / scalefactor = (500+180) / 1.5 = 680 / 1.5 = 453.33. In other words, the position Ps(1000,500) clicked by the user on the screen corresponds to the position Pv(666.67,453.33) in the detected image.

[0049] Furthermore, in order to make the position of the screen clicked by the user belong to a valid position, that is, the position of the screen clicked by the user belongs to the position of the detection image displayed that includes the detection image content. In an embodiment of the present application, the clickable range of the user on the screen can be limited so that the position clicked by the user is in the visible area of ​​the detection image. Specifically, if As>Av, the horizontal limit coordinate xs is xs=max(offsetx,min(xs,offsetx+Wv′)), and the vertical limit coordinate ys is ys=max(0,min(ys,Hs)). If As≤Av, the horizontal limit coordinate xs is xs=max(0,min(xs,Ws)), and the vertical limit coordinate ys is ys=max(offsety,min(ys,offsety+Hv′))

[0050] Step 150: Setting a label for the target in the detection image according to the confirmation result.

[0051] In some embodiments, the annotation includes a first annotation and a second annotation. If the confirmation result is that the predicted category in the corrected recognition result is correct, a first annotation (e.g., a positive sample) is set for the target in the detection image. If the confirmation result is that the predicted category in the corrected recognition result is wrong, a second annotation (e.g., a negative sample) is set for the target in the detection image. For example, if the confirmation result is that the first target 101 belongs to a drone, a first annotation is set for the first target 101; if the confirmation result is that the first target 101 does not belong to a drone, a second annotation is set for the first target 101.

[0052] Step 160: using the annotated detection image as training data to train the image recognition result correction model, and updating the parameters of the image recognition result correction model.

[0053] Among them, the image recognition result correction model is trained by using the annotated detection image, and the model's weights, biases and other parameters are updated to improve the accuracy of the model, so that when the predicted category in the preliminary recognition result sent by the radar and visual device 100 is subsequently corrected, the accuracy of the predicted category in the corrected recognition result can be improved.

[0054] In the embodiment of the present application, after obtaining the preliminary recognition result of the detection image, the predicted category in the preliminary recognition result is corrected by the image recognition result correction model, and then the detection image is annotated according to the user's feedback information (i.e., the confirmation result of whether the predicted category in the corrected recognition result input by the user is correct or not). Since the annotation is set based on the confirmation result input by the user, the quality of the annotated data is relatively high. Then, the image recognition result correction model is trained and updated using the annotated detection image, so that the model can be dynamically adjusted to improve the accuracy of the model's correction of the preliminary recognition result, thereby improving the adaptability, flexibility and accuracy of drone detection. In addition, by effectively utilizing the data fed back by users, the model can be continuously optimized to enhance the model's ability to correct the preliminary recognition results of the detection images obtained in different scenarios, thereby reducing the false positives and false negatives in the displayed recognition results.

[0055] If there are many targets in the detected image, and the user only confirms the predicted categories of some of the targets, but does not confirm the predicted categories of some of the targets, in order to use the targets whose predicted categories the user has not confirmed as training data to further improve the accuracy of the model. In the embodiment of the present application, the method further includes the following steps:

[0056] Step a1: If the confirmation result input by the user for some of the multiple targets is not obtained, a pseudo mark is set for the part of the target in the detection image according to the confidence corresponding to the predicted category of the part of the target.

[0057] The preliminary recognition result also includes the confidence level corresponding to the predicted category. The confidence level corresponding to the predicted category refers to the probability that the predicted category of the target is correct.

[0058] In some embodiments, the pseudo-identifier includes a first pseudo-identifier and a second pseudo-identifier. For each target for which no confirmation result is obtained: if the confidence corresponding to the predicted category of the target is greater than or equal to the first confidence threshold, a first pseudo-identifier (e.g., a pseudo-positive sample) is set for the target in the detected image; if the confidence corresponding to the predicted category of the target is less than or equal to the second confidence threshold, a second pseudo-identifier (e.g., a pseudo-negative sample) is set for the target in the detected image. Among them, the second confidence threshold is less than the first confidence threshold, and the first confidence threshold and the second confidence threshold can be set as needed, for example, the first confidence threshold is set to 90% or 95%, and the second confidence threshold is set to 40% or 50%.

[0059] For example, for a target x that has not obtained a determination result input by the user, the pseudo-identifier y^ is determined by the formula y^=argymaxP(y|x), where P(y|x) is the confidence level. If P(y|x)≥threshold, a pseudo-positive sample identifier is set for the target, where threshold is the first confidence threshold, such as 90%. If P(y|x)≤low threshold, a pseudo-negative sample identifier is set for the target or the target is ignored, where low threshold is the second confidence threshold, such as 50%.

[0060] It can be understood that if the confidence of the predicted category is high (the confidence is greater than or equal to the first confidence threshold), it means that the accuracy of the predicted category is high, so a pseudo-positive sample identifier is set for the target. If the confidence of the predicted category is low (the confidence is less than or equal to the second confidence threshold), it means that the accuracy of the predicted category is low, so a pseudo-negative sample identifier is set for the target. For targets whose confidence in the predicted category is greater than the second confidence threshold and less than the first confidence threshold, it means that the predicted category cannot be fully trusted, so a pseudo-positive sample identifier is not set for the target. Since the confidence of the target is not low enough, a pseudo-negative sample identifier is not set for the target, that is, the target will not be used as training data, avoiding the introduction of too much noise (such as erroneous pseudo-positive samples or pseudo-negative samples) into the training data, so that most of the data in the training data is high-quality data, thereby improving the accuracy of the training data and improving the effect of training the model using the training data.

[0061] Step a2: using the detection images with annotations and pseudo-labels as training data, training the image recognition result correction model, and updating the parameters of the image recognition result correction model.

[0062] In this step, the total loss function Ltotal is composed of the supervised learning part Lsupervised and the unsupervised learning part Lunsupervised, with a balance coefficient α. Among them, Ltotal = αLsupervised + (1-α), Lsupervised is trained using labeled data, and Lunsupervised is trained using pseudo-labeled data.

[0063] It is understandable that if the quality of the training data meets the requirements, the more training data there is, the better the effect of training the model and the higher the accuracy of the model. In an embodiment of the present application, for a target that has not obtained the confirmation result of the user input, a corresponding pseudo-identification is set for the target according to the relationship between the confidence corresponding to the predicted category of the target and the first confidence threshold, and the detection image after setting the annotation and pseudo-identification is used as the training data, and then all the targets in the detection image can be used as the training data, thereby increasing the amount of training data and further improving the accuracy of the model. In addition, in an embodiment of the present application, by setting annotations for the detection image in combination with user feedback, and using semi-supervised learning to generate pseudo-identification data, the workload of the user inputting feedback information can be reduced, and the detection performance of the model can be improved. Specifically, semi-supervised learning is embodied in training the model using training data with annotations and training data without annotations (i.e., training data with pseudo-identification). Incremental learning is reflected in the addition of new incremental data, that is, setting labeled training data and unlabeled pseudo-identification data based on the confirmation results input by the user to enhance the model. The specific implementation method can be to impose constraints through the loss function of the new task to protect old knowledge from being overwritten by new knowledge, or to prevent the model from forgetting old knowledge through a playback mechanism, etc. The weights and parameters of the model are continuously fine-tuned and updated in the above manner, so that when the model corrects the preliminary recognition results, it can improve the accuracy of the corrected recognition results. The optimized model will be more accurate in target classification, positioning and confidence estimation, reducing false positives and false negatives, especially in the judgment of targets with low confidence.

[0064] Moreover, after each drone inspection of the control area, new training data is determined based on the new detection images and the model is trained, which can gradually improve the accuracy of the model and enable the model to adapt to changes in the environment and drones, thereby improving the accuracy of the corrected recognition results obtained by correcting the initial recognition results.

[0065] In some embodiments, the preliminary recognition result also includes a target frame, which is used to indicate the location information of the target in the detected image. Figure 4 As shown, taking the first target 101 as an example, the first target 101 is within the target frame 11, so that the position information of the first target 101 in the detection image can be reflected by the position information of the target frame 11 in the detection image. In the embodiment of the present application, before step 130, the method further includes: adjusting the target frame by correcting the model of the image recognition result, so that the adjusted target frame correctly represents the position information of the target in the detection image.

[0066] The image recognition result correction model adjusts the target frame, which may be to adjust the position and / or size of the target frame, so that the target frame can more accurately represent the position information of the target in the detected image. For example, before the image recognition result correction model adjusts the target frame, only part of the target area is located in the target frame, and another part of the target area is located outside the target frame. In this step, the image recognition result correction model adjusts the position and / or size of the target frame so that the target is completely located in the target frame.

[0067] After executing this step, go to step 130.

[0068] It should be noted that, in some embodiments, after executing step 130 to display the detected image and the corrected recognition result, since the displayed corrected recognition result includes a target frame, the user can also input a confirmation result for the target frame to determine whether the position and size of the target frame in the corrected recognition result are correct. The user can also adjust the position and / or size of the target frame so that the target frame can accurately represent the position information of the target in the detected image. In step 150, a label is also set for the target frame based on the confirmation result of the user input for the target frame and the user's operation of adjusting the target frame. Therefore, in step 160, after the model is trained and the parameters of the model are updated using the detected image after the label is set as training data, the subsequent model can more accurately adjust the target frame when adjusting the target frame in the preliminary recognition result, so that the target frame can more accurately represent the position information of the target in the detected image.

[0069] It is understandable that if the lighting in the control area is dim, it is difficult to distinguish the background and the target in the detection image. For example, when a drone is used to detect a control area at night, the background color in the detection image obtained may be close to black, and it is difficult for the user to distinguish the background and the target in the detection image. Therefore, in an embodiment of the present application, the position information of the target in the detection image is represented by using a target frame 11, for example, a target frame 11 of a color different from the background color of the detection image is used to frame the target in the detection image, so that the user can quickly identify the target in the detection image based on the target frame 11.

[0070] In some embodiments, step 130 includes: displaying the detected image and displaying the target box and the predicted category according to the confidence level. The preliminary recognition result also includes the confidence level corresponding to the target box and the predicted category. As described above, the confidence level corresponding to the predicted category refers to the probability that the predicted category is correct. The higher the confidence level, the more accurate the predicted category.

[0071] Specifically, in this step, when displaying the detection image, it can be based on the display of the detection image, only the target frame and the corresponding prediction category of the target with high confidence can be displayed, or only the target frame and the corresponding prediction category of the target with high confidence and predicted category as drone can be displayed. For example, the predicted category of the first target is drone, and the corresponding confidence is 50%, the predicted category of the second target is drone, and the corresponding confidence is 90%, and the predicted category of the third target is bird, and the corresponding confidence is 95%. In this step, since the confidence of the predicted category corresponding to the first target is low, when displaying the detection image, the predicted category of the second target and the target frame of the second target, as well as the predicted category of the third target and the target frame of the third target can be displayed at the same time; since the predicted category of the third target is not drone, only the predicted category of the second target and the target frame of the second target can be displayed. Among them, the level of confidence is based on the relationship between the confidence and the third confidence threshold. For example, if the confidence level corresponding to the prediction category of a target is greater than or equal to the third confidence threshold, it means that the confidence level of the prediction category of the target is high; if the confidence level corresponding to the prediction category of a target is less than the third confidence threshold, it means that the confidence level of the prediction category of the target is low. The third confidence threshold can be determined as needed, for example, the third confidence threshold is 80%, 90% or 95%.

[0072] It is understandable that when there are many targets in the detection image, if each target is framed with a target frame, the target frames in the detection image displayed on the screen may be dense, which is not conducive to the user to match the target with the target frame one by one. Therefore, in an embodiment of the present application, when displaying the detection image, by displaying the target frame and the predicted category according to the confidence level, it is possible to avoid the situation where the target frames in the displayed detection image are dense, thereby facilitating the user to focus on the key targets more efficiently from more targets. The optimized model will be more accurate in target classification, positioning and confidence estimation, reducing false positives and negatives, especially in the determination of target frames with low confidence levels.

[0073] Targets with higher confidence levels corresponding to the predicted categories are usually more important detection results. By only displaying the target frames of these targets in the detection image of the current frame, users can quickly determine whether the displayed predicted categories are accurate and focus on marking these key targets. In addition, if all target frames are displayed on the interface at the same time, users may not be able to quickly locate the targets they are really interested in. When displaying target frames, by only displaying the target frames of targets with higher confidence levels corresponding to the predicted categories, the display interface can reduce the interference caused by irrelevant target frames to users.

[0074] In some embodiments, step 120 includes: the image recognition result correction model corrects the predicted category according to the confidence level. Among them, the preliminary recognition result includes the confidence level corresponding to the predicted category. Specifically, in the case of low confidence (the confidence level corresponding to the predicted category is less than the third confidence threshold), the image recognition result correction model can correct the predicted category in the preliminary recognition result in combination with the confirmation result input by the user at the historical moment and the record of correcting the displayed corrected recognition result. For example, at a historical moment, if the user changes the low-confidence "bird" to "drone" many times, the image recognition result correction model will learn the user's correction pattern, and then the image recognition result correction model will directly correct the "bird" to "drone" for similar situations.

[0075] In some embodiments, by continuously detecting the control area and executing the updating method of the image recognition result correction model provided by the embodiments of the present application, the model is incrementally updated by obtaining feedback information continuously provided by the user and training the model based on the feedback information to gradually update the model in small scales without completely retraining the entire model from scratch. This method is particularly suitable for online or real-time drone detection scenarios because it can adapt to new data distribution, target types, and detection requirements based on changes in the real-time environment, while maintaining memory of old data. This is particularly important in drone detection because the type of drone, flight mode, and environmental conditions may change at any time.

[0076] Whenever the detection image is annotated based on user feedback (such as confirmation or correction of the predicted category), these newly annotated data and the pseudo-labeled data generated by semi-supervised learning will be input into the model as incremental data to further optimize the performance of the model. Incremental learning reduces the consumption of computing resources by gradually updating model parameters instead of completely retraining the entire model, while allowing the system to quickly adapt to new target features.

[0077] Among them, the core of incremental learning is how to balance the learning of new data and old data. Each time the model is incrementally updated, it combines the previously learned data to perform small-scale training and adjustments. In order to prevent the model from over-adapting to new data and forgetting old knowledge, the loss function used in incremental learning includes the loss of old data and the loss of new data, and adjusts the weights of the two through the balance coefficient α\alphaα. Among them, Lincremental=αLold+(1-α)Lnew, where Lold represents the loss of old data, ensuring that the model will not forget the previously accumulated knowledge due to learning new data. Lnew represents the loss of new data, ensuring that the model can adapt to new detection needs based on new user feedback and environment. α is the balance coefficient for adjusting the learning ratio of old data and new data. This loss function design enables the model to retain old knowledge while gradually learning new knowledge, reducing the forgetting effect.

[0078] In deep learning models, it is usually not necessary to make large-scale adjustments to the entire network structure, but only to make small-scale updates to the last few layers related to high-order features. In this way, the model retains its memory of old features while absorbing new knowledge. Specifically, the first few layers of the network can be frozen. These layers usually learn common features (such as edges and textures of images) and do not need to be updated frequently. Only fine-tuning the high-level layers of the network (such as the last few layers) enables the model to learn new target features, such as new drone shapes or flight modes.

[0079] After each new detection image is acquired, the model is adjusted in small scale by updating the model. Online learning is well suited for real-time scenarios of drone detection because it can process new data quickly and does not require the accumulation of large amounts of new data for training. In each detection image, the feedback information entered by the user is immediately used to adjust the model parameters. The model can prevent the model from forgetting the features of old data when learning new data through a "memory playback" mechanism. In this mechanism, when the model is trained on new data, it randomly replays a portion of the old data to ensure that the model retains the old knowledge.

[0080] In summary, incremental learning ensures that the model can continuously improve its ability to correct initial recognition results and adapt to new targets and environmental characteristics by gradually updating the model in small scales. By combining user feedback and pseudo-identification data, an adaptive feedback loop can be formed for the model, so that the model is gradually optimized in each round of detection. This incremental update mechanism makes the drone detection system more efficient and accurate, while reducing the amount of manual intervention by users and improving the robustness and scalability of the model.

[0081] In some embodiments, when the model is updated, semi-supervised learning and incremental learning are used simultaneously to optimize the performance of the model. Setting a pseudo-label for the target according to the confidence level is the semi-supervised learning stage. For targets that have not obtained feedback information input by the user, a pseudo-label is set for the target according to the confidence level of the target, wherein a pseudo-positive sample label is set for a target with high confidence, and a pseudo-negative sample label is set or ignored for a target with low confidence. The model performs incremental learning by using the labeled data and the pseudo-label data generated by semi-supervised learning. Incremental learning enables the model to learn new data without forgetting old knowledge by gradually fine-tuning the parameters of the model. The loss function of incremental learning includes two parts. The first part is the supervised learning part: setting labels for the detection images based on the user's feedback information, and using the detection image data after the labels are set for supervised learning to ensure that the model learns the labeled positive and negative samples. The second part is the unsupervised learning part: using the pseudo-label data generated by semi-supervised learning for unsupervised learning to expand the training data set. The loss function is expressed as Ltotal=αLsupervised+(1 - α)Lunsupervised, where Lsupervised represents the supervised learning loss of labeled data and Lunsupervised represents the unsupervised learning loss of pseudo-labeled data. By adjusting α, the data ratio of supervised learning and unsupervised learning can be balanced.

[0082] In this way, for the target recognition model, the combination of semi-supervised learning and incremental learning achieves continuous optimization of the model by effectively utilizing user feedback, pseudo-identification data, and a gradual update strategy. Semi-supervised learning reduces the reliance on user input confirmation results, while incremental learning ensures that the model can adapt to new data and continuously improve detection accuracy while maintaining memory of old data. This combination mechanism shows great advantages in complex drone detection scenarios, especially when dealing with dynamically changing targets and environments.

[0083] An embodiment of the present application provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, an embodiment of the method for updating the image recognition result correction model is implemented.

[0084] An embodiment of the present application provides a computer program, which can be executed by a processor to implement the above-mentioned method embodiment for updating the image recognition result correction model.

[0085] An embodiment of the present application provides a computer program product, which includes a computer program. When the computer program is executed by a processor, an embodiment of the method for updating the image recognition result correction model is implemented.

[0086] In several embodiments provided in the present application, if any function is implemented in the form of a software function module / unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, all or part of the technical solution of the present application can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions for a computer device (which can be a personal computer, server or other electronic device) to perform all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (Read-Only Memory, ROM), random access memory (RandomAccess Memory, RAM), disk or optical disk and other media that can store computer program codes.

[0087] The algorithm or display provided here are not inherently related to any specific computer, virtual system or other equipment. Various general systems can also be used together with the teaching based on this. According to the above description, it is obvious to construct the structure required for this type of system. In addition, the present application embodiment is not directed to any specific programming language yet. It should be understood that various programming languages ​​can be utilized to realize the content of the present application described here, and the above description of specific languages ​​is to disclose the best mode of implementation of the present application.

[0088] It should be noted that the above embodiments illustrate the present application rather than limit the present application, and that those skilled in the art may design alternative embodiments without departing from the scope of the appended claims. In the claims, any reference symbol between brackets shall not be constructed as a limitation on the claims. The word "comprising" does not exclude the presence of elements or steps not listed in the claims. The word "one" or "an" preceding an element does not exclude the presence of multiple such elements. The present application may be implemented by means of hardware including several different elements and by means of a suitably programmed computer. In the claims that list several devices, several units or modules in these devices may be embodied by the same hardware item. The use of the words first, second, and third, etc. does not indicate any order. These words may be interpreted as names. The steps in the above embodiments, unless otherwise specified, should not be understood as limitations on the order of execution.

[0089] The above-described embodiments only express several implementation methods of the present application, and the descriptions thereof are relatively specific and detailed, but they cannot be understood as limiting the scope of the present application. It should be pointed out that, for a person of ordinary skill in the art, several variations and improvements can be made without departing from the concept of the present application, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the attached claims.

Claims

1. A method for updating an image recognition result correction model, characterized in that: The method comprises: Acquire a detection image and a preliminary recognition result of the detection image, wherein the preliminary recognition result includes a target in the detection image and a predicted category of the target; Correcting the predicted category by using an image recognition result correction model to obtain a corrected recognition result; Displaying the detected image and the corrected recognition result; Obtaining a confirmation result input by the user, wherein the confirmation result is used to indicate whether the predicted category in the corrected recognition result is correct; setting a label for the target in the detection image according to the confirmation result; The annotated detection image is set as training data to train the image recognition result correction model, and the parameters of the image recognition result correction model are updated.

2. The method according to claim 1, characterized in that The annotation includes a first annotation and a second annotation, and setting the annotation for the target in the detection image according to the confirmation result includes: If the confirmation result is that the predicted category in the corrected recognition result is correct, setting the first annotation for the target in the detection image; If the confirmation result is that the predicted category in the corrected recognition result is wrong, the second annotation is set for the target in the detection image.

3. The method according to claim 1, characterized in that The detected image includes a plurality of targets, the preliminary recognition result also includes a confidence level corresponding to the predicted category, and the method further includes: If the confirmation result input by the user for some of the multiple targets is not obtained, setting a pseudo mark for the part of the targets in the detection image according to the confidence corresponding to the predicted category of the part of the targets; The detected image after being set with the annotation and the pseudo-identification is used as training data to train the image recognition result correction model, and the parameters of the image recognition result correction model are updated.

4. The method according to claim 3, characterized in that The pseudo identifier includes a first pseudo identifier and a second pseudo identifier, and setting the pseudo identifier for the partial target in the detected image according to the confidence corresponding to the predicted category of the partial target includes: For each target in the stated section: If the confidence corresponding to the predicted category of the target in the preliminary recognition result is greater than or equal to a first confidence threshold, setting the first pseudo-identification for the target in the detection image; If the confidence corresponding to the predicted category of the target in the preliminary recognition result is less than or equal to a second confidence threshold, the second pseudo-identification is set for the target in the detection image, wherein the second confidence threshold is less than the first confidence threshold.

5. The method according to claim 1, characterized in that The preliminary recognition result further includes a target frame, and the target frame is used to indicate the position information of the target in the detection image. Before displaying the detection image and the corrected recognition result, the method further includes: The target frame is adjusted by the image recognition result correction model so that the adjusted target frame correctly represents the position information of the target in the detected image.

6. The method according to claim 1, characterized in that The preliminary recognition result also includes a target box and a confidence level corresponding to the predicted category, wherein the target box is used to indicate the position information of the target in the detected image; The displaying of the detected image and the corrected recognition result includes: The detected image is displayed and the target box and the predicted category are displayed according to the confidence level.

7. The method according to claim 1, characterized in that The preliminary recognition result also includes a confidence level corresponding to the predicted category, and the correcting the predicted category by using the image recognition result correction model includes: The image recognition result correction model corrects the predicted category according to the confidence level.

8. A device for updating an image recognition result correction model, comprising a memory, a processor, and a computer program stored in the memory, characterized in that: The processor executes the computer program to implement the method for updating the image recognition result correction model according to any one of claims 1 to 7.

9. A target recognition and display system, characterized in that: The system includes radar and visual equipment and an updating device for the image recognition result correction model as described in claim 8, wherein the radar and visual equipment are used to detect the control area to obtain a detection image, identify the detection image to obtain a preliminary recognition result, and transmit the detection image and the recognition result to the updating device for the image recognition result correction model, and the preliminary recognition result includes a target in the detection image and a predicted category of the target.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method for updating the image recognition result correction model described in any one of claims 1 to 7 is implemented.