Explanatable target detection method of end side equipment
By designing a lightweight object detection model on the end-side equipment and introducing a language model, the problem of high-precision detection and interpretation in the end-side equipment is solved, efficient and accurate object detection is achieved and detailed explanation is provided.
Patent Information
- Application Number
- CN202510374437.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-27
- Publication Date
- 2025-07-25
AI Technical Summary
The existing target detection methods provide reasonable explanations while the end-side device cannot provide high-precision detection results when computing resources and storage space are limited.
A lightweight object detection model is designed to combine non-maximum suppression algorithms and language big models to generate bounding boxes, category labels and confidence scores, and provide detailed explanations through language big models.
Implement high-precision object detection on resource-constrained devices, while providing semantic rich interpretation information to enhance user trust and system operability.
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of computer vision and deep learning, and more specifically, to an interpretable object detection method for edge devices. Background Art
[0002] With the rapid development of deep learning and computer vision technologies, object detection technology has achieved remarkable results in multiple fields. However, existing object detection methods are usually "black box" models, that is, they only give detection results and cannot explain why the model makes a certain decision. Interpretable object detection can not only enhance the credibility of the model but also improve users' understanding of the model results in practical applications. In edge devices, due to the limitations of computing resources and storage space, how to provide reasonable explanations while ensuring detection accuracy is still an important technical challenge. Summary of the Invention
[0003] Aiming at the deficiencies of the existing technology, the purpose of the present invention is to provide an interpretable object detection method for edge devices, which can not only detect objects in images but also give explanatory information for object detection under limited computing resources. Based on an advanced deep learning model and combined with the feature extraction and reasoning process of images, this method provides a lightweight and highly interpretable object detection solution.
[0004] To achieve the above purpose, the present invention provides the following technical solution: An interpretable object detection method for edge devices, including the following steps:
[0005] Step 1, design an object detection model and preprocess the input image. After the image is scaled and normalized, it is input into the object detection model. The object detection model identifies different objects in the image and generates a bounding box and a class label for each object to describe the position of the object in the image. At the same time, a confidence score is also attached to indicate the probability that the object is correctly identified;
[0006] Step 2, through the non-maximum suppression algorithm, remove duplicate detection boxes and retain the most accurate object detection results;
[0007] Among them, on the basis of object detection, a language large model is introduced to generate detailed explanations to convert the output of the object detection model into semantically rich and easy-to-understand explanations.
[0008] As a further improvement of the present invention patent, the specific steps of the object detection model in Step 1 for identifying different objects in the image and generating a bounding box and a class label for each object to describe the position of the object in the image, and at the same time attaching a confidence score are as follows:
[0009] Step 1-1: Divide the image into multiple grids through the object detection model, and each grid is responsible for predicting the category and bounding box of the objects within that area;
[0010] Step 1-2: For each bounding box divided in Step 1-1, the object detection model outputs a confidence score, which is usually a value between 0 and 1, indicating the probability that a certain object is contained within that bounding box;
[0011] Among them, a high confidence score indicates that the model is more certain that the position contains the object, while a low score indicates that the model is more skeptical that there is an object in that area. The confidence score is calculated by combining the weights and features learned by the model during training with the output information of the object category.
[0012] As a further improvement of this invention patent, the specific steps of removing duplicate detection boxes through the non-maximum suppression algorithm in Step 2 and retaining the most accurate object detection results are as follows:
[0013] Step 2-1: First, sort all candidate boxes according to the confidence scores of each bounding box, and the box with the highest score is ranked at the front;
[0014] Step 2-2: From the sorted candidate boxes, select the box with the highest confidence score as the current "best box";
[0015] Step 2-3: Calculate the degree of overlap between the current "best box" and other candidate boxes, represented by the intersection over union (IoU), which is the ratio of the intersection area to the union area of the two boxes;
[0016] Step 2-4: For other boxes whose overlap with the "best box" exceeds the set threshold, remove them from the candidate boxes;
[0017] Step 2-5: Repeat Steps 2-2 to 2-4 until there are no more candidate boxes to process;
[0018] Step 2-6: Retain the remaining boxes as the best detection results after processing.
[0019] As a further improvement of this invention patent, the specific steps of introducing a language large model to generate detailed explanations are as follows:
[0020] Step 3: Based on the boxes in the object detection model results, use a feature extraction network to extract the features of the screenshot, and also use the feature extraction network to extract the features of the entire image. Then fuse the two features into a set of features and input them into the language large model, so that the language large model can reason according to the context to generate a more detailed natural language description. At the same time, the language large model combines knowledge bases in different fields to provide domain-specific explanations for the object detection results.
[0021] As a further improvement of this invention patent, it also includes the step of training the object detection model in combination with a large language model, specifically as follows: Use image + text data to train the object detection model. The training method is to fix the object detection module and the large language model and only train the feature extraction network;
[0022] Among them, the text is based on the context information of the object.
[0023] The beneficial effects of this invention patent: Compared with the existing "black box" object detection model, it has the following advantages:
[0024] 1. Edge-side operation: The object detection model designed in the present invention runs on edge-side devices, reducing the dependence on network bandwidth and being suitable for low-latency and high-performance application scenarios.
[0025] 2. Interpretability: By integrating an interpretability mechanism, users can understand the reason for a certain detection result of the model, increasing trust in the model.
[0026] 3. Lightweight design: Adopting a lightweight network structure ensures the efficient operation of the model on devices with limited computing resources.
[0027] 4. High precision and low computational cost: Through the optimized feature extraction network and object detection head, the present invention can provide high-precision object detection results on edge-side devices while maintaining a low computational cost. Specific implementation mode
[0028] The following will further detail the present invention with the given embodiments.
[0029] An interpretable object detection method for an edge-side device in this embodiment is used to implement the detection of interpretable objects and is specifically composed of the following steps:
[0030] 1. Object detection model design and input preprocessing
[0031] · Image input: First, construct a basic object detection model. The object detection model receives an input image. After the image undergoes standard preprocessing steps (such as scaling and normalization), it is input into the object detection model. The model identifies different objects in the image and generates a bounding box and a class label for each object.
[0032] 2. Object detection result generation
[0033] · Object category and location: The object detection model generates a class label (such as "cat", "car") and a bounding box for each object, describing the location of the object in the image.
[0034] · Confidence Score: Each object detection result is also accompanied by a confidence score, which represents the probability that the object is correctly recognized. The generation process is as follows: (1) The object detection model usually divides the image into multiple grids, and each grid is responsible for predicting the object category and bounding box within that area. (2) For each bounding box, the model outputs a confidence score, which is usually a value between 0 and 1, representing the probability that there is an object within that bounding box. A high confidence score indicates that the model is more certain that the position contains an object, while a low score indicates that the model is more skeptical about the presence of an object in that area. (3) The confidence score is calculated by combining the weights and features learned by the model during training with the output information of the object category. This score takes into account not only whether the object category is correct but also whether the bounding box of the object is accurate.
[0035] · Candidate Box Filtering: Duplicate detection boxes are removed through the Non-Maximum Suppression (NMS) algorithm, and the most accurate object detection results are retained. The specific steps of the NMS algorithm are as follows: (1) Sorting: First, all candidate boxes are sorted according to the confidence scores of each bounding box, and the box with the highest score is ranked at the front. (2) Selecting the Best Box: From the sorted candidate boxes, the box with the highest confidence score is selected as the current "best box". (3) Calculating the Overlap Degree (IOU): Calculate the overlap degree between the current "best box" and other candidate boxes. The commonly used metric is the Intersection over Union (IoU), which represents the ratio of the intersection area to the union area of two boxes. (4) Suppressing Duplicate Boxes: For other boxes whose overlap degree (IoU) with the "best box" exceeds the set threshold (usually the threshold is set to 0.5), they are removed from the candidate boxes. (5) Repeating the Process: Repeat steps 2 to 4 until there are no more candidate boxes to process. (6) Retaining the Final Boxes: Finally, the remaining boxes are retained, and they are all the best detection results after NMS processing.
[0036] 3. Reasoning Combining with Large Language Model (LLM)
[0037] Based on object detection, we introduce a large language model to generate detailed explanations. The LLM has powerful natural language generation capabilities and can transform the output of the object detection model into semantically rich and easy-to-understand explanations.
[0038] · Format Input to the LLM: Based on the boxes of the od results, use the feature extraction network to extract the features of the screenshot, and then use the feature extraction network to extract the features of the entire image. The two features are fused into a set of features and input into the LLM
[0039] · Context information supplementation: The LLM not only receives the basic information of object detection but can also reason based on the context (such as the scene of the image, the presence of other objects, etc.) to generate a more detailed natural language description. For example, the LLM can identify the relationships between multiple objects in the image and give corresponding explanations, such as "The cat in the image is on the windowsill, near the car."
[0040] · Customized explanation: The LLM can combine knowledge bases in different fields to provide domain-specific explanations for object detection results. For example, in medical image processing, the LLM can provide medical explanations for tumor detection results, or in industrial monitoring, it can explain the detection of equipment failures.
[0041] 4. Training in combination with the Language Model (LLM)
[0042] (1) Data: Images + text, where the text is mainly based on the context information of the object
[0043] (2) Training method: Fix the od module and the LLM, and only train the feature extraction network.
[0044] In this embodiment, an application scenario of the method is provided. Assume that the method of the present invention is applied in an intelligent monitoring camera. The video stream captured by the camera is input into the processing unit of the device. The object detection model on the device will real-time detect objects such as pedestrians, vehicles, and animals in the image, and at the same time generate a heat map, showing the key areas of the model's decision-making. These information can not only help users understand the specific situation in the monitoring video but also provide effective feedback on the system's decision-making process, enhancing the operability and trust of the system.
[0045] In summary, the interpretable object detection method for edge devices provided by the present invention can achieve efficient and accurate object detection on resource-constrained devices through a lightweight deep learning model and an interpretability mechanism, and at the same time provide clear explanation information for each detection result. This technology has broad application prospects in the fields of intelligent security, autonomous driving, industrial detection, etc.
[0046] The above is only the preferred embodiment of the present invention. The protection scope of the present invention is not limited to the above embodiments. All technical solutions falling within the idea of the present invention belong to the protection scope of the present invention. It should be noted that for those of ordinary skill in the art, several improvements and refinements made without departing from the principle of the present invention should also be regarded as within the protection scope of the present invention.
Claims
1. An interpretable object detection method for edge devices, characterized in that: It includes the following steps: Step 1: Design a target detection model and preprocess the input image. After the image is scaled and normalized, it is input into the target detection model. The target detection model identifies different targets in the image and generates a bounding box and a class label for each target to describe the position of the object in the image. At the same time, it also comes with a confidence score to indicate the probability that the target is correctly identified; Step 2: Use the non-maximum suppression algorithm to remove duplicate detection boxes and retain the most accurate target detection results; Among them, based on target detection, a language large model is introduced to generate detailed explanations to convert the output of the target detection model into semantically rich and easy-to-understand explanations.
2. The interpretable object detection method for edge devices according to claim 1, wherein: The specific steps in Step 1 where the target detection model identifies different targets in the image and generates a bounding box and a class label for each target to describe the position of the object in the image, and also comes with a confidence score are as follows: Step 1-1: Divide the image into multiple grids through the target detection model, and each grid is responsible for predicting the class and bounding box of the object in that area; Step 1-2: For each bounding box divided in Step 1-1, the target detection model outputs a confidence score. Usually, this score is a value between 0 and 1, indicating the probability that there is a certain target within the bounding box; Among them, a high confidence score indicates that the model is more certain that the position contains a target, while a low score indicates that the model is more skeptical that there is a target in that area. The confidence score is calculated by the weights and features learned by the model during training, combined with the output information of the target class.
3. The interpretable object detection method for the edge device according to claim 1 or 2, characterized in that: The specific steps in Step 2 where the non-maximum suppression algorithm is used to remove duplicate detection boxes and retain the most accurate target detection results are as follows: Step 2-1: First, sort all candidate boxes according to the confidence score of each bounding box, and the box with the highest score is ranked at the front; Step 2-2: From the sorted candidate boxes, select the box with the highest confidence score as the current "best box"; Step 2-3: Calculate the overlap degree between the current "best box" and other candidate boxes, represented by the intersection over union (IoU), which is the ratio of the intersection area to the union area of the two boxes; Step 2-4: Remove other boxes whose overlap degree with the "best box" exceeds the set threshold from the candidate boxes; Step 2-5: Repeat Steps 2-2 to 2-4 until there are no more candidate boxes to process; Step 2-6: Retain the remaining boxes as the best detection results after processing.
4. The interpretable object detection method for the edge device according to claim 1 or 2, characterized in that: The specific steps for introducing the language large model to generate detailed explanations are as follows: Step 3: Based on the boxes in the results of the target detection model, use a feature extraction network to extract the features of the screenshot. Then use the feature extraction network to extract the features of the entire image, and fuse the two features into a set of features and input them into the language large model. In this way, the language large model can perform reasoning based on the context to generate a more detailed natural language description. At the same time, the language large model combines knowledge bases in different fields to provide domain-specific explanations for the target detection results.
5. The interpretable object detection method for the edge device according to claim 1 or 2, characterized in that: It further includes the step of training the object detection model in combination with a language large model, specifically as follows: training the object detection model with image + text data, and the training method is to fix the object detection module and the language large model and only train the feature extraction network; Among them, the text is based on the context information of the object.