Inspection robot image recognition method and inspection robot
By extracting key image features in the inspection robot and calculating confidence in combination with object detection algorithms and context information, the problem of insufficient recognition accuracy in traditional inspection robots in complex environments is solved, achieving higher recognition accuracy and robustness.
Patent Information
- Application Number
- CN202510142373.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-10
- Publication Date
- 2025-05-30
AI Technical Summary
Traditional inspection robots lack image recognition accuracy and robustness in complex and changing inspection environments, making it difficult to cope with changes in lighting conditions, occlusions and target diversity.
By collecting patrol images in the patrol robot, preprocessing is performed to extract key feature information, positioning the target to be detected using the target detection algorithm, and calculating the confidence of the recognition result in combination with the feature matching degree and context information, setting a confidence threshold to confirm or mark the target to be detected.
It improves the image recognition accuracy and robustness of patrol robots in complex environments, significantly reduces false alarms and missed alarms, and ensures a high recognition accuracy.
Smart Images

Figure CN120071283A_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of robot patrol inspection, and particularly relates to a patrol inspection robot image recognition method and a patrol inspection robot. Background Art
[0002] With the rapid development of industrial automation and intelligent technologies, patrol inspection robots have been widely used in various industries, commercial and public facilities. These robots can autonomously or remotely perform patrol inspection tasks by integrating cameras and other sensors, effectively replacing the traditional manual patrol inspection method and improving the patrol inspection efficiency and safety.
[0003] However, there are still many challenges in image recognition for traditional patrol inspection robots. On the one hand, due to the complex and changeable patrol inspection environment, such as factors like lighting conditions, occlusions, and target diversity, the quality of the collected patrol inspection images is uneven, bringing great difficulties to image recognition. On the other hand, traditional image recognition methods often rely on manually designed feature descriptors and classifiers, and it is difficult to handle complex and changeable patrol inspection scenarios, resulting in insufficient recognition accuracy and robustness. Summary of the Invention
[0004] The purpose of this application is to overcome the above technical problems in the prior art and provide a patrol inspection robot image recognition method and a patrol inspection robot.
[0005] This application provides a patrol inspection robot image recognition method, including:
[0006] The patrol inspection robot collects patrol inspection images;
[0007] After preprocessing the collected patrol inspection images, extract the key feature information in the images;
[0008] Based on the extracted key feature information, use a target detection algorithm to locate the target to be detected in the image, and calculate the position, size, and shape of the target to be detected;
[0009] Match the position, size, and shape of the target to be detected with a predefined target feature library, and calculate the feature matching degree;
[0010] Calculate the confidence level of the recognition result according to the feature matching degree combined with the context information in the image;
[0011] Set a confidence level threshold. When the confidence level is higher than the confidence level threshold, confirm the target to be detected; otherwise, mark it as a suspected target and conduct secondary verification or manual confirmation.
[0012] Optionally, extracting the key feature information in the image includes:
[0013] Use edge detection or contour extraction to obtain the key features in the image.
[0014] Optionally, the object detection algorithm includes: an algorithm based on a region proposal network and the Fast R-CNN or Faster R-CNN framework.
[0015] Optionally, the context information includes:
[0016] background information in the image, adjacent object information, or the relative positional relationship between objects.
[0017] Optionally, it further includes:
[0018] Updating and optimizing a predefined object feature library according to the results of manual confirmation to improve the accuracy and efficiency of subsequent recognition.
[0019] This application also provides an inspection robot, including:
[0020] an image acquisition device for acquiring images during the inspection process;
[0021] an image processing module for preprocessing the acquired images and extracting key feature information in the images; an object detection module for locating the object to be detected in the image using an object detection algorithm and calculating the position, size, and shape of the object;
[0022] a feature matching module for matching the position, size, and shape of the object to be detected with a predefined object feature library and calculating the feature matching degree;
[0023] a confidence calculation module for calculating the confidence of the recognition result based on the feature matching degree in combination with the context information in the image;
[0024] a decision-making module for setting a confidence threshold and determining whether the object to be detected is a real object based on the confidence, or marking it as a suspected object for further processing.
[0025] Optionally, the image processing module uses a convolutional neural network or a deep residual network to extract features from the image.
[0026] Optionally, the object detection module uses an algorithm based on a region proposal network and the Fast R-CNN or Faster R-CNN framework for object detection.
[0027] Optionally, the context information considered by the confidence calculation module includes background information in the image, adjacent object information, or the relative positional relationship between objects.
[0028] Optionally, it further includes a manual confirmation interface and a feature library update interface for receiving the results of manual confirmation and updating and optimizing the predefined object feature library according to the results.
[0029] The beneficial effects of this application are as follows:
[0030] Inventive point 1: Calculation of confidence
[0031] Inventive point 2: Selection of background information parameters when calculating confidence
[0032] Inventive point 3: Integrated application of coordinates, size parameters, and shape description
[0033] This application provides a method for image recognition of an inspection robot, including: the inspection robot collects inspection images; after preprocessing the collected inspection images, key feature information in the images is extracted; based on the extracted key feature information, a target detection algorithm is used to locate the target to be detected in the image, and the position, size, and shape of the target to be detected are calculated; the position, size, and shape of the target to be detected are matched with a predefined target feature library to calculate the feature matching degree; according to the feature matching degree and combined with the context information in the image, the confidence of the recognition result is calculated; a confidence threshold is set, and when the confidence is higher than the confidence threshold, the target to be detected is confirmed; otherwise, it is marked as a suspected target for secondary verification or manual confirmation. By calculating the feature matching degree and comprehensively considering other relevant information in the image, this application makes the recognition result more reliable, enhances the robustness of the recognition, and can maintain a high recognition accuracy even in a complex and changeable inspection environment. By setting a confidence threshold and making an intelligent decision based on the confidence of the recognition result, this technical solution can significantly reduce the situations of false alarms and missed detections. BRIEF DESCRIPTION OF THE DRAWINGS
[0034] Figure 1 is a schematic diagram of the image recognition process of the inspection robot in this application;
[0035] Figure 2 is a schematic diagram of the structure of the inspection robot in this application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0036] Hereinafter, exemplary embodiments of the present disclosure will be described in more detail with reference to the accompanying drawings. Although the exemplary embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be limited by the embodiments described herein. On the contrary, the described embodiments are provided so that this disclosure can be more thoroughly understood and the scope of this disclosure can be fully conveyed to those skilled in the art.
[0037] Please refer to Figure 1 as shown, this application provides a method for image recognition of an inspection robot, including:
[0038] S101. The inspection robot collects inspection images;
[0039] Before the inspection begins, the inspection robot is deployed in the designated inspection area. According to the requirements of the inspection task, the robot will perform path planning to ensure that all key inspection points can be comprehensively and efficiently covered.
[0040] The inspection robot is equipped with high-performance image acquisition devices, including high-definition cameras.
[0041] During the inspection, the robot will automatically trigger the image acquisition function according to the preset path and inspection points. When the robot reaches a certain inspection point, it will adjust parameters such as camera focus and exposure to ensure that the captured images are clear and accurate.
[0042] The captured inspection images will be stored in the robot's built-in storage device in real time. At the same time, the robot will also transmit the image data to the background monitoring center or cloud platform in real time through wireless network or wired connection.
[0043] S102. After preprocessing the captured inspection images, extract the key feature information in the images;
[0044] Image preprocessing is a necessary step before image analysis, aiming to improve image quality, reduce noise and interference, so as to enhance the useful information in the image. The preprocessing of inspection images includes the following links:
[0045] 1. Denoising processing: During the inspection, due to the influence of various factors such as the environment and equipment, the captured images will contain noise. The purpose of denoising processing is to remove this noise and make the image clearer. Commonly used denoising methods include mean filtering, median filtering, Gaussian filtering, etc.
[0046] 2. Image enhancement: For images with insufficient light or low contrast, image enhancement technology is used to improve their contrast, brightness and other attributes, so that the details in the image are clearer.
[0047] 3. Image transformation: In order to better extract the features in the image, it is necessary to transform the image, such as Fourier transform, wavelet transform, etc. These transformations convert the image from the spatial domain to the frequency domain, making it easier to analyze and process the information in the image.
[0048] After the image preprocessing is completed, extract the key feature information in the image. These feature information include shape, texture, and color.
[0049] Shape is an important feature in an image, which is extracted by methods such as edge detection and contour extraction. Edge detection finds the edge points in the image, while contour extraction further obtains the contour lines of the object. For example: Use an edge detection algorithm (such as Canny edge detection) to extract the edge information in the image, and use a contour extraction algorithm (such as the findContours function in OpenCV) to extract the contour of the vehicle. Parameters such as the perimeter, area, and rectangularity of the contour are used as shape features.
[0050] Texture is a repetitive pattern of pixel gray values or color values in a local area of the image. Through texture analysis, texture features in the image are extracted, such as the thickness and direction of the texture. For example: Use the Local Binary Pattern (LBP) algorithm to extract the texture information of the image. The LBP algorithm generates a binary pattern by comparing the gray value of each pixel with that of its neighboring pixels, and then statistically processes these patterns to obtain a texture histogram as the texture feature.
[0051] Color is also an important feature in the image. Through methods such as color space conversion and color histogram statistics, color features in the image are extracted. For example: Convert the image to the HSV color space and calculate the hue (H), saturation (S), and value (V) values of each pixel. Then, statistically process these values to obtain a color histogram as the color feature.
[0052] S103. Based on the extracted key feature information, use an object detection algorithm to locate the object to be detected in the image, and calculate the position, size, and shape of the object to be detected;
[0053] The object detection algorithm includes: an algorithm based on the Region Proposal Network and the Fast R-CNN or Faster R-CNN framework.
[0054] Use methods such as simple splicing or weighted summation to also fuse the color, shape, and texture into a composite feature vector. Then, use a machine learning algorithm (such as Support Vector Machine SVM) or a deep learning algorithm (such as Convolutional Neural Network CNN) to perform classification training on this composite feature vector.
[0055] Input the composite feature vector into the trained classifier for classification. If the classifier output indicates the existence of the target, determine the position of the target in the image according to the classifier output (such as bounding box coordinates, target category, etc.). Further, calculate the size of the target (such as width and height) according to the coordinates of the bounding box, and calculate the shape of the target (such as rectangularity, circularity, etc.) according to the contour information.
[0056] S104. Match the position, size, and shape of the object to be detected with a predefined object feature library, and calculate the feature matching degree;
[0057] Construct a feature library containing multiple known target categories. Each category has a set of predefined features, which are obtained by analyzing images of a large number of targets of the same kind. Each entry in the feature library includes the category label, location features, size features, and shape features of the target.
[0058] Compare the features of the target to be detected with each entry in the feature library one by one, and calculate the feature matching degree between the target to be detected and each target in the feature library.
[0059] The matching degree is a quantitative index used to measure the similarity between the target to be detected and the targets in the library. The higher the matching degree, the greater the similarity between the target to be detected and the targets in the library.
[0060] Assume that the feature vector of the target to be detected is F det =[f 1 ,f 2 ,…,f n , where f i represents the i-th feature (such as location, size, shape). The feature vector of a certain target in the feature library is F lib =[f′ 1 ,f′ 2 ,…,f′ n .
[0061] The feature matching degree M is measured by calculating the distance or similarity between two feature vectors. This application uses the reciprocal of the Euclidean distance as the similarity index.
[0062] Taking the reciprocal of the Euclidean distance as an example, the feature matching degree M is expressed as:
[0063]
[0064] where n is the number of dimensions of the feature vector.
[0065] In the formula, calculates the Euclidean distance between two feature vectors. The smaller the Euclidean distance, the closer the positions of the two feature vectors in the multi-dimensional space, that is, the higher the similarity.
[0066] To convert the distance into a similarity index, take the reciprocal of the Euclidean distance. In this way, when the distance is smaller (the similarity is higher), the matching degree M is larger.
[0067] n in the formula represents the number of dimensions of the feature vector, that is, how many features are considered for matching. The feature dimensions include position coordinates, size parameters, and shape descriptions.
[0068] S105. Calculate the confidence level of the recognition result based on the feature matching degree and the context information in the image;
[0069] The context information includes: background information in the image, adjacent target information, or relative positional relationship between targets.
[0070] The confidence level calculation formula is as follows:
[0071]
[0072] Among them, C is the confidence level, M is the feature matching degree, α i is the weight coefficient of the i-th target detection attribute, A pi is the accuracy score of the i-th target detection attribute, S ci is the consistency score of the i-th target detection attribute, β is the target detection score adjustment coefficient, D ci is the target detection inconsistency score, γ is the adjustment coefficient of the context information inconsistency weight, W ij is the weight of the j-th context information inconsistency factor, I dj is the score of the j-th context information inconsistency factor, δ is the context information deviation adjustment coefficient, ∈ k is the reliability coefficient of the k-th context information factor, C kk′ is the consistency score between the k-th context information factor and the k'-th expected context information factor.
[0073] Except for the matching degree, each parameter is divided into four categories:
[0074] The first category, parameters related to target detection attributes:
[0075] α i , used to adjust the influence of different attributes on the target detection score. Obtained through expert scoring or machine learning model training.
[0076] A pi , reflecting the detection accuracy of the target in the image for this attribute. Based on the output of the detection algorithm, calculate the accuracy in combination with the true label. The calculation formula is as follows:
[0077]
[0078] Among them, TP represents the true positive example, that is, the number of positive samples correctly detected. FP represents the false positive example, that is, the number of negative samples misdetected as positive samples
[0079] S ci , measuring the stability of the target for this attribute in multiple detections. Evaluate by detecting the same target multiple times and calculating the standard deviation or coefficient of variation of the attribute scores. The calculation formula is as follows:
[0080]
[0081] Among them, σ represents the standard deviation of the attribute scores in multiple detection results, and μ represents the average value of the attribute scores in multiple detection results.
[0082] The second category, the target detection inconsistency score:
[0083] D ci , which reflects the degree of inconsistency between attributes when the target is detected in the image. It is obtained by calculating the correlation or difference degree between different attribute scores. The calculation formula is as follows:
[0084]
[0085] Among them, |ρ| is the absolute value of the Pearson correlation coefficient, and the Pearson correlation coefficient measures the linear correlation degree between two variables.
[0086] The third category, the context information inconsistency related parameters:
[0087] W ij , which is used to adjust the influence of different factors on the confidence calculation. It is obtained through expert scoring and data-based importance evaluation.
[0088] I dj , which measures the degree of inconsistency between the context information related to the target in the image and the expected environment or state of the target. Based on the extraction and comparison of context information, the inconsistency score is calculated. The calculation formula is as follows:
[0089]
[0090] Among them, C actual represents the actually extracted context information, C expected represents the expected environment or state of the target, similarity(C actual , C expected ) represents the similarity between the two, and max similarity is the maximum value of the similarity.
[0091] ∈k, which reflects the reliability of the context information factor for the confidence calculation. It is evaluated by verifying the correlation degree between the context information and the target detection result. The calculation formula is as follows:
[0092]
[0093] Among them, C represents the context information factor, T represents the true label or verification data related to C, accuracy(C, T) represents the accuracy of C, and max accuracy is the maximum value of the accuracy.
[0094] C kk , measure the consistency between the actual context information and the expected context information. Compare the actually extracted context information with the expected context information and calculate the similarity or consistency score. The calculation formula is as follows:
[0095] C kk′ = similarity(C actual , C expected ')
[0096] where C actual represents the actually extracted context information, and C expected ' represents a form of expected context information, and C consistency = similarity(C actual , C expected ') represents the similarity between the two.
[0097] The fourth category, adjustment coefficients:
[0098] γ: The adjustment coefficient for the weight of context information inconsistency.
[0099] δ: The adjustment coefficient for context information deviation.
[0100] β: The adjustment coefficient for the target detection score.
[0101] S106. Set a confidence threshold. When the confidence is higher than the confidence threshold, confirm the target to be detected; otherwise, mark it as a suspected target and conduct secondary verification or manual confirmation.
[0102] During the process of target detection or recognition, assign a confidence score to each detected target. This score reflects the confidence level of the algorithm in the detection result, that is, the probability that the target is indeed the expected target.
[0103] First, set a confidence threshold. This threshold is a predefined value used to distinguish high-confidence detection results from low-confidence detection results. The selection of the confidence threshold is usually based on various factors, including the accuracy of the algorithm, false alarm rate, missed detection rate requirements, and the specific requirements of the application scenario, etc.
[0104] Evaluate the confidence score of each detection result according to the threshold.
[0105] If the confidence score of a certain detection result is higher than the set threshold, then it can be considered that this detection result is reliable, that is, confirm that the target is the target to be detected.
[0106] If the confidence score of a certain detection result is lower than the set threshold, then it can be considered that this detection result may not be accurate or reliable enough, so it is marked as a suspected target. For suspected targets, further secondary verification or manual confirmation is carried out to improve the accuracy of detection.
[0107] Detection results with high confidence can be directly confirmed, thus saving time and resources; while detection results with low confidence need to be further verified to ensure the accuracy of detection.
[0108] Furthermore, during the target recognition process, when the system cannot determine the identity or category of a certain target, it will be marked as a target to be confirmed. These targets to be confirmed will be submitted to human reviewers for confirmation. The reviewers judge the true identity or category of the target based on professional knowledge, experience or auxiliary tools (such as magnifying glasses, measuring tools, etc.). The results of manual confirmation will be recorded, including information such as the correct identity, category, and feature description of the target.
[0109] Update the predefined target feature library according to the results of manual confirmation.
[0110] If the target to be confirmed is identified as a new target type or category, its feature information will be added to the feature library. If the target to be confirmed is identified as a known target but the feature description is incorrect or incomplete, the corresponding entry in the feature library will be corrected or supplemented.
[0111] While updating the feature library, the features can also be optimized to improve the accuracy and efficiency of recognition.
[0112] Optimization methods include but are not limited to: extracting more refined features, using more advanced feature extraction algorithms, performing dimensionality reduction on features, etc. By optimizing the features, noise interference can be reduced, the discrimination between features can be improved, and thus the robustness of the recognition system can be enhanced.
[0113] As Figure 2 shown, the present application also provides a patrol robot, including:
[0114] An image acquisition device 201 for acquiring images during the patrol;
[0115] An image processing module 202 for preprocessing the acquired images and extracting key feature information in the images;
[0116] A target detection module 203 for locating the target to be detected in the image using a target detection algorithm and calculating the position, size and shape of the target;
[0117] A feature matching module 204 for matching the position, size and shape of the target to be detected with the predefined target feature library and calculating the feature matching degree;
[0118] The confidence calculation module 205 calculates the confidence of the recognition result according to the feature matching degree by combining the context information in the image;
[0119] The decision-making module 206 sets a confidence threshold, and determines whether the target to be detected is a real target according to the confidence, or marks it as a suspected target for further processing.
[0120] Furthermore, the image processing module extracts features from the image by using a convolutional neural network or a deep residual network.
[0121] Furthermore, the target detection module adopts an algorithm based on the region proposal network and the Fast R-CNN or Faster R-CNN framework for target detection.
[0122] Furthermore, the context information considered by the confidence calculation module includes background information, adjacent target information or relative position relationship between targets in the image.
[0123] Furthermore, it further includes a manual confirmation interface and a feature library update interface, which are used to receive the result of manual confirmation, and update and optimize the predefined target feature library according to this result.
Claims
1. A patrol robot image recognition method, characterized in that: include: The inspection robot collects inspection images; After preprocessing the collected inspection images, key feature information in the images is extracted; Based on the extracted key feature information, a target detection algorithm is used to locate the target to be detected in the image, and the position, size and shape of the target to be detected are calculated; Matching the position, size and shape of the target to be detected with a predefined target feature library, and calculating the feature matching degree; Calculating the confidence of the recognition result according to the feature matching degree combined with the context information in the image; Setting a confidence threshold, and when the confidence is higher than the confidence threshold, confirming the target to be detected; Otherwise, it is marked as a suspected target for secondary verification or manual confirmation.
2. The inspection robot image recognition method according to claim 1, characterized in that: Extract key feature information from the image, including: Use edge detection or contour extraction to obtain key features in the image.
3. The inspection robot image recognition method according to claim 1, characterized in that: The target detection algorithm includes: an algorithm based on a region proposal network and a Fast R-CNN or Faster R-CNN framework.
4. The inspection robot image recognition method according to claim 1, characterized in that: The context information includes: Background information in the image, information about adjacent targets, or the relative position relationship between targets.
5. The inspection robot image recognition method according to claim 1, characterized in that: Also includes: The predefined target feature library is updated and optimized based on the results of manual confirmation to improve the accuracy and efficiency of subsequent identification.
6. A patrol robot, characterized in that: include: An image acquisition device, used to acquire images during the inspection process; The image processing module pre-processes the collected images and extracts key feature information from the images; The target detection module uses the target detection algorithm to locate the target to be detected in the image and calculate the position, size and shape of the target; The feature matching module matches the position, size and shape of the target to be detected with the predefined target feature library and calculates the feature matching degree; The confidence calculation module combines the context information in the image and calculates the confidence of the recognition result according to the feature matching degree; The decision module sets the confidence threshold and determines whether the target to be detected is a real target based on the confidence, or marks it as a suspected target for further processing.
7. The inspection robot according to claim 6, characterized in that: The image processing module uses a convolutional neural network or a deep residual network to extract features from the image.
8. The inspection robot according to claim 6, characterized in that: The target detection module uses an algorithm based on a region proposal network and a Fast R-CNN or Faster R-CNN framework to perform target detection.
9. The inspection robot according to claim 6, characterized in that: The context information considered by the confidence calculation module includes background information in the image, adjacent target information or relative position relationship between targets.
10. The inspection robot according to claim 6, characterized in that: It also includes a manual confirmation interface and a feature library update interface, which are used to receive the results of manual confirmation and update and optimize the predefined target feature library based on the results.
Citation Information
Patent Citations
Method for property protection based on video image local feature matching
CN104298988A
Vein recognition method based on statistical information
CN114782715A
Equipment inspection method and device, electronic equipment and computer program product
CN116503807A
Cited By
Robot inspection method and system for container data room
CN121121676A
Abnormality detection method and device applied to inspection robot, equipment and medium
CN121708535A