Power equipment non-contact distance measurement method based on rotating target detection and infrared image

Through the non-contact ranging method of power equipment based on rotation target detection and infrared image, the problems of insufficient accuracy and poor environmental adaptability in power equipment identification and ranging technology are solved, and high-precision and low-cost power equipment identification and ranging are realized, which is suitable for intelligent inspection under complex working conditions.

CN120339732AActive Publication Date: 2025-07-18ZHUHAI JINRUI ELECTRIC POWER TECH CO LTD

Patent Information

Application Number
CN202510813626.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-18
Publication Date
2025-07-18
Estimated Expiration
2045-06-18

AI Technical Summary

Technical Problem

The existing power equipment identification and ranging technology has problems such as low recognition accuracy, poor adaptability to tilted targets, dependence on external sensors for ranging, and insufficient environmental adaptability. It is especially difficult to achieve efficient and low-cost intelligent identification and ranging under complex working conditions.

Method used

The non-contact ranging method of power equipment based on rotation target detection and infrared images is adopted. By collecting infrared image data sets of power equipment, a YOLOv5 object detection model is built, rotating target detection is optimized, spatial distance is calculated using rotation bounding box information, and the distance measurement results are corrected in combination with dynamic update algorithms to achieve high-precision identification and stable distance measurement of power equipment.

Benefits of technology

High-precision identification and stable distance measurement of power equipment are achieved in complex environments. It has no dependence on sensors, supports target attitude adaptation, and high distance measurement accuracy. It is suitable for all-weather operations and long-distance monitoring, reducing system complexity and cost.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120339732A_ABST
    Figure CN120339732A_ABST
Patent Text Reader

Abstract

The invention provides a power equipment non-contact distance measurement method based on rotating target detection and an infrared image, and the method comprises the steps: collecting the infrared image of power equipment, and generating an infrared image data set containing the target information of the power equipment; a YOLOv5 target detection model is established, and the model optimizes rotation target detection and is used for detecting rotation bounding box information of a power equipment target in the infrared image; performing target detection on the input infrared image by using the trained YOLOv5 target detection model, and outputting a non-contact distance measurement result; and based on a historical target detection result, correcting the current target position information and the distance measurement result by adopting a dynamic updating algorithm. According to the invention, high-precision identification and stable distance measurement of electrical equipment such as a terminal and a lightning arrester can be realized, and the method has the advantages of being independent of a sensor, supporting target attitude adaptation, being high in distance measurement precision and the like.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of power equipment image recognition and ranging, and particularly relates to a non-contact ranging method for power equipment based on rotating target detection and infrared images, which is applicable to scenarios such as power line inspection, substation equipment status monitoring, remote identification and positioning of high-voltage facilities, etc., and is especially suitable for efficiently identifying and accurately ranging power components such as cable terminals and lightning arresters in complex environments, at long distances, or in areas that are difficult to access manually. Background Art

[0002] With the continuous improvement of the intelligent level of the power system, the demand for automatic inspection and intelligent recognition of power equipment is increasing day by day. Especially in scenarios such as power transmission and transformation lines and substations, the status recognition and spatial positioning of key components such as cable terminals and lightning arresters are of great significance to the safe operation of equipment. At present, the recognition and ranging of power equipment mostly rely on manual inspection or the use of hardware devices such as lidar and ultrasonic sensors for auxiliary detection, and there are the following deficiencies: The recognition method relies on manual or traditional algorithms, with low efficiency and strong subjectivity: Traditional image recognition methods have low recognition accuracy when facing complex backgrounds, light changes, or equipment occlusion, and do not have self-adaptive capabilities; Ranging relies on dedicated hardware devices, with complex systems and high costs: Common ranging methods such as laser ranging and ultrasonic ranging require additional installation of dedicated sensors, which not only increases the cost but also improves the maintenance complexity; The target detection framework is insensitive to angle changes: Most of the currently widely used target detection models adopt horizontal bounding boxes (HBox), and when facing situations such as equipment posture tilt and perspective deviation, the detection accuracy significantly decreases, and false negatives or false positives are likely to occur; The ranging accuracy is greatly affected by target detection errors and lacks a compensation mechanism: Most existing image ranging methods calculate based on a single dimension (such as width or height), without considering the imaging differences caused by target angle changes, which is likely to lead to the accumulation of ranging errors.

[0003] In summary, the existing technologies still have problems such as insufficient accuracy, poor environmental adaptability, and high system costs in the recognition and ranging of power equipment, and there is an urgent need for a more efficient, low-cost, and intelligent recognition and ranging solution suitable for complex working conditions. Summary of the Invention

[0004] Aiming at the problems existing in the existing power equipment identification and ranging technologies, such as low identification accuracy, poor adaptability to inclined targets, ranging relying on external sensors, and insufficient environmental adaptability, the purpose of the present invention is to provide a non-contact ranging method for power equipment based on rotating target detection and infrared images applicable to intelligent inspection and equipment positioning applications in power systems under complex environments. This method can achieve high-precision identification and stable ranging of power equipment such as terminal heads and lightning arresters, and has the advantages of not relying on sensors, supporting target attitude adaptation, and high ranging accuracy.

[0005] The present invention realizes the above purpose through the following technical solutions: A non-contact ranging method for power equipment based on rotating target detection and infrared images, comprising the following steps: Collect infrared images of power equipment to generate an infrared image dataset containing power equipment target information; Build a YOLOv5 target detection model, which is optimized for rotating target detection and used to detect the rotating bounding box information of power equipment targets in infrared images; among them, the loss function of the YOLOv5 target detection model L includes classification loss L cls , bounding box regression loss L reg and rotation angle prediction loss L θ ; Use the trained YOLOv5 target detection model to perform target detection on the input infrared image to obtain the rotating bounding box information of the power equipment target B =( x c , y c , w , h , θ ), determine the coordinate of the feature point of the target on the image plane based on the rotating bounding box information, and calculate the spatial distance according to the distance relationship between the target feature point and the camera origin in the world coordinate system D , and output the non-contact ranging result; where ([[]] x c , y c ) is the center coordinate of the bounding box, w is the width of the bounding box, h is the height of the bounding box, θ is the rotation angle of the bounding box; Based on the historical target detection results, use a dynamic update algorithm to correct the current target position information and ranging results; among them, the dynamic update algorithm comprehensively considers the target detection results of the current frame and the previous nThe historical detection results of the frame are used to smooth the target position and ranging value through a weighted average algorithm.

[0006] According to a non-contact ranging method for power equipment based on rotating target detection and infrared images provided by the present invention, in the process of making an infrared image data set and preprocessing the data set, the following steps are included: Design cable terminal acquisition tasks for various actual usage scenarios, where the actual usage scenarios at least include scenarios under different environmental temperatures, different lighting conditions, different weather conditions, and different operating states of the cable terminal; Uniformly format all the obtained infrared images, and classify and store the formatted infrared images according to the time stamp, device number, and scenario number; Use a rotating rectangle to label the targets in the infrared images, so that the rotating rectangle fits the outer contour of the target. The labeling information includes the four vertex coordinates of the rotating target box in the image coordinate system and the rotation angle θ; at the same time, label the category information for each rotating target box, and the category information includes cable terminals and lightning arresters; Split all the labeled samples into a training set, a validation set, and a test set according to a preset ratio; among them, introduce a variety of data augmentation operations on the training set, at least including angle rotation, contrast change, mirror flipping, and cropping methods.

[0007] According to a non-contact ranging method for power equipment based on rotating target detection and infrared images provided by the present invention, determining the characteristic point coordinates of the target on the image plane based on the rotating bounding box information includes: For the rotating bounding box B =( x c , y c , w , h , θ ), select the four vertices of the rotating bounding box as the characteristic points of the target on the image plane. Let the rotating bounding box be centered at the center point ( x c , y c ), with the width direction as the x axis direction and the height direction as the y axis direction to establish a local coordinate system. In the local coordinate system, the coordinates of the four vertices are respectively P 1(- w / 2,- h / 2)、 P 2( w / 2,- h / 2)、 P 3( w / 2, h / 2), and P 4(- w / 2, h / 2); According to the rotation angle θ perform a rotation transformation on the vertex coordinates in the local coordinate system, using the rotation matrix R θ perform a rotation transformation on the vertex coordinates in the local coordinate system, and the rotation matrix R θ has the following expression:

[0008] For the vertex coordinates ( x local , y local ) in the local coordinate system, the rotated coordinates ( x rot , y rot ) are calculated as follows:

[0009] Convert the local coordinates of the four vertices P 1, P 2, P 3, P 4 to the rotated coordinates in the local coordinate system P 1rot ( x 1rot , y 1rot ), P 2rot ( x 2rot , y 2rot ), P 3rot ( x 3rot , y 3rot ), P 4rot ( x 4rot , y 4rot ); Convert the vertex coordinates in the rotated local coordinate system to the global coordinate system of the image plane, and obtain the coordinates ( x i , y i ) of the four feature points on the image plane ( i= 1, 2, 3, 4), the conversion formula is: x i = x rot + x c 、 y i = y rot + y c ; Based on P 1rot 、 P 2rot 、 P 3rot 、 P 4rot the coordinates of four feature points on the image plane are obtained based on P 1img ( x 1, y 1), P 2img ( x 2, y 2), P 3img ( x 3, y 3), P 4img ( x 4, y 4).

[0010] According to a non-contact ranging method for power equipment based on rotating target detection and infrared images provided by the present invention, calculating the spatial distance according to the distance relationship between the target feature points and the camera origin in the world coordinate system D, includes: Setting the internal parameter matrix of the infrared camera K as: where f x and f y are the focal lengths of the camera in the x axis and y axis directions respectively, c x and c y are the coordinates of the image principal point in the x axis and y axis directions respectively; For the coordinates of the feature points on the image plane ([[]] x i , yi ), perform normalization to obtain the normalized coordinates ([ x x inorm , y inorm ): [[ x inorm = ([ x x i - [[ c c x ) / [ f f x y inorm = ([ y y i - [[ c c y ) [ / f / f y At this time, assuming that [ Z Z ci = 1, [[ Z Z ci is the [[ z z - axis coordinate of the feature point in the camera coordinate system, then the normalized coordinates ([ x x inorm , [[ y y inorm , 1) approximately represent the coordinates in the camera coordinate system; [[ For [[ n n feature points, let the coordinates of the feature points in the camera coordinate system be ([ X X ci , [[ Y Y ci , [[ Z Z ci ), satisfying the following relationship: [[

[0011] Then there is: [[ x x i = [[ f f x u i + [[ c c x ; [[ y y i = [[ f f y v i + [[ c c y ; [[ u u i = ([ x x i - [[ c c x) / f x ; v i = ( y i - c y ) / f y ; Establish a system of equations through the coordinates of multiple feature points and solve it using the least squares method Z ci , Let the equation of the i th feature point be:

[0012] Combine the equations of the n th feature point and write them in matrix form Ax = b , where A is the coefficient matrix, x is the unknown vector, and this vector contains u i , v i and Z ci related information, b is the constant vector, and solve x through the least squares method to obtain Z ci value, and then obtain the coordinates of the feature points in the camera coordinate system ([[]] X ci , Y ci , Z ci ).

[0013] According to a non-contact ranging method for power equipment based on rotational target detection and infrared images provided by the present invention, set the extrinsic parameter matrix of the infrared camera R | T , where R is a 3×3 rotation matrix, T is a 3×1 translation vector; The coordinates of the feature points in the camera coordinate system ([[]] X ci , Y ci , Z ci ) and the coordinates of the feature points in the world coordinate system ([[]] X wi , Y wi , Z wi) The conversion relationship is as follows:

[0014] Convert the coordinates of the feature points in the camera coordinate system to the coordinates of the feature points in the world coordinate system; Calculate the distance between the target feature point in the world coordinate system and the camera origin D i , for the i th feature point, the distance D i The calculation formula is:

[0015] For the distances calculated for the four feature points D i ( i = 1, 2, 3, 4) take the average value to obtain the final spatial distance D , the calculation formula is:

[0016] Output the calculated spatial distance D as the non-contact ranging result.

[0017] According to a non-contact ranging method for power equipment based on rotating target detection and infrared images provided by the present invention, the loss function of the YOLOv5 target detection model L is expressed as the following formula: L = λ 1 L cls + λ 2 L reg + λ 3 L θ where λ 1, λ 2, λ 3 are weight coefficients; during the training process, use the validation set to evaluate the network model during training, and judge the fitting situation of the network according to the evaluation index on the validation set.

[0018] A non-contact ranging method for power equipment based on rotated object detection and infrared images provided by the present invention, wherein the YOLOv5 object detection model includes an input layer, a backbone network, a neck network, and a detection head. The backbone network adopts the CSPDarknet architecture and is used to extract multi-scale image features from infrared images. An SPPF module is embedded between the backbone network and the detection head. The SPPF module is composed of multiple max-pooling layers with different scales in parallel. By performing pooling operations on the input feature map at different scales and splicing the pooling results. The PANet structure is introduced into the neck network. The PANet structure fuses and transfers feature maps at different levels through top-down and bottom-up bidirectional feature fusion paths. Among them, the detection head predicts the target bounding boxes for feature maps of different scales. Specifically, it is divided into three detection heads with different scales, which are respectively applicable to the recognition scenarios of large, medium, and small-sized targets. Each detection head processes the input feature map through convolutional operations, and the output includes at least the class probability of the target, the center coordinates of the rotated bounding box ( x c , y c ), width w , height h , and rotation angle θ information, so as to achieve the positioning and classification of the target.

[0019] A non-contact ranging method for power equipment based on rotated object detection and infrared images provided by the present invention. In each round of training, samples are randomly drawn from the training set in batches and input into the network. The samples first undergo preprocessing through the input layer of the network, and then sequentially pass through the backbone network, the SPPF module, the neck network, and the detection head to obtain the output results of the network encoder. Compare the output results of the network encoder with the corresponding ground truth labels, calculate the loss value through the designed loss function L , and use the backpropagation algorithm to update and optimize the network parameters according to the loss value, so that the network continuously approaches the optimal solution. When the last batch of the training set is drawn and the training is completed, all batches in the validation set are sequentially input into the network for validation. During the validation process, the network predicts the validation set samples to obtain the prediction results. Compare the prediction results with the ground truth labels of the validation set, calculate a series of evaluation metrics to evaluate the performance of the current model on unseen samples, and determine whether there is overfitting or underfitting in the network. Repeat the training and validation steps until all samples in the training set complete the iterative update of the set cycle, complete the entire model training process, and obtain the trained YOLOv5 object detection model.

[0020] According to a non-contact ranging method for power equipment based on rotating object detection and infrared images provided by the present invention, a target buffer is established. After each object detection is completed and the ranging result is calculated, the object category, position information, ranging result, and the current frame number are recorded in the target buffer. For each object in the newly detected frame, calculate the intersection over union (IoU) value between its rotating bounding box and all rotating bounding boxes of the same category objects in the target buffer. Set an IoU matching threshold. When the IoU value of an object in the newly detected frame and a certain object of the same category in the target buffer is greater than or equal to the matching threshold, it is determined that these two objects are the same object; when the IoU value is less than the matching threshold or the category information is inconsistent, it is determined that the object in the newly detected frame is a new object. For the case where the objects are determined to be the same object, use the weighted average algorithm to update the ranging result.

[0021] According to a non-contact ranging method for power equipment based on rotating object detection and infrared images provided by the present invention, the calculation of the intersection over union (IoU) value of the rotating bounding box includes: First, according to the parameters of the two rotating bounding boxes, calculate their intersection area; among them, convert the rotating bounding box into a polygon, and then calculate the area of the intersection polygon through the polygon intersection algorithm. Calculate the areas of the two rotating bounding boxes respectively. Divide the area of the intersection polygon by the sum of the areas of the two rotating bounding boxes minus the area of the intersection polygon to obtain the IoU value, that is IoU = S intersection / S box1 + S box2 - S intersection , where S intersection is the area of the intersection polygon, S box1 and S box2 are the areas of the two rotating bounding boxes respectively.

[0022] It can be seen that, compared with the prior art, the method provided by the present invention has the following beneficial effects: 1. The present invention identifies and ranges power equipment based on infrared images. Utilizing the characteristics that infrared imaging is not restricted by light and can stably image at night or in bad weather, it can stably detect targets such as terminal heads and lightning arresters in various complex environments. Compared with traditional identification systems relying on visible light images, the present invention has stronger applicability and stability in scenarios such as all-weather operation and long-distance monitoring.

[0023] 2. The present invention adopts a rotation target detection method based on YOLOv5, which optimizes the rotation target detection and can accurately detect the rotation bounding box information of power equipment targets in infrared images. In the actual power equipment detection scenario, the poses of power equipment in images are often arbitrary. The traditional rectangular bounding box detection method is difficult to accurately describe the shape and position of the target, while the rotation bounding box can better fit the actual contour of the target, greatly improving the accuracy of target detection. For example, for an inclined cable terminal, the rotation bounding box can more accurately select the target, avoiding the situation that the traditional rectangular bounding box may contain too much background information or cannot completely cover the target, and providing more reliable target position information for subsequent ranging calculations; 3. The loss function L of the YOLOv5 target detection model of the present invention includes the classification loss L cls , the bounding box regression loss L reg and the rotation angle prediction loss L θ . This way of collaborative optimization of multiple loss functions enables the model to simultaneously focus on the class classification of the target, the regression of the position and size of the bounding box, and the prediction of the rotation angle. By jointly optimizing these loss functions, the model can more comprehensively learn the features of power equipment targets, further improving the accuracy and robustness of target detection. In practical applications, even if the power equipment is affected by factors such as light changes and occlusion, the model can accurately detect the target and obtain its rotation bounding box information; 4. The present invention uses the obtained rotation bounding box information of the power equipment target to determine the coordinate of the feature point of the target on the image plane based on the rotation bounding box information, and calculates the spatial distance according to the distance relationship between the target feature point and the camera origin in the world coordinate system, and outputs a non-contact ranging result. This ranging method based on image geometric features does not require additional sensors or markers to be installed on the power equipment, realizing true non-contact ranging and avoiding the potential safety hazards and equipment damage risks brought by traditional contact ranging methods. At the same time, this method makes full use of the geometric information in the image and obtains the spatial distance of the target through precise mathematical calculations, with high ranging accuracy.

[0024] 5. Based on the historical object detection results, the present invention uses a dynamic update algorithm to correct the current object position information and ranging results. This dynamic update algorithm comprehensively considers the object detection results of the current frame and the historical detection results of the previous n frames, and smooths the object position and ranging values through a weighted average algorithm. In practical applications, due to factors such as image noise and temporary occlusion of the object, there may be certain errors in the object detection results of a single frame. By comprehensively considering the historical detection results, the dynamic update algorithm can effectively reduce the influence of these errors and improve the stability and accuracy of the object position and ranging results. In addition, for the moving power equipment object, the dynamic update algorithm can real-time track the position and ranging changes of the object, and perform smoothing processing according to historical data, making the ranging results more in line with the actual movement trajectory of the object. For example, during the inspection of power equipment, the object may change its position and posture as the inspection equipment moves, and the dynamic update algorithm can timely adjust the object position and ranging results, providing more accurate data support for subsequent analysis and decision-making.

[0025] The following further elaborates on the present invention in detail in conjunction with the accompanying drawings and specific embodiments. Description of the Drawings

[0026] Figure 1 is a flowchart of an embodiment of a non-contact ranging method for power equipment based on rotational object detection and infrared images of the present invention.

[0027] Figure 2 is a flowchart of the processing of the infrared image dataset in an embodiment of a non-contact ranging method for power equipment based on rotational object detection and infrared images of the present invention.

[0028] Figure 3 is a flowchart of determining the coordinates of the feature points of the object on the image plane in an embodiment of a non-contact ranging method for power equipment based on rotational object detection and infrared images of the present invention.

[0029] Figure 4 is about calculating the spatial distance in an embodiment of a non-contact ranging method for power equipment based on rotational object detection and infrared images of the present invention. D The first flowchart.

[0030] Figure 5 is about calculating the spatial distance in an embodiment of a non-contact ranging method for power equipment based on rotational object detection and infrared images of the present invention. D The second flowchart.

[0031] Figure 6 is a schematic diagram of the first prediction result of the YOLOv5 object detection model in an embodiment of a non-contact ranging method for power equipment based on rotational object detection and infrared images of the present invention.

[0032] Figure 7 It is a schematic diagram of the second prediction result of the YOLOv5 target detection model in an embodiment of a non-contact ranging method for power equipment based on rotated target detection and infrared images according to the present invention.

[0033] Figure 8 It is a schematic diagram of the process of the dynamic update mechanism in an embodiment of a non-contact ranging method for power equipment based on rotated target detection and infrared images according to the present invention.

[0034] Figure 9 It is a schematic diagram of the result of multi-target ranging in an embodiment of a non-contact ranging method for power equipment based on rotated target detection and infrared images according to the present invention. Detailed implementation manners

[0035] To make the objectives, technical solutions and advantages of the present invention clearer, the technical solutions in the present invention will be clearly and completely described below with reference to the accompanying drawings in the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art without making creative efforts based on the embodiments in the present invention belong to the scope of protection of the present invention.

[0036] Reference to "embodiment" herein means that a particular feature, structure, or characteristic described in connection with the embodiment can be included in at least one embodiment of the present application. The phrase appears in various places in the specification and does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. Those skilled in the art will explicitly and implicitly understand that the embodiments described herein can be combined with other embodiments.

[0037] Refer to Figure 1 , this embodiment provides a non-contact ranging method for power equipment based on rotated target detection and infrared images, and the method includes the following steps: Step S1, collect infrared images of power equipment to generate an infrared image dataset containing target information of the power equipment; Step S2, build a YOLOv5 target detection model, which is optimized for rotated target detection and is used to detect the rotated bounding box information of the power equipment target in the infrared image; wherein, the loss function of the YOLOv5 target detection model L includes classification loss L cls , bounding box regression loss L reg and rotation angle prediction loss L θ; During the training process, use the validation set to evaluate the network model being trained. Based on the evaluation metrics on the validation set, such as average precision (AP), recall, etc., judge the fitting situation of the network. If underfitting or overfitting occurs, adjust the model structure or training parameters in a timely manner.

[0038] Step S3: Use the trained YOLOv5 object detection model to perform object detection on the input infrared image, and obtain the rotated bounding box information of the power equipment target. B =( x c , y c , w , h , θ ), Based on the rotated bounding box information, determine the coordinates of the feature points of the target on the image plane, and calculate the spatial distance according to the distance relationship between the target feature points and the camera origin in the world coordinate system. D , Output the non-contact ranging result; where ([[]] x c , y c ) is the center coordinate of the bounding box, w is the width of the bounding box, h is the height of the bounding box, θ is the rotation angle of the bounding box; Step S4: Based on the historical object detection results, use a dynamic update algorithm to correct the current object position information and ranging results; among them, the dynamic update algorithm comprehensively considers the object detection results of the current frame and the previous n frame of historical detection results, and smooths the object position and ranging values through a weighted average algorithm.

[0039] In the above step S1, during the process of making the infrared image dataset and preprocessing the dataset, the following steps are included: Use OpenCV image processing technology to clean, preprocess, and extract features from the dataset, so as to organize the dataset into a format suitable for model training; combine Python language for programming experiments and write Python code to implement the division of the dataset.

[0040] In practical applications, such as Figure 2As shown in the figure, first, cable terminal acquisition tasks for various actual usage scenarios are designed, covering typical scenarios such as substations, ground piles, platforms, and towers. The actual usage scenarios include at least scenarios under different environmental temperatures (such as high temperature, low temperature, normal temperature), different lighting conditions (such as strong light, weak light, backlight), different weather conditions (such as sunny, cloudy, rainy), and different operating states of the cable terminal (such as normal operation, minor fault, severe fault). During the data acquisition process, multiple cable terminals in the same scenario are repeatedly photographed to ensure diverse images are obtained under different times, angles, distances, and lighting conditions, so as to improve the comprehensiveness and representativeness of the dataset.

[0041] All the acquired infrared images are uniformly formatted and classified and stored according to the acquisition timestamp, device number, and scenario number. This helps to achieve rapid retrieval and effective management of the images, providing an orderly data foundation for subsequent stages such as data annotation and model training. Among them, the timestamp is used as the first-level directory and sorted and stored in the order of shooting time; under each timestamp directory, the device number is used as the second-level directory to distinguish the images corresponding to different cable terminal devices; under each device number directory, the scenario number is used as the third-level directory to further classify and store the images under different actual usage scenarios, facilitating the management and retrieval of subsequent data.

[0042] To improve the image quality and enhance the recognizability of the cable terminal in the image, the collected original infrared images are processed using an image enhancement algorithm.

[0043] Considering the angle and direction changes of the cable terminal in the image, a rotated rectangular box is used to label the target in the infrared image. For example, the Rolabelimg annotation tool is used to perform rotated target box annotation on the cable terminal or lightning arrester in the infrared image, so that the rotated rectangular box fits the outer contour of the target. The annotation information includes the four vertex coordinates of the rotated target box in the image coordinate system and the rotation angle θ; at the same time, the category information of each rotated target box is labeled, and the category information includes the cable terminal and the lightning arrester; All the labeled samples are split into a training set, a validation set, and a test set according to a preset ratio; among them, the training set is used for model training, the validation set is used to judge the fitting situation of the network model during the training process, and the test set is used to finally evaluate the model performance.

[0044] Furthermore, a variety of data augmentation operations are introduced on the training set, including at least angle rotation, contrast change, mirror flipping, and cropping methods.

[0045] The data - augmented dataset is divided into a training set, a validation set, and a test set in a ratio of 8:1:1, obtaining a total of 42,655 training samples, 3,416 validation samples, and 344 test samples.

[0046] In the above - mentioned step S3, the characteristic point coordinates of the target on the image plane are determined based on the rotated bounding box information, as Figure 3 shown, including: For the rotated bounding box B =( x c , y c , w , h , θ ), select the four vertices of the rotated bounding box as the characteristic points of the target on the image plane. Let the rotated bounding box be centered at the center point ( x c , y c ), with the width direction as the x axis direction and the height direction as the y axis direction to establish a local coordinate system. In the local coordinate system, the coordinates of the four vertices are respectively P 1(- w / 2,- h / 2)、 P 2( w / 2,- h / 2)、 P 3( w / 2, h / 2) and P 4(- w / 2, h / 2); According to the rotation angle θ perform a rotation transformation on the vertex coordinates in the local coordinate system. Use the rotation matrix R θ to perform a rotation transformation on the vertex coordinates in the local coordinate system. The expression of the rotation matrix R θ is:

[0047] For the vertex coordinates ( x local , y local ) in the local coordinate system, the rotated coordinates ( x rot , y rot ) calculation formula is:

[0048] Convert the local coordinates of the four vertices P 1、 P 2、 P 3、 P 4 to the coordinates after rotation in the local coordinate system P 1rot ( x 1rot , y 1rot )、 P 2rot ( x 2rot , y 2rot )、 P 3rot ( x 3rot , y 3rot )、 P 4rot ( x 4rot , y 4rot ); Convert the vertex coordinates in the rotated local coordinate system to the global coordinate system of the image plane to obtain the coordinates of the four feature points on the image plane( x i , y i )( i =1,2,3,4), and the conversion formula is: x i = x rot + x c 、 y i = y rot + y c ; Based on P 1rot 、 P 2rot 、 P 3rot 、 P 4rot to obtain the coordinates of the four feature points on the image plane P 1img ( x 1, y 1)、 P 2img ( x 2, y 2)、P 3img ( x 3, y 3), P 4img ( x 4, y 4).

[0049] In the above step S3, as Figure 4 shown, calculate the spatial distance according to the distance relationship between the target feature point and the camera origin in the world coordinate system D, including: Set the internal parameter matrix of the infrared camera K as: where f x and f y are the focal lengths of the camera in the x axis and y axis directions respectively, c x and c y are the coordinates of the image principal point in the x axis and y axis directions respectively; For the feature point coordinates ([[]] x i , y i ) on the image plane, perform normalization processing to obtain the normalized coordinates ([[]] x inorm , y inorm ): x inorm = ([[]] x i - c x ) / f x y inorm = ([[]] y i - c y ) / / f y At this time, assuming that Z ci = 1 in the camera coordinate system, Z ci is the z axis coordinate of the feature point in the camera coordinate system, then the normalized coordinates ([[]]x inorm , y inorm , 1) Approximately represent the coordinates in the camera coordinate system; For n feature points, let the coordinates of the feature points in the camera coordinate system be ([[]] X ci , Y ci , Z ci ), satisfying the following relationship:

[0050] Then there is: x i = f x u i + c x ; y i = f y v i + c y ; u i = ([[]] x i - c x ) / f x ; v i = ([[]] y i - c y ) / f y ; Establish a system of equations through the coordinates of multiple feature points and solve using the least squares method Z ci , Let the equation of the i th feature point be:

[0051] Combine the equations of the n feature points and write them in matrix form Ax = b , where A is the coefficient matrix, x is the unknown vector, and this vector contains u i , v i andZ ci The relevant information of b is a constant vector, which is solved by the least squares method x , and Z ci the value of is obtained, and then the coordinates of the feature points in the camera coordinate system are obtained ([[]] X ci , Y ci , Z ci ).

[0052] As Figure 5 shown, set the extrinsic parameter matrix of the infrared camera R | T , where R is a 3×3 rotation matrix, T is a 3×1 translation vector; The coordinates of the feature points in the camera coordinate system ([[]] X ci , Y ci , Z ci ) and the coordinates of the feature points in the world coordinate system ([[]] X wi , Y wi , Z wi ) have the following transformation relationship:

[0053] Convert the coordinates of the feature points in the camera coordinate system to the coordinates of the feature points in the world coordinate system; Calculate the distance between the target feature point in the world coordinate system and the camera origin (the camera origin coordinates in the world coordinate system are (0,0,0)) D i , for the i th feature point, the distance D i is calculated as follows:

[0054] The distances D i calculated for the four feature points ( i =1,2,3,4) are averaged to obtain the final spatial distance D , and the calculation formula is:

[0055] Output the calculated spatial distance D as the non-contact ranging result.

[0056] Therefore, this embodiment provides the whole process of obtaining the information of the rotation bounding box to calculate the spatial distance, including steps such as determining the characteristic point coordinates, converting the image coordinates to the camera coordinates, converting the camera coordinates to the world coordinates, and calculating the spatial distance, etc., which can effectively calculate the spatial distance between the power equipment and the camera. During the calculation process, geometric parameters such as the internal parameter matrix and external parameter matrix of the camera, as well as various parameters of the rotation bounding box (center coordinates, width, height, rotation angle) are fully considered, ensuring the accuracy and reliability of the calculation results.

[0057] In the target detection stage, an improved YOLOv5 model is used to optimize the rotation target detection, which can more accurately detect the information of the rotation bounding box of the power equipment target in the infrared image. By using the information of the rotation bounding box to determine the characteristic point coordinates of the target on the image plane, and through more accurately detecting the rotation target bounding box and combining the geometric projection principle to calculate the distance, the accuracy of non-contact ranging can be improved, providing more reliable data support for the monitoring and maintenance of power equipment.

[0058] In this embodiment, the loss function of the YOLOv5 target detection model L is expressed as the following formula: L = λ 1 L cls + λ 2 L reg + λ 3 L θ where λ 1, λ 2, λ 3 are weight coefficients; during the training process, the validation set is used to evaluate the network model in training, and the fitting situation of the network is judged according to the evaluation indexes on the validation set.

[0059] The classification loss of this embodiment is calculated using binary cross-entropy. In order to improve the sensitivity of the bounding box to geometric factors during the regression process, a positioning loss F-CIOU is designed, and the formula is as follows:

[0060] where, B and B mb are the predicted box and the target box respectively, b and b mb represent the center points of the predicted box and the target box respectively, p (·) represents the Euclidean distance between the predicted center point and the target bounding box,c is the diagonal length of the smallest bounding box covering two boxes, w mb and h mb are the width and height of the target bounding box, w and h are the width and height of the predicted bounding box. F-CIOU retains the advantages of CIOU loss.

[0061] In this embodiment, an improved YOLOv5 network is constructed using the PyTorch deep learning framework, and the functions and tools provided by this framework are used to implement model initialization, training, parameter tuning, and performance evaluation; programming experiments are carried out by combining Python and Java languages. Python code is written to implement model construction and training, and Java code is written to implement multi-target distance calculation and dynamic update of multi-target ranging results. The entire setup process of the non-contact ranging system for power equipment is completed by executing the code in the Python and Java runtime environments.

[0062] In this embodiment, the YOLOv5 object detection model includes an input layer, a backbone network, a neck network, and a detection head to achieve end-to-end object detection capabilities, ensuring accurate positioning and recognition of cable terminals in infrared images at multiple angles and distances, and providing an accurate and robust feature basis for subsequent analysis tasks.

[0063] Specifically, the backbone network adopts the CSPDarknet architecture. This architecture divides the feature map into two parts for independent processing and then merges them through the CrossStage Partial (CSP) technology, effectively reducing the computational amount while enhancing the feature extraction ability. It can efficiently extract multi-scale image features from infrared images and provide rich semantic information for subsequent network processing; an SPPF module is embedded between the backbone network and the detection head. The SPPF module consists of multiple max-pooling layers of different scales in parallel. By performing pooling operations on the input feature map at different scales and splicing the pooling results, the receptive field information of the network for targets of different scales is significantly enhanced, enabling the network to better capture the context information of the target and improving the accuracy of object detection; the PANet structure is introduced into the neck network. The PANet structure fuses and transmits feature maps of different levels through a top-down and bottom-up two-way feature fusion path. Specifically, the top-down path transmits high-level semantic information to the low level, enhancing the semantic expression ability of the low-level feature map; the bottom-up path transmits the low-level position information to the high level, making up for the lack of position information in the high-level feature map, realizing the full fusion of upper and lower layer features, and improving the detection performance of the network for targets of different scales.

[0064] Among them, the detection head predicts the target bounding boxes for feature maps of different scales, specifically divided into three detection heads of different scales, which are respectively applicable to the recognition scenarios of large, medium, and small-sized targets; each detection head processes the input feature map through convolutional operations, and the output includes at least the class probability of the target, the center coordinates of the rotated bounding box ( x c , y c ), width w , height h and rotation angle θ information, so as to realize the positioning and classification of the target.

[0065] Further, load the improved YOLOv5 network, set the network parameters for training, the network training cycle is 300, the batch size is 8, use the Adam optimizer, set the initial learning rate of the network to 1e-3, the learning rate for the network to stop updating is 0.1, and the weight decay coefficient is 0.001.

[0066] In this example, the Adam optimizer and the cosine annealing decay method are used to update the learning rate, and the formula is as follows:

[0067] Among them, lr is the learning rate, E is the total number of training epochs, lr f is the learning rate for stopping updating.

[0068] In each round of training, samples are randomly drawn from the training set in batches and input into the network. The samples are first preprocessed through the input layer of the network, and then sequentially pass through the backbone network, SPPF module, neck network and detection head to obtain the output results of the network encoder; Compare the output results of the network encoder with the corresponding true labels, calculate the loss value through the designed loss function L , and use the backpropagation algorithm to update and optimize the network parameters according to the loss value, so that the network continuously approaches the optimal solution; When the last batch of the training set is drawn and the training is completed, all batches in the validation set are sequentially input into the network for verification, so as to judge the fitting situation of the network; during the verification process, the network makes predictions on the validation set samples to obtain the prediction results; if the loss value of the training set passing through the network is basically the same as the loss value of the validation set passing through the network, the network performance is better; if the loss value of the training set passing through the network is much smaller than the loss value of the validation set passing through the network, the overfitting degree of the network is serious.

[0069] Compare the prediction results with the true labels of the validation set, calculate a series of evaluation metrics to evaluate the performance of the current model on unseen samples, and determine whether there is overfitting or underfitting in the network; Repeat the training and validation steps until all samples in the training set complete the iterative update for the set number of cycles, complete the entire model training process, and obtain the trained YOLOv5 object detection model. The prediction results of the improved YOLOv5 object detection model are as Figure 6 and 7 shown.

[0070] In the above step S4, establish a target buffer. After each object detection is completed and the ranging result is calculated, record the object category, location information, ranging result, and the current frame number into the target buffer; if there is already a record with the same object number in the target buffer, update the location information, ranging result, and the most recent update frame number in this record; if not, assign a unique object number to the new object and create a new data record for storage.

[0071] To avoid storing too much invalid or outdated object information in the target buffer, set a maximum buffer duration threshold. Regularly check the difference between the most recent update frame number of each object in the target buffer and the current frame number. If this difference exceeds the maximum buffer duration threshold, it is considered that the object has left the detection area, and it will be deleted from the target buffer to release the buffer space and ensure the validity of the object information stored in the buffer.

[0072] For each object in the new detection frame, calculate the intersection over union (IoU) value between its rotated bounding box and all the rotated bounding boxes of the same category objects in the target buffer; Set an IoU matching threshold (e.g., 0.5). When the IoU value of an object in the new detection frame and a certain object of the same category in the target buffer is greater than or equal to this matching threshold, it is determined that these two objects are the same object; when the IoU value is less than this matching threshold, or the category information is inconsistent, it is determined that the object in the new detection frame is a new object; For the case where it is determined that they are the same object, use the weighted average algorithm to update the ranging result.

[0073] Among them, the weighted average formula is:

[0074] Among them D new is the updated ranging result, D current is the ranging result of the current frame, D i-1 is the ranging result of the previous i -1 frame,α is the weight coefficient of the ranging result for the current frame, where 0 < α < 1; through this weighted average algorithm, the ranging fluctuations caused by factors such as single-frame image noise or temporary target occlusion are eliminated, thereby improving the ranging stability and robustness of the system.

[0075] For the case where a new target is determined, a new unique target number is assigned to this target, a new data record is created according to the data record structure of the target buffer area, and information such as the category, location information, ranging result, and current frame sequence number of the new target are stored in this record and added to the target buffer area. After that, this new target will participate in the target matching and ranging result update process of subsequent detection frames.

[0076] Specifically, the calculation of the intersection over union (IoU) value of the rotated bounding boxes includes: First, according to the parameters of the two rotated bounding boxes, calculate their intersection area; among them, the rotated bounding box is converted into a polygon, and then the area of the intersection polygon is calculated through the polygon intersection algorithm; Calculate the areas of the two rotated bounding boxes respectively; Divide the area of the intersection polygon by the sum of the areas of the two rotated bounding boxes minus the area of the intersection polygon to obtain the IoU value, that is IoU = S intersection / S box1 + S box2 - S intersection , where S intersection is the area of the intersection polygon, S box1 and S box2 are the areas of the two rotated bounding boxes respectively.

[0077] This embodiment further proposes a dynamic update mechanism based on the historical detection status for improving the stability of target ranging, as Figure 8 shown. When the system first detects an image, it makes a confidence judgment on all the identified targets, only retains the targets with a confidence higher than the preset threshold as valid targets for ranging, and caches their pixel positions, category labels, and ranging results in the system memory for use as a reference for subsequent images.

[0078] When the system enters the subsequent detection cycle, the system will first attempt to perform the target detection operation. If the target detection module fails to identify any target, or the confidence level of the detection result is lower than the set threshold, the system will automatically call the previously cached target information and output it as the recognition result of the current image, avoiding the loss of the target caused by temporary image occlusion, thereby improving the ranging continuity and output stability.

[0079] When a target is detected in the new image, the system matches the newly detected target with the cached target of the previous image. To improve the matching accuracy, in this embodiment, a matching strategy based on the IoU (Intersection over Union) overlap degree is adopted, and the overlap ratio of the current target and the historical target bounding box is calculated. The specific formula is as follows:

[0080] Among them, B new represents the target bounding box detected in the current image, B old represents the target bounding box cached in the previous image, A inter is the area of the intersection region, A union is the area of the union region.

[0081] If the IoU of the current target and the cached target is greater than or equal to the set matching threshold and the category information is the same, it is determined to be the same target, and the system updates its position and ranging information; otherwise, it is determined to be a new target, and a new number is re-allocated and enters the ranging process. The results of multi-target ranging are as Figure 9 shown.

[0082] Through the above mechanism, the present invention realizes the effective tracking and dynamic update of the target, has the characteristics of exception fault tolerance and ranging continuity optimization, and is particularly suitable for application scenarios with unstable factors such as occlusion, light change, and jitter shooting in power inspection. It can realize the spatial positioning of multiple power equipment, provide effective distance information support for intelligent inspection, equipment layout analysis and warning systems in high-voltage environments, and has wide engineering application value.

[0083] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered to be within the scope described in this specification.

[0084] The above embodiments are only the preferred embodiments of the present invention, and the scope of protection of the present invention cannot be limited thereby. Any non-substantive changes and substitutions made by those skilled in the art based on the present invention fall within the scope of protection required by the present invention.

Claims

1. A non-contact ranging method for power equipment based on rotating target detection and infrared images, characterized in that It includes the following steps: Collect infrared images of power equipment to generate an infrared image dataset containing target information of power equipment; Build a YOLOv5 object detection model, which is used to detect the rotated bounding box information of power equipment targets in infrared images; among them, the loss function of the YOLOv5 object detection model L includes classification loss L cls , bounding box regression loss L reg and rotation angle prediction loss L θ ; Use the trained YOLOv5 object detection model to perform object detection on the input infrared image, and obtain the rotated bounding box information of the power equipment object B =( x c , y c , w , h , θ ), determine the coordinates of the feature points of the object on the image plane based on the rotated bounding box information, and calculate the spatial distance according to the distance relationship between the object feature points and the camera origin in the world coordinate system D , and output the non-contact ranging result; where ( x c , y c ) is the center coordinate of the bounding box, w is the width of the bounding box, h is the height of the bounding box, θ is the rotation angle of the bounding box; Based on the historical object detection results, a dynamic update algorithm is used to correct the current object position information and ranging results; among them, the dynamic update algorithm comprehensively considers the object detection results of the current frame and the historical detection results of the previous n frames, and smooths the object position and ranging values through a weighted average algorithm.

2. The method according to claim 1, wherein In the process of making the infrared image dataset and preprocessing the dataset, it includes the following steps: Design cable terminal acquisition tasks for various actual usage scenarios, and the actual usage scenarios at least include scenarios under different environmental temperatures, different lighting conditions, different weather conditions, and different operating states of the cable terminal; Uniformly format all the obtained infrared images, and classify and store the formatted infrared images according to the timestamp, device number, and scenario number; Use a rotated rectangle box to label the targets in the infrared image, so that the rotated rectangle box fits the outer contour of the target. The labeling information includes the four vertex coordinates of the rotated target box in the image coordinate system and the rotation angle θ; at the same time, label the category information for each rotated target box, and the category information includes cable terminals and lightning arresters; Split all the labeled samples into a training set, a validation set, and a test set according to a preset ratio; among them, introduce a variety of data augmentation operations on the training set, at least including angle rotation, contrast change, mirror flipping, and cropping methods.

3. The method according to claim 1, wherein The determination of the characteristic point coordinates of the target on the image plane based on the rotated bounding box information includes: For the rotated bounding box B =( x c , y c , w , h , θ ), select the four vertices of the rotated bounding box as the feature points of the target on the image plane. Let the rotated bounding box be centered at the center point ( x c , y c ), with the width direction being the x -axis direction and the height direction being the y -axis direction to establish a local coordinate system. In the local coordinate system, the coordinates of the four vertices are respectively P 1(- w / 2,- h / 2)、 P 2( w / 2,- h / 2)、 P 3( w / 2, h / 2) and P 4(- w / 2, h / 2); According to the rotation angle θ Perform a rotation transformation on the vertex coordinates in the local coordinate system using the rotation matrix R θ Perform a rotation transformation on the vertex coordinates in the local coordinate system. The rotation matrix R θ has the following expression: For the vertex coordinates in the local coordinate system ( x local , y local ), the calculation formula for the rotated coordinates ( x rot , y rot ) is as follows: Convert the local coordinates of the four vertices P 1、 P 2、 P 3、 P 4 to obtain the coordinates after rotation in the local coordinate system P 1rot ( x 1rot , y 1rot )、 P 2rot ( x 2rot , y 2rot )、 P 3rot ( x 3rot , y 3rot )、 P 4rot ( x 4rot , y 4rot ); Convert the vertex coordinates in the rotated local coordinate system to the global coordinate system of the image plane to obtain the coordinates of the four feature points on the image plane ( x i , y i ). ( i = 1, 2, 3, 4), and the conversion formula is: x i = x rot + x c 、 y i = y rot + y c ; Based on P 1rot 、 P 2rot 、 P 3rot 、 P 4rot the coordinates of four feature points on the image plane are obtained P 1img ( x 1, y 1), P 2img ( x 2, y 2), P 3img ( x 3, y 3), P 4img ( x 4, y 4).

4. The method according to claim 1, characterized in that, Calculating the spatial distance according to the distance relationship between the target feature point and the camera origin in the world coordinate system D, including: Set the internal parameter matrix of the infrared camera K as follows: wherein f x and f y are respectively the focal lengths of the camera in the x axis and y axis directions, c x and c y are respectively the coordinates of the principal point of the image in the x axis and y axis directions; For the feature point coordinates ([ x i ], [ y i ]) on the image plane, perform normalization to obtain the normalized coordinates ([ x inorm ], [ y inorm ]) as follows: x i , y i ) x inorm , y inorm ) x inorm =( x i - c x ) / f x y inorm =( y i - c y ) / f y At this time, assuming that in the camera coordinate system Z ci = 1, Z ci is the z axis coordinate of the feature point in the camera coordinate system, then the normalized coordinate ([[]] x inorm , y inorm , 1) approximately represents the coordinate in the camera coordinate system; For n feature points, let the coordinates of the feature points in the camera coordinate system be ( X ci , Y ci , Z ci ), satisfying the following relationship: Then there is: x i = f x u i + c x ; y i = f y v i + c y ; u i = ( x i - c x ) / f x ; v i = ( y i - c y ) / f y ; Establish a system of equations based on the coordinates of multiple feature points and solve it using the least squares method Z ci , let the equation of the i th feature point be: Will n The equations of the characteristic points are combined and written in matrix form Ax = b ,in A is the coefficient matrix, x is a vector of unknown numbers, which contains u i , v i and Z ci Related information, b is a constant vector, and is solved by the least squares method x ,get Z ci The value of , and then get the feature point coordinates in the camera coordinate system ( X ci , Y ci , Z ci ).

5. The method according to claim 4, wherein: Set the external parameter matrix of the infrared camera R | T , where R is a 3×3 rotation matrix, T is a 3×1 translation vector; Feature point coordinates in the camera coordinate system ( X ci , Y ci , Z ci ) and the conversion relationship with the feature point coordinates in the world coordinate system ( X wi , Y wi , Z wi ) is as follows: Convert the characteristic point coordinates in the camera coordinate system into the characteristic point coordinates in the world coordinate system; Calculate the distance between the target feature point in the world coordinate system and the camera origin D i , for the i th feature point, the distance D i is calculated by the formula: The distances calculated for the four feature points D i ( i = 1, 2, 3, 4) Take the average to obtain the final spatial distance D , and the calculation formula is: Output the calculated spatial distance D As the non-contact ranging result.

6. The method according to claim 2, wherein: The loss function of the YOLOv5 object detection model L is expressed as the following formula: L = λ 1 L cls + λ 2 L reg + λ 3 L θ Among them λ 1、 λ 2、 λ 3 is the weight coefficient; During the training process, use the validation set to evaluate the network model during training, and judge the fitting situation of the network according to the evaluation metrics on the validation set.

7. The method according to any one of claims 1 to 6, wherein: The YOLOv5 target detection model includes an input layer, a backbone network, a neck network, and a detection head. The backbone network adopts the CSPDarknet architecture and is used to extract multi-scale image features from infrared images; an SPPF module is embedded between the backbone network and the detection head. The SPPF module consists of multiple max-pooling layers with different scales in parallel. By performing pooling operations on the input feature map at different scales and splicing the pooling results; Introduce the PANet structure in the neck network. The PANet structure fuses and transmits feature maps at different levels through a top-down and bottom-up two-way feature fusion path; Among them, the detection head predicts the target bounding boxes for feature maps of different scales. Specifically, it is divided into three detection heads of different scales, which are respectively applicable to the recognition scenarios of large, medium, and small-sized targets; each detection head processes the input feature map through convolution operations, and the output includes at least the class probability of the target, the center coordinates ( x c , y c ) of the rotated bounding box, the width w , the height h , and the rotation angle θ information, so as to realize the positioning and classification of the target.

8. The method according to claim 7, wherein: In each round of training, randomly extract samples from the training set in batches and input them into the network. The samples are first preprocessed through the input layer of the network, and then sequentially pass through the backbone network, the SPPF module, the neck network, and the detection head to obtain the output results of the network encoder; Compare the output result of the network encoder with the corresponding true label, and calculate the loss value through the designed loss function L Calculate the loss value, and use the backpropagation algorithm to update and optimize the network parameters according to the loss value, so that the network continuously approaches the optimal solution; When the last batch of the training set is extracted and the training is completed, input all batches in the validation set into the network in sequence for verification; During the verification process, the network makes predictions on the validation set samples to obtain prediction results; Compare the prediction results with the true labels of the validation set, and calculate a series of evaluation metrics to evaluate the performance of the current model on unseen samples, and determine whether there is overfitting or underfitting in the network; Repeat the training and validation steps until all samples in the training set complete the iterative update for the set number of cycles, complete the entire model training process, and obtain the trained YOLOv5 object detection model.

9. The method according to any one of claims 1 to 6, characterized in that: Establish an object buffer area. After each object detection is completed and the ranging result is calculated, record the object category, location information, ranging result, and the current frame number into the object buffer area; For each object in the new detection frame, calculate the intersection over union (IoU) value between its rotated bounding box and all rotated bounding boxes of the same category objects in the object buffer area; Set an IoU matching threshold. When the IoU value of an object in the new detection frame and a rotated bounding box of a certain object of the same category in the object buffer area is greater than or equal to the matching threshold, then determine that these two objects are the same object; when the IoU value is less than the matching threshold, or the category information is inconsistent, then determine that the object in the new detection frame is a new object; For the case where the objects are determined to be the same object, use the weighted average algorithm to update the ranging result.

10. The method according to claim 9, characterized in that: The calculation of the intersection over union (IoU) value of the rotated bounding box includes: First, according to the parameters of the two rotated bounding boxes, calculate their intersection area; wherein, convert the rotated bounding box into a polygon, and then calculate the area of the intersection polygon through the polygon intersection algorithm; Calculate the areas of the two rotated bounding boxes respectively; Divide the area of the intersection polygon by the sum of the areas of the two rotated bounding boxes minus the area of the intersection polygon to obtain the IoU value, that is IoU = S intersection / S box1 + S box2 - S intersection , where S intersection is the area of the intersection polygon, S box1 and S box2 are the areas of the two rotated bounding boxes respectively.

Citation Information

Patent Citations

  • Identification detection method and device and electronic system

    CN109670503A

Cited By

  • Language guidance feature decoupling infrared target detection method

    CN121330249A