Construction method of intelligent Agent for detecting collapsed transmission tower in remote sensing image

By combining the improved YOLOV 10 object detection small model and multimodal large model, the accuracy of transmission tower collapse detection in remote sensing images is solved, and higher detection accuracy and reliability are achieved.

CN120356086APending Publication Date: 2025-07-22INFORMATION & COMM BRANCH OF STATE GRID JIANGSU ELECTRIC POWER
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510223441.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-27
Publication Date
2025-07-22

Smart Images

  • Figure CN120356086A_ABST
    Figure CN120356086A_ABST
Patent Text Reader

Abstract

The invention discloses a method for constructing an intelligent Agent for detecting a collapsed transmission tower in a few-sample remote sensing image. The method comprises the following steps: training a model to detect the transmission tower; calculating and analyzing typical transmission tower characteristics; and the intelligent Agent is used for carrying out collapse detection on the selected towers not exceeding F, and a user is reported in time when a collapsed tower exists. According to the method, the power transmission towers in the remote sensing image are detected by using the improved YOLOV 10 target detection small model, and then the collapsed power transmission towers are distinguished from the detected power transmission towers by using the multi-mode large model and the user is notified in time, so that the detail difference between the collapsed power transmission towers and the non-collapsed power transmission towers can be more concerned, meanwhile, the method has relatively good few-sample learning ability, and the method is suitable for large-scale popularization and application. Compared with a single small model or a large model, the method achieves a better effect, and greatly improves the detection accuracy of the collapse of the transmission tower under the condition that the number of samples is small.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of artificial intelligence large models, and specifically relates to a method for constructing an intelligent Agent for detecting collapsed transmission towers in remote sensing images.

[0002] The target detection small model is combined with the multi-modal large model for detecting collapsed transmission towers. First, the improved YOLOV 10 target detection small model is used to detect the transmission towers in the remote sensing image, and then the multi-modal large model is used to distinguish the collapsed transmission towers from the detected transmission towers and notify the user in time. This method for detecting the collapse of transmission towers can pay more attention to the detailed differences between collapsed and non-collapsed towers, and at the same time has good few-shot learning ability, obtaining better results than using only the small model or the large model, and greatly improving the accuracy of tower collapse detection. Background Art

[0003] The detection of collapsed transmission towers in remote sensing images is of great significance for the post-disaster reconstruction of power companies. It can help power companies discover collapsed towers in the first time, so as to assist power companies in formulating scientific and reasonable post-disaster reconstruction plans for power facilities. The object detection in remote sensing images is an application field of ordinary image object detection. According to the different detection strategies adopted, the existing deep learning object detection methods for remote sensing images can be roughly divided into object detection methods based on convolutional neural networks and those based on Transformers. Among the methods based on convolutional neural networks, they can be further divided into one-stage object detection methods based on regression analysis and two-stage object detection methods based on candidate regions. The one-stage object detection method based on regression analysis is an end-to-end object detection algorithm. Since it does not need to extract object candidate boxes, it is faster, and the shortcoming of its detection accuracy has been greatly improved in continuous improvement and has been widely applied. The YOLO algorithm proposed in the literature [You only look once: unified, real-time object detection] is a typical one-stage object detection method, which divides the image into small grids of the same size and directly regresses the category and bounding box of the target on these small regions. Later, some researchers have continuously improved it, and the latest version is the YOLO V10 algorithm proposed in the literature [YOLOv10: Real-Time End-to-End Object Detection]. Compared with ordinary images, the object detection task of remote sensing images has the characteristics of high resolution, large size, complex background, and often contains smaller targets, which increases the detection difficulty. The detection of collapsed transmission towers in remote sensing images has the characteristic of scarce samples compared with other remote sensing object detections.

[0004] In the past three years, the rapid development of large language model technology led by ChatGPT of OpenAI has profoundly changed the way people access and utilize information. Large language models such as ChatGPT, LLAMA, and Qwen have been widely used. These LLMs have demonstrated extraordinary emerging capabilities and shown strong zero-shot / few-shot reasoning performance in most natural language processing (NLP) tasks and complex real-world applications. However, they are essentially "blind" to vision because they can only understand discrete text. On the other hand, large visual models (LVMs) can "see clearly", but they are not as good as large language models in reasoning. The combination of large language models and large visual models has given rise to a new field of multimodal large models. Multimodal large models have the ability to receive, reason, and output multimodal information and have strong zero-shot / few-shot reasoning capabilities. However, at present, the hallucination phenomenon existing in multimodal large models and their poor performance in processing data in fine-grained fields such as remote sensing images make it difficult to obtain the expected results when directly applying them to the task of detecting the collapse of remote sensing transmission towers. Multimodal hallucination refers to the situation where the responses generated by multimodal large language models (MLLMs) do not match the image content. This phenomenon significantly affects the reliability and accuracy of multimodal large models in actual industrial applications and has become an important obstacle restricting their wide application. Summary of the Invention

[0005] The object of the present invention is to provide a method for constructing an intelligent agent for detecting collapsed transmission towers in remote sensing images. First, an improved YOLOV 10 small target detection model is used to detect transmission towers in remote sensing images, and then a multimodal large model is used to distinguish collapsed transmission towers from the detected transmission towers and notify users in a timely manner. It can pay more attention to the detailed differences between collapsed and non-collapsed towers, and at the same time has good few-shot learning ability, obtaining better results than using only small models or large models, and greatly improving the accuracy of tower collapse detection.

[0006] The object of the present invention is achieved through the following technical solutions:

[0007] A method for constructing an intelligent agent for detecting collapsed transmission towers in remote sensing images, characterized by including the following steps:

[0008] 1) Design and train a model to detect transmission towers;

[0009] 2) Calculate and analyze the characteristics of typical transmission towers;

[0010] 3) Use an intelligent agent to detect collapsed transmission towers in each remote sensing image and report to the user in a timely manner when there are collapsed towers.

[0011] The step of designing and training a model to detect transmission towers is as follows:

[0012] Step 1: Collect remote sensing images and label transmission towers on the collected sample images.

[0013] Step 2: Modify the YOLOV10 algorithm by replacing the 7×7 large kernel in the second depth convolution of the compression conversion block CIB with two parallel 3×3 dilated convolutions, one with a dilation rate of 1 and the other with a dilation rate of 2. After expanding to the same shape as the result of the 3×3 convolution, add it to the result of the 3×3 convolution, which is expressed by the following formula:

[0014]

[0015] where w r (i,j) is obtained by training and learning. X is the matrix to be convolved, W is the convolution kernel weight matrix, and m and n are the rows and columns of the matrix elements respectively.

[0016] Step 3: Train on the improved YOLO V10 algorithm to obtain a remote sensing image transmission tower detection model.

[0017] The steps for calculating and analyzing the characteristics of typical transmission towers are as follows:

[0018] Step 1: Extract the DINOV2768-dimensional embedding feature tr i =(trf i 1,trf i 2,...trf i 768 )

[0019] Step 2: Calculate the average value of the features tr i of all transmission tower areas in the training samples to obtain the characteristics of typical transmission towers

[0020] The steps for using an intelligent Agent to detect collapsed transmission towers in each remote sensing image and report to the user in a timely manner when there is a collapsed tower are as follows:

[0021] Step 1: Build an intelligent Agent for detecting collapsed transmission towers, using a multi-modal large model as the base and decision center, and using a persistent database as its memory storage medium.

[0022] Step 2 constructs a knowledge base for the Agent. The knowledge base is directly provided to the Agent in the form of a txt file. If there are two small images of transmission towers detected by YOLOV10 that contain collapsed towers, denoted as 1.jpg and 2.jpg, and there are two other small images that do not contain collapsed towers, then the knowledge base is constructed as follows: 'The following two images are of collapsed transmission towers: 1.jpg, 2.jpg. The following two images are of non-collapsed transmission towers: 3.jpg, 4.jpg.' Generally speaking, since there are few remotely sensed images of collapsed transmission towers that can be collected, the number of images of collapsed towers included in the knowledge base <= 10 images.

[0023] Step 3 accepts each question from the user and uses the Embedding algorithm to convert the question into a question vector, denoted as vector UQ E =(v i,1 , v i,2 ,......v i,N ,). Compare vector UQ E with a pre-defined question vector DQ E =(v'1, v'2,......v' N ). The pre-defined question is the sentence 'Are there any collapsed transmission towers in the following figure?' Calculate the cosine similarity of these two vectors

[0024]

[0025] where ||v i || and ||v'|| are the norms of vectors UQ E and DQ E respectively.

[0026] Step 4 If simi(UQ E , DQ E ) < T, it means that the question asked by the user in the current round has nothing to do with the detection of collapsed transmission towers. Therefore, this question is directly submitted to the multi-modal large model for answering. If simi(UQ E , DQ E ) >= T, that is, the user question vector UQ E is similar to the preset question vector DQ E , indicating that the question asked by the user in this round is related to tower collapse, and continue to the next step

[0027] Step 5 When they are similar, call the transmission tower detection model trained in the foregoing claims to detect all the transmission towers in the current remotely sensed image and store them in a directory in the form of one picture file for each detected tower. If there are many picture files, set a limit value F for the total number of files. When the number of pictures is less than this value, process all pictures; otherwise, process F pictures.

[0028] Step 6 When the number of pictures is greater than F, extract the features of each small picture, and select F small pictures that are least similar to the typical transmission tower features obtained in the foregoing claims, that is where K represents the number of transmission towers detected in the current picture.

[0029] Step 7 After selecting at most F small tower pictures, construct a transmission tower collapse detection problem and submit it to the multi-modal large model for answering. The constructed question is as follows: "When a transmission tower collapses, it generally lies obliquely on the ground and may also be somewhat bent. Carefully observe the pictures in the knowledge base and then answer the following question: Is this picture more similar to a collapsed transmission tower or an uncollapsed transmission tower?" + pic_path, where pic_path is the temporary storage path of the small picture. If there are f (f <= F) small pictures, the above question will be submitted to the intelligent agent f times.

[0030] Step 8 When the large model in the intelligent agent determines that there is a collapsed transmission tower in a certain small picture, synthesize the answers returned by the large model in the intelligent agent and feedback the synthesized result to the user; if it is determined that there is no collapsed tower, and if the judgment result is connected to an alarm or notification program, the alarm or notification will not be triggered.

[0031] The beneficial effects of the present invention are as follows:

[0032] The present invention first uses an improved YOLOV 10 object detection small model to detect transmission towers in remote sensing images, and then uses a multi-modal large model to distinguish collapsed transmission towers from the detected transmission towers and notify the user in time. The present invention combines large and small models, uses the accurate object detection ability of the small model to make up for the deficiency of the large model in comprehensively and accurately positioning objects, and uses the zero-shot and few-shot learning ability of the large model to make up for the deficiency of the small model in requiring sufficient training samples, so as to accurately detect the target of transmission tower collapse.

[0033] Compared with ordinary object detection algorithms, the present invention can obtain better detection effects with fewer samples, and achieves better effects than using only small models or large models. Brief Description of the Drawings

[0034] Figure 1 It is the overall flow chart of the present invention.

[0035] Figure 2 It is the structural diagram of the improved YOLOV 10CIB of the present invention.

[0036] Figure 3 It is the detection effect diagram of the present invention. Detailed Embodiments

[0037] An intelligent Agent construction method for detecting collapsed transmission towers in remote sensing images, as follows Figure 1 and Figure 2 shown, including the following steps:

[0038] 1) Design and train a model to detect transmission towers, the steps are as follows:

[0039] Step 1: Collect remote sensing images and label the transmission towers in the collected sample images.

[0040] Step 2: Modify the YOLOV10 algorithm. Replace the 7×7 large kernel in the second depth convolution of the compression conversion block CIB with two parallel 3×3 dilated convolutions. One has a dilation rate of 1 and the other has a dilation rate of 2. After expanding to the same shape as the result of the 3×3 convolution and then adding it to the result of the 3×3 convolution, it is represented by the following formula:

[0041]

[0042] where w r (i,j) is obtained by training and learning. X is the matrix to be convolved, W is the convolution kernel weight matrix, and m and n are the rows and columns of the matrix elements respectively.

[0043] Step 3: Train on the improved YOLO V10 algorithm to obtain a remote sensing image transmission tower detection model.

[0044] 2) Calculate and analyze the typical transmission tower features, the steps are as follows:

[0045] Step 1: Extract the DINOV2768-dimensional embedding feature tr i =(trf i 1,trf i 2,...trf i 768 ) for the transmission tower areas marked in the training samples.

[0046] Step 2: Calculate the average value of the features tr i of all the transmission tower areas in the training samples to obtain the typical transmission tower features

[0047] 3) Use the intelligent Agent to detect collapsed transmission towers in each remote sensing image and report to the user in time when there are collapsed towers. Figure 3 This is the detection effect diagram of the present invention, and the steps are as follows:

[0048] Step 1: Construct an intelligent Agent for detecting the collapse of transmission towers, using a multimodal large model as the base and decision center, and using a persistent database as its memory storage medium.

[0049] Step 2 constructs a knowledge base for the Agent. The knowledge base is directly provided to the Agent in the form of a txt file. If there are two small images of transmission towers detected by YOLOV10 that contain collapsed towers, denoted as 1.jpg and 2.jpg, and there are another two small images that do not contain collapsed towers, then the knowledge base is constructed as follows: 'The following two images are of collapsed transmission towers: 1.jpg, 2.jpg. The following two images are of non-collapsed transmission towers: 3.jpg, 4.jpg.' Generally speaking, since there are few remotely sensed images of collapsed transmission towers that can be collected, the number of images of collapsed towers contained in the knowledge base <= 10 images.

[0050] Step 3 accepts each question from the user and uses the Embedding embedding algorithm to convert the question into a question vector, denoted as vector UQ E =(v i,1 , v i,2 ,......v i,N ,). Compare the vector UQ E with the pre-defined question vector DQ E =(v'1, v'2,......v' N ). The pre-defined question is the sentence 'Are there any collapsed transmission towers in the following figure?' Calculate the cosine similarity of these two vectors

[0051]

[0052] where ||v i || and ||v'|| are the norms of the vectors UQ E and DQ E respectively.

[0053] Step 4 If simi(UQ E , DQ E ) < T (here T represents a similarity threshold, generally taken as 0.5), it means that the question raised by the user in the current round has nothing to do with the detection of collapsed transmission towers. Therefore, this question is directly submitted to the multi-modal large model for answering. If simi(UQ E , DQ E ) >= T, that is, the user question vector UQ E is similar to the preset question vector DQ E , indicating that the question asked by the user in this round is related to the collapse of the tower, and continue to the next step of processing

[0054] Step 5 When they are similar, call the transmission tower detection model trained in the foregoing claims to detect all the transmission towers in the current remotely sensed image and store them in the directory in the form of one picture file for each detected tower. If there are many picture files, set a limit value F for the total number of files. When the number of pictures is less than this value, process all the pictures; otherwise, process F pictures.

[0055] Step 6 When the number of pictures is greater than F, extract the features of each small picture, and select F small pictures that are least similar to the typical transmission tower features obtained in the foregoing claims, that is 1 ≤ k ≤ K, where K represents the number of transmission towers detected in the current picture.

[0056] Step 7 After selecting at most F small tower pictures, construct a transmission tower collapse detection problem and submit it to the multi-modal large model for answering. The constructed question is as follows: "When a transmission tower collapses, it generally lies obliquely on the ground and may also be somewhat bent. Carefully observe the pictures in the knowledge base and then answer the following question: Is this picture more similar to a collapsed transmission tower or a non-collapsed transmission tower?" + pic_path, where pic_path is the temporary storage path of the small picture. If there are f (f <= F) small pictures, the above question will be submitted to the intelligent agent f times.

[0057] Step 8 When the large model in the intelligent agent determines that there is a collapsed transmission tower in a certain small picture, synthesize the answers returned by the large model in the intelligent agent and feedback the synthesized result to the user; if it is determined that there is no collapsed tower, and if the judgment result is connected to an alarm or notification program, the alarm or notification is not triggered.

Claims

1. A method for constructing an intelligent agent for detecting collapsed transmission towers in remote sensing images, characterized in that, It includes the following steps: 1) Design and train a model to detect transmission towers; specifically as follows: Step 11: Collect remote sensing images and label the transmission towers in the collected sample images; Step 12: Modify the YOLOV10 algorithm. Replace the 7×7 large kernel in the second depth convolution of the compression conversion block CIB with two parallel 3×3 dilated convolutions. One has a dilation rate of 1 and the other has a dilation rate of 2. After expanding to the same shape as the result of the 3×3 convolution and then adding it to the result of the 3×3 convolution, it is expressed by the following formula: w r (i, j), where w r (i, j) is obtained by training and learning, X is the matrix to be convolved, W is the convolutional kernel weight matrix, and m and n are the rows and columns of the matrix elements respectively; Step 13: Train on the improved YOLO V10 algorithm to obtain a remote sensing image transmission tower detection model; 2) Calculate and analyze the characteristics of typical transmission towers; Specifically as follows: Step 21: Extract the 768-dimensional embedded feature tr of the power transmission tower area marked in the training samples i =(trf i 1, trf i 2,... trf i 768 ); Step 22: Calculate the average value of the features tr of all transmission tower regions in the training samples to obtain the typical transmission tower features i ​ 3) Use an intelligent Agent to detect collapsed transmission towers in each remote sensing image and report to the user in time when a collapsed tower is found; specifically as follows: Step 31: Build an intelligent Agent for detecting the collapse of transmission towers, using a multimodal large model as the base and decision center, and using a persistent database as its memory storage medium; Step 32: Construct a knowledge base for the Agent. The knowledge base is directly provided to the Agent in the form of a txt file. If there are two small images of transmission towers detected by YOLOV10 that contain collapsed towers, set as 1.jpg and 2.jpg, and there are another two small images that do not contain collapsed towers, then the knowledge base is constructed as follows: 'The following two pictures are of collapsed transmission towers: 1.jpg, 2.jpg. The following two pictures are of non-collapsed transmission towers: 3.pg, 4.jpg.'; Step 33: Receive each question from the user, and use the Embedding algorithm to convert the question into a question vector, denoted as vector UQ E =(v i,1 , v i,2 ,......v i,N ,), and compare vector UQ E with the predefined question vector DQ E =(v'1, v'2,......v' N ), where the predefined question is the sentence "Is there a collapsed transmission tower in the following figure?". Calculate the cosine similarity between these two vectors where ||v i || and ||v'|| are the magnitudes of vectors UQ E and DQ E respectively; Step 34, if simi(UQ E , DQ E ) < T, it indicates that the question raised by the user in the current round has nothing to do with the detection of transmission tower collapse, and this question is directly submitted to the multi-modal large model for answering; here, T represents a similarity threshold, taking 0.5; if simi(UQ E , DQ E ) >= T, that is, the user question vector UQE is similar to the preset question vector DQE, indicating that the question asked by the user in this round is related to tower collapse, and continue to the next step of processing; Step 35: When they are similar, call the trained transmission tower detection model to detect all the transmission towers in the current remote sensing image and store them in the directory in the form of a picture file for each detected tower; if there are many picture files, set a limit value F for the total number of files. When the number of pictures is less than this value, process all the pictures, otherwise process F pictures; Step 36: When the number of pictures is greater than F, extract the features of each small picture, and select F small pictures that are least similar to the features of typical transmission towers, that is where K represents the number of transmission towers detected in the current picture; Step 37: After selecting at most F small pictures of towers, construct a question about the collapse of the tower and submit it to the multimodal large model for answering; the constructed question is as follows: "When a transmission tower collapses, it generally lies obliquely on the ground and may also be a bit bent. Carefully observe the pictures in the knowledge base and then answer the following question: Is this picture more similar to a collapsed transmission tower or a non-collapsed transmission tower?" + pic_path, where pic_path is the temporary storage path of the small picture. If there are f (f <= F) small pictures, the above question will be submitted to the intelligent Agent f times; Step 38: When the large model in the intelligent Agent believes that there is a collapsed transmission tower in a certain small picture, synthesize the answers returned by the large model in the intelligent Agent and feedback the synthesized result to the user; if it is judged that there is no collapsed tower, and if the judgment result is connected to an alarm or notification program, the alarm or notification is not triggered.

2. The method for constructing an intelligent agent for detecting collapsed transmission towers in remote sensing images according to claim 1, wherein In Step 32, the number of collapsed tower images included in the knowledge base <= 10 pictures.