A method for intelligent detection of ground objects based on the fusion of reinforcement learning and Transformer features
By combining reinforcement learning with Transformer features in intelligent detection of land objects, the problem of not being able to obtain structured data in the prior art is solved, and a more fine-grained analysis of regional land objects is achieved, and the accuracy and efficiency of transmission line engineering design is improved.
Patent Information
- Application Number
- CN202310686309.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-06-10
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2043-06-10
AI Technical Summary
The prior art cannot obtain structured data in regional land object analysis, and ignores the spatial geographical information of the planned area, resulting in unreasonable transmission line engineering design and affecting the allocation of power supply capacity requirements.
Using intelligent detection method based on reinforcement learning and Transformer features, a three-dimensional landform model is constructed through BIM technology, combining VisionTransformer model and reinforcement learning evaluation network, identify and adjust the landform position to obtain a finer-grained landform position map.
A more fine-grained analysis of the planned area of land objects is achieved, the accuracy and efficiency of transmission line engineering design is improved, the ability to identify various types of land objects is enhanced, and the appearance characteristics of different land objects are adapted to.
Smart Images

Figure CN116863330B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer vision, and particularly to a method for intelligent detection of ground objects based on the fusion of reinforcement learning and Transformer features. Background Art
[0002] The development of society and cities is inseparable from the supply and replenishment of power resources. In order to better meet the sustainable development of the regional social economy, detailed and comprehensive power planning work needs to be carried out in specific regions to ensure the power supply demand in the later stage of regional development and improve the power supply efficiency of regional transmission lines. For this reason, it is necessary to conduct a ground object survey on the target area in advance and analyze the specific ground object conditions of the planning area from the field of geospatial to efficiently and accurately carry out the layout project of transmission line facilities.
[0003] In the current process of regional ground object analysis, a three-dimensional unstructured ground object map is often constructed through BIM technology, and it is impossible to study ground objects by obtaining structured data, ignoring many spatial geographical information of the planning area, resulting in certain irrationality in the design of transmission line projects and affecting the distribution of power supply capacity requirements in the planning area. Summary of the Invention
[0004] The purpose of the present invention is to provide a method for intelligent detection of ground objects based on the fusion of reinforcement learning and Transformer features, which can analyze the specific ground object conditions of the planning area in a finer granularity.
[0005] To achieve the above purpose, the technical solution of the present invention is: a method for intelligent detection of ground objects based on the fusion of reinforcement learning and Transformer features, including the following steps:
[0006] Step S1: Based on BIM technology, with the help of a digital three-dimensional design platform, construct an unstructured BIM three-dimensional ground object model; classify the BIM three-dimensional ground object map, and sort out the ground object type labels to obtain various ground object target data sets;
[0007] Step S2: Construct a Vision Transformer model with feature fusion, separate and identify various objects in the target scene, obtain ground object features, and detect the initial positions of ground objects;
[0008] Step S3: Input the initial positions of ground objects and the ground object features into the reinforcement learning evaluation network, and the reinforcement learning evaluation network outputs a ground object position map with a finer granularity;
[0009] Step S4: Visualize the algorithm results and provide an external usage interface.
[0010] In an embodiment of the present invention, the specific content of step S1 is:
[0011] Step S11: Based on BIM technology, with the help of a digital 3D design platform, draw the BIM 3D topographic map of various topographic scenes;
[0012] Step S12: Classify and organize the drawn BIM 3D topographic map to obtain various topographic target data sets;
[0013] Step S13: Make training labels annotation for each BIM 3D image image in the topographic target data set; image ; Using the manual annotation method, label various objects on the ground, that is, manually draw a detection box, and obtain the upper left point of the detection box and the lower right point of the two-dimensional coordinate values; where object i refers to the i-th ground object in the BIM 3D image image;
[0014] Step S14: Organize the labels to obtain the data set labels required for the training model.
[0015] In an embodiment of the present invention, the specific steps of step S2 are as follows:
[0016] Step S21: Input the image image in the topographic target data set into the VisionTransformer model U-N, and use a feature fusion module to enhance the learning ability of the model; the feature fusion module is composed of three layers of image features in the model: the second layer of feature map feature_image with a size of 1 / 8 2 、the third layer of feature map feature_image with a size of 1 / 16 3 and the last layer of feature map feature_image with a size of 1 / 32 4 ; The feature fusion formula is as follows:
[0017] fusion_feature_image
[0018] = feature_image 2 + upsample(feature_image 3 ) + upsample(feature_image 4 )
[0019] where fusion_feature_image is the final image feature map after fusion, and upsample() is the upsampling operation. Since the sizes of the three feature maps are different, the third layer of feature map feature_image with a size of 1 / 16 3and the feature map feature_image of the last layer with a size of 1 / 32 4 are all upsampled to a size of 1 / 8;
[0020] Step S22: Use the cross-entropy loss function CrossEntropy_Loss to calculate the model loss value Loss between the output result of the feature fusion VisionTransformer model U-N and the image label. The loss function formula is as follows:
[0021]
[0022] In the formula, number is the total number of image pixels, index is the pixel subscript, representing the index-th pixel, prob() represents the event probability, represents the category to which the index-th pixel pixel in the image image in the label belongs, represents the category to which the index-th pixel pixel in the image image output by the model belongs, represents the probability that the corresponding pixel belongs to the label category;
[0023] Since the target ground objects have multiple categories, in order to accurately distinguish various ground objects and reduce the probability that the U-Net model misclassifies one ground object as another, the Focal_Loss loss function needs to be used. The formula is as follows:
[0024]
[0025]
[0026] In the formula, γ is a hyperparameter and is adjusted according to specific circumstances.
[0027] Step S23: After calculating the loss value by the cross-entropy loss function, use the stochastic gradient descent optimization method to use the loss value to train the neural network model until the network model converges to achieve the highest accuracy.
[0028] Step S24: Use the regression head composed of fully connected layers to obtain the target detection map detect_map output by the converged model image , and the target detection map detect_map image is a two-dimensional plane graph, which records the upper left point coordinates and the lower right point coordinates of each ground object detection box.
[0029] In an embodiment of the present invention, the specific content of the step S3 is as follows:
[0030] Step S31: Input the fused final image feature map fusion_feature_image and the target detection map detect_map image into the reinforcement learning evaluation network composed of multi-layer perceptrons. The reinforcement learning evaluation network determines whether the position of each detection box in the target detection map detect_map image needs to be further adjusted based on the detection box coordinate records in it and the fused final image feature map, and outputs an action map action_map for adjusting the target detection map detect_map image ; image image ;
[0031] Step S32: Add the action map action_map image and the target detection map detect_map image two-dimensionally to obtain the final target detection map final_detect_map image , and the formula is as follows:
[0032] final_detect_map image = action_map image + detect_map image
[0033] The final target detection map final_detect_map image records the specific coordinate positions of each ground object detection box, and identifies the ground objects in a more fine-grained manner.
[0034] In an embodiment of the present invention, the step S4 is specifically as follows:
[0035] Step S41: Implement visual graphical display of the ground object detection effect based on Python;
[0036] Step S42: Provide an external interface for processing data for easy use.
[0037] Compared with the prior art, the present invention has the following beneficial effects:
[0038] 1. The present invention applies the advanced Vision Transformer model technology to the analysis of ground objects and conditions. According to the geospatial environment characteristics, the model is used to efficiently and quickly identify various ground objects, and a feature fusion module is added to strengthen the learning degree of the model for ground objects. At the same time, two loss functions are used to train the network model, which efficiently improves the fitting degree and recognition ability of the model for multi-class and multi-feature ground objects, and intelligently adapts to the appearance characteristics of different ground objects in the planning area.
[0039] 2. Traditional methods only implement object detection tasks based on convolutional neural networks or Transformer models. The present invention creatively uses reinforcement learning to perform fine-grained adjustment on the results output by the Transformer model, making up for the limitations of the Transformer model in object detection tasks, strengthening the recognition ability of the network model, and obtaining more accurate object detection results than the Transformer model.
[0040] 3. The present invention graphically displays the planning results and provides an external interface for convenient use. Description of the Drawings
[0041] Figure 1 It is a flowchart of the method of the present invention. Detailed Embodiments
[0042] The following combines the drawings to specifically describe the technical solutions of the present invention.
[0043] Please refer to Figure 1 , the present invention provides a method for intelligent detection of ground objects based on the fusion of reinforcement learning and Transformer features, including the following steps:
[0044] Step S1: Based on BIM technology, with the help of a digital 3D design platform, construct an unstructured BIM 3D ground object model. Classify the BIM 3D ground object map and sort out the ground object type labels to obtain various ground object target datasets;
[0045] Step S2: Construct a Vision Transformer model with feature fusion, separate and identify various objects in the target scene, obtain ground object features, and detect the initial positions of ground objects;
[0046] Step S3: Input the initial positions of ground objects and ground object features into the reinforcement learning module, and the evaluation network outputs a more fine-grained ground object position map;
[0047] Step S4: Visualize the algorithm results and provide an external usage interface.
[0048] In this embodiment, step S1 specifically includes the following steps:
[0049] Step S11: Based on BIM technology, with the help of a digital 3D design platform, draw various ground object scenes, such as farmland, mountains, rivers, roads, etc.;
[0050] Step S12: Classify and sort out the drawn BIM 3D ground object map to obtain various ground object target datasets;
[0051] Step S13: Generate training labels annotation for each BIM three-dimensional image image in the ground object image dataset image Using the manual annotation method, various objects on the ground are annotated. For objects such as farmland, rivers and lakes, mountains, buildings, etc., it is necessary to manually draw detection frames. And obtain the upper left point of the detection frame and the lower right point of the two-dimensional coordinate values. Where object i refers to the i-th ground object in the picture.
[0052] Step S14: Organize the labels to obtain the dataset labels required for the training model.
[0053] In this embodiment, step S2 specifically includes the following steps:
[0054] Step S21: Input the image image in the preprocessed dataset into the VisionTransformer model U-N, and adopt a feature fusion module to enhance the learning ability of the model. The feature fusion module is composed of three layers of image features in the model: the feature map feature_image of size 1 / 8 in the second layer 2 , the feature map feature_image of size 1 / 16 in the third layer 3 and the feature map feature_image of size 1 / 32 in the last layer 4 . The feature fusion formula is as follows:
[0055] fusion_feature_image
[0056] = feature_image 2 + upsample(feature_image 3 ) + upsample(feature_image 4 )
[0057] where fusion_feature_image is the final image feature map after fusion, and upsample() is the upsampling operation (since the sizes of the three feature maps are different, it is necessary to perform upsampling operations on the feature map feature_image of size 1 / 16 in the third layer 3 and the feature map feature_image of size 1 / 32 in the last layer 4 to the size of 1 / 8).
[0058] Step S22: Use the cross - entropy loss function CrossEntropy_Loss to calculate the model loss value Loss between the output result of the feature - fusion VisionTransformer model and the image label. The loss function formula is as follows:
[0059]
[0060] In the formula, number is the total number of image pixels, index is the pixel subscript, representing the index - th pixel, prob() represents the event probability, represents the category to which the index - th pixel pixel in the image image in the label belongs, represents the category to which the index - th pixel pixel in the image image output by the model belongs. Then represents the probability that this pixel belongs to the label category.
[0061] Furthermore, since the target ground objects have multiple categories, in order to accurately distinguish various ground objects and reduce the probability that the U - Net model misjudges one ground object as another, the Focal_Loss loss function needs to be used. The formula is as follows:
[0062]
[0063] In the formula, γ is a hyper - parameter and is adjusted according to specific circumstances.
[0064] Step S23: After calculating the loss value by the cross - entropy loss function, use the stochastic gradient descent optimization method to use the loss value to train the neural network model until the network model converges to achieve the highest accuracy.
[0065] Step S24: Use the regression head composed of fully - connected layers to obtain the target detection map detect_map output by the converged model image , the target detection map detect_map image is a two - dimensional plane graph, which records the upper - left point coordinates and lower - right point coordinates of each ground object detection box.
[0066] In this embodiment, step S3 specifically includes the following steps:
[0067] Step S31: Input the final feature fusion_feature_image of the original image and the target detection map detect_map image into the reinforcement learning evaluation network composed of multi - layer perceptrons. The evaluation network determines the target detection map detect_map according to the detection box coordinate records in the target detection map detect_map image and the final features of the original image.image Whether the position of each detection box in image needs to be further adjusted, and an image detect_map for adjusting the object detection is output image ;
[0068] Step S32: Add the action map image and the object detection map detect_map image two-dimensionally to obtain the final object detection map final_detect_map image , and the formula is as follows:
[0069] final_detect_map image = action_map image + detect_map image
[0070] The final object detection map final_detect_map image records the specific coordinate positions of each ground object detection box, and identifies the ground objects in a more fine-grained manner.
[0071] In this embodiment, step S4 specifically includes the following steps:
[0072] Step S41: Implement visual graphical display of the ground object detection effect based on Python;
[0073] Step S42: Provide an external interface for processing data for easy use.
[0074] The above are the preferred embodiments of the present invention. All changes made according to the technical solutions of the present invention, when the functions and effects generated do not exceed the scope of the technical solutions of the present invention, fall within the protection scope of the present invention.
Claims
1. A method for intelligent detection of ground objects based on the fusion of reinforcement learning and Transformer features, characterized in that, it includes the following steps: Step S1: Based on BIM technology, with the help of a digital 3D design platform, construct an unstructured BIM 3D ground object model; classify the BIM 3D ground object map, and sort out the ground object type labels to obtain various ground object target datasets; Step S2: Construct a VisionTransformer model with feature fusion, separate and identify various objects in the target scene, obtain ground object features, and detect the initial position of the ground object; the specific implementation is as follows: Step S21: Input the image in the ground object target dataset into the VisionTransformer model U-N, and use the feature fusion module to enhance the learning ability of the model; the feature fusion module is formed by fusing the image features of three layers in the model: the feature map feature_image with a size of 1 / 8 in the second layer 2 , the feature map feature_image with a size of 1 / 16 in the third layer 3 , and the feature map feature_image with a size of 1 / 32 in the last layer 4 ; the feature fusion formula is as follows: fusion_feature_image =feature_image 2 +upsample(feature_image 3 ) +upsample(feature_image 4 ) Among them, fusion_feature_image is the final image feature map after fusion, and upsample() is the upsampling operation. Since the sizes of the three feature maps are different, the feature map feature_image of the third layer with a size of 1 / 16 and 3 the feature map feature_image of the last layer with a size of 1 / 32 4 both need to be upsampled to a size of 1 / 8; Step S22: Use the cross-entropy loss function CrossEntropy_Loss to calculate the model loss value Loss between the output result of the feature fusion VisionTransformer model U-N and the image label. The loss function formula is as follows: where number is the total number of image pixels, index is the pixel subscript, representing the index-th pixel, prob() represents the event probability, representing the category to which the index-th pixel pixel in the image image in the label belongs, representing the category to which the index-th pixel pixel in the image image output by the model belongs, representing the probability that the corresponding pixel belongs to the label category; Since there are multiple categories of target ground objects, in order to accurately distinguish various ground objects and reduce the probability that the U-Net model misclassifies one ground object as another, the Focal_Loss loss function needs to be used. The formula is as follows: In the formula, γ belongs to a hyperparameter and is adjusted according to specific circumstances; Step S23: After calculating the loss value by the cross-entropy loss function, use the stochastic gradient descent optimization method to use the loss value to train the neural network model until the network model converges to achieve the highest accuracy; Step S24. Use a regression head composed of fully connected layers to obtain the target detection map detect_map output by the converged model image , the target detection map detect_map image is a two-dimensional planar graph, which records the coordinates of the upper left point and the lower right point of each ground object detection box; Step S3: Input the initial position of the ground object and the ground object features into the reinforcement learning evaluation network, and the reinforcement learning evaluation network outputs a more fine-grained ground object position map; Step S4: Visualize the algorithm results and provide an external usage interface.
2. A method for intelligent detection of ground objects based on the fusion of reinforcement learning and Transformer features according to claim 1, characterized in that, the specific content of step S1 is: Step S11: Based on BIM technology, with the help of a digital 3D design platform, draw the BIM 3D ground object map of various ground object scenes; Step S12: Classify and sort out the drawn BIM 3D ground object map to obtain various ground object target datasets; Step S13: Generate training labels annotation for each BIM three-dimensional image image in the ground object dataset image ; Using the manual annotation method, annotate various objects on the ground, that is, manually draw a detection box and obtain the two-dimensional coordinate values of the upper left point and the lower right point ; where object i refers to the i-th ground object in the BIM three-dimensional image image; Step S14: Sort out the labels to obtain the dataset labels required for training the model.
3. A method for intelligent detection of ground objects based on the fusion of reinforcement learning and Transformer features according to claim 1, characterized in that, the specific content of step S3 is: Step S31: Input the fused final image feature map fusion_feature_image and the target detection map detect_map image into the reinforcement learning evaluation network composed of a multi-layer perceptron. The reinforcement learning evaluation network determines whether the position of each detection box in the target detection map detect_map image needs to be further adjusted based on the detection box coordinate records in it and the fused final image feature map, and outputs an action map action_map for adjusting the target detection map detect_map image ; image image ; Step S32: Add the action map action_map image and the target detection map detect_map image by two-dimensional value addition to obtain the final target detection map final_detect_map image , and the formula is as follows: final_detect_map image = action_map image + detect_map image Final target detection map final_detect_map image Records the specific coordinate positions of each ground object detection box, and identifies ground objects in a more fine-grained manner.
4. A method for intelligent detection of ground objects based on the fusion of reinforcement learning and Transformer features according to claim 1, characterized in that, the specific content of step S4 is: Step S41: Based on Python, implement visual graphical display of the ground object detection effect; Step S42: Provide an external interface for processing data for easy use.
Citation Information
Patent Citations
Remote sensing ground object classification post-processing method based on iterative superpixel segmentation
CN111553222A
Three-dimensional space molecule generation method and device based on multi-task pre-training inverse reinforcement learning
CN115831261A