Power transmission channel image ground object intelligent segmentation method and device
By annotating and rasterizing the image data of power transmission channels and performing sample enhancement processing, combined with a semantic segmentation model that is dynamically optimized in multiple rounds, the efficiency and accuracy problems in ground feature segmentation of power transmission channel images have been solved. This has enabled high-precision automated segmentation and vector result optimization, which is suitable for power transmission channel operation and maintenance.
Patent Information
- Application Number
- CN202511784891.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-01
- Publication Date
- 2026-02-27
- Estimated Expiration
- 2045-12-01
AI Technical Summary
Existing technologies suffer from low efficiency and low accuracy in the segmentation of ground features in power transmission channel images. In particular, the processing of small fragmented areas caused by high-resolution image data is complicated, and multispectral image classification methods are not applicable, resulting in poor engineering application effects.
A semantic segmentation model employs labeled rasterization, sample augmentation, and data slicing, combined with a sliding window transformer backbone network, feature pyramid, and dual attention mechanism. The model optimizes the feature segmentation kernel parameters through a dynamic kernel generator, outputs high-precision segmentation mask data, and performs raster vectorization transformation and topology optimization to form continuous vector surface results.
It has achieved automated and high-precision segmentation of ground features in power transmission channel images, improving segmentation efficiency and accuracy, solving the problem of small fragmented surfaces, and is applicable to operation and maintenance tasks such as power transmission channel hazard analysis and floating object monitoring.
Smart Images

Figure CN121259459B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of power transmission channel image processing, in particular to a power transmission channel image feature intelligent segmentation method and device. BACKGROUND
[0002] This section is intended to provide background or context to the inventive embodiments recited in the claims. The description herein does not constitute admission that the background art is prior art nor does it constitute an indication that the background art is relevant to determining inventiveness.
[0003] In power grid inspection operations, the power transmission channel orthophoto data (DOM) collected by the unmanned aerial vehicle is mainly used for three-dimensional digital twin scene construction, and the traditional feature segmentation relies on manual interpretation, which is inefficient.
[0004] The prior art adopts a remote sensing image land use classification method, including a supervised classification algorithm, an object-oriented classification, a random forest machine learning, and a CNN / U-Net deep learning algorithm, and uses software tools such as eCognition and GEE to perform feature segmentation. However, these methods have significant defects: the spatial resolution of the power transmission channel orthophoto is as high as 0.1 meters, resulting in a large number of small fragments in the segmentation results that require complex manual post-processing; and because the data is only in RGB three bands, the efficient classification method of multi-spectral images cannot be applied, and the existing scheme has poor effect in engineering application. SUMMARY
[0005] The embodiments of the present application provide a power transmission channel image feature intelligent segmentation method to improve the feature segmentation efficiency, precision and accuracy of the power transmission channel image, which comprises:
[0006] The power transmission channel digital orthophoto data and the corresponding vector label are subjected to label rasterization processing, label automatic verification processing and sample enhancement processing to obtain the power transmission channel digital orthophoto data and the corresponding label mask data;
[0007] The power transmission channel digital orthophoto data and the corresponding label mask data are subjected to data slicing processing to obtain a power transmission channel feature segmentation model sample set;
[0008] Based on the power transmission channel feature segmentation model sample set, a multi-round dynamic optimization strategy is adopted to train a pre-constructed power transmission channel feature semantic segmentation model network structure to obtain a trained power transmission channel feature segmentation model; the power transmission channel feature semantic segmentation model network structure adopts a sliding window transformer backbone network to extract multi-scale power transmission channel feature characteristics, optimizes semantic expression by fusing a feature pyramid and a double attention mechanism, and uses a dynamic kernel generator to adaptively generate feature segmentation kernel parameters; after iterative updating by a dynamic kernel optimizer, the power transmission channel digital orthophoto data corresponding to the segmentation mask data is output by a segmentation head;
[0009] The vector data integration module is used for generating target segmentation mask data corresponding to target power transmission channel digital orthographic image data based on the power transmission channel ground object segmentation model, performing raster vectorization conversion on the target segmentation mask data to obtain corresponding frame vector data, and integrating the frame vector data by using a topological optimization algorithm to obtain a continuous vector surface result of the entire power transmission channel.
[0010] The embodiment of the present application also provides a power transmission channel image ground object intelligent segmentation device for improving the ground object segmentation efficiency, precision and accuracy of the power transmission channel image, and the device comprises:
[0011] The annotation mask data generation module is used for performing annotation rasterization processing, annotation automatic verification processing and sample enhancement processing on the power transmission channel digital orthographic image data and corresponding vector annotation to obtain the power transmission channel digital orthographic image data and corresponding annotation mask data.
[0012] The sample set generation module is used for performing data slicing processing on the power transmission channel digital orthographic image data and corresponding annotation mask data to obtain a power transmission channel ground object segmentation model sample set.
[0013] The power transmission channel ground object segmentation model output module is used for training a pre-constructed power transmission channel ground object semantic segmentation model network structure by adopting a multi-round dynamic optimization strategy based on the power transmission channel ground object segmentation model sample set to obtain a trained power transmission channel ground object segmentation model, wherein the power transmission channel ground object semantic segmentation model network structure adopts a sliding window transformer backbone network to extract multi-scale power transmission channel ground object features, optimizes semantic expression by fusing a feature pyramid and a double attention mechanism, and adaptively generates ground object segmentation kernel parameters by using a dynamic kernel generator; after iterative updating by a dynamic kernel optimizer, the power transmission channel digital orthographic image data corresponding segmentation mask data is output by a segmentation head.
[0014] The vector data integration module is used for generating target segmentation mask data corresponding to target power transmission channel digital orthographic image data based on the power transmission channel ground object segmentation model, performing raster vectorization conversion on the target segmentation mask data to obtain corresponding frame vector data, and integrating the frame vector data by using a topological optimization algorithm to obtain a continuous vector surface result of the entire power transmission channel.
[0015] The embodiment of the present application also provides a computer device, which comprises a memory, a processor and a computer program stored on the memory and capable of running on the processor, and the processor implements the power transmission channel image ground object intelligent segmentation method when executing the computer program.
[0016] The embodiment of the present application also provides a computer readable storage medium, which stores a computer program, and the computer program is executed by a processor to implement the power transmission channel image ground object intelligent segmentation method.
[0017] The embodiment of the present application also provides a computer program product, which comprises a computer program, and the computer program realizes the power transmission channel image feature intelligent segmentation method when executed by a processor.
[0018] The embodiment of the present application solves the problem of low efficiency of traditional artificial domestication sample construction and insufficient model generalization ability caused by uneven sample distribution by performing annotation rasterization processing on power transmission channel digital orthographic image data and vector annotation, combining with automatic coordinate verification to ensure data space alignment, and sample enhancement processing; when training a dedicated segmentation model based on the sample library, a sliding window transformer backbone network is used to extract multi-scale features in layers, a feature pyramid network and a channel-spatial dual attention mechanism are fused to accurately optimize semantic expression, a dynamic kernel generator is used to adaptively generate feature-specific segmentation kernel parameters, and after iteration and update by a dynamic kernel optimizer, a high-precision segmentation mask is output, which can improve the power transmission channel feature intelligent segmentation capability and enhance the edge segmentation precision of typical features, and solves the adaptability defects of traditional technologies in the power transmission channel scene; the target segmentation mask data is converted into raster and vector, and integrated by using a topological optimization algorithm, and the continuous vector surface results formed can be used for power transmission channel hazard analysis, floating object monitoring and other operation and maintenance businesses, solving the problem of small broken surfaces caused by high-resolution images in the prior art, thereby realizing automatic high-precision segmentation of power transmission channel features and intelligent optimization of vector results, and improving the feature segmentation efficiency, precision and accuracy of power transmission channel images. BRIEF DESCRIPTION OF DRAWINGS
[0019] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings needed to be used in the embodiments or prior art description. Obviously, the drawings in the following description are only some embodiments of the present application, and for those skilled in the art, other drawings can also be obtained without creative labor. In the drawings:
[0020] Figure 1 It is a flowchart of a power transmission channel image feature intelligent segmentation method in an embodiment of the present application;
[0021] Figure 2 It is a specific example diagram of a power transmission channel feature segmentation training sample library construction technology flow in an embodiment of the present application;
[0022] Figure 3 It is a specific example diagram of a power transmission channel feature segmentation model structure design and model training technology flow in an embodiment of the present application;
[0023] Figure 4 It is a specific example diagram of a model inference post-processing technology flow in an embodiment of the present application;
[0024] Figure 5 A specific schematic diagram of a power transmission channel ground object segmentation model network structure in an embodiment of the present application is shown in the figure.
[0025] Figure 6 A structural example diagram of a power transmission channel image ground object intelligent segmentation device in an embodiment of the present application is shown in the figure.
[0026] Figure 7 A computer device schematic diagram provided in an embodiment of the present application is shown in the figure. DETAILED DESCRIPTION
[0027] To make the purpose, technical scheme and advantages of the embodiments of the present application clearer and more apparent, the embodiments of the present application are further described in detail below with reference to the accompanying drawings. Herein, the schematic embodiments of the present application and the descriptions thereof are used to explain the present application, but not as a limitation of the present application.
[0028] The term “and / or” in this document is merely used to describe an associated relationship, which means that there can be three relationships, for example, A and / or B can mean that A exists alone, A and B exist together, and B exists alone. In addition, the term “at least one” in this document means any one of a plurality of combinations or any combination of at least two of a plurality of combinations, for example, including at least one of A, B and C can mean including any one or more elements selected from the set consisting of A, B and C.
[0029] In the description of the present application, “including”, “containing”, “having”, “comprising” and the like are all open terms, which means including but not limited to. The description of the terms “one embodiment”, “one specific embodiment”, “some embodiments”, “for example” and the like means that the specific features, structures or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present application. In the present specification, the schematic description of the above terms does not necessarily mean the same embodiment or example. Moreover, the specific features, structures or characteristics described can be combined in any one or more embodiments or examples in a suitable manner. The order of steps involved in each embodiment is used to illustrate the implementation of the present application, and the order of steps is not limited, which can be adjusted as needed.
[0030] The technical terms involved in the present application are explained as follows:
[0031] Swin-Transformer: Sliding Window Transformer; FNP: Feature Pyramid Network; DOM: Digital Orthophoto Map; IoU: Intersection over Union; Backbone: Backbone Network; Neck: Feature Fusion Neck; Kernel: Dynamic Kernel Module; Head: Segmentation Head; TIFF / TIF: Tagged Image File Format; SHP: Shapefile; RGBA: Red Green Blue Alpha; CNN: Convolutional Neural Network; U-Net: U-shaped Network.
[0032] Specifically, with the wide popularity of unmanned aerial vehicle technology, a large amount of power transmission line laser point cloud data, tower and line detail photo data, and power transmission channel orthophoto data will be collected in power grid inspection operation. The point cloud data and tower photo data are the main application data in the current power grid inspection business, and the channel orthophoto data is more applied to the twin base of the power grid three-dimensional digital twin scene. In recent years, with the development of artificial intelligence technology and the continuous extension of power grid inspection business, the application direction of power transmission channel orthophoto data gradually increases, among which the identification and segmentation of typical features of power transmission channel are the main business application direction. At present, in the power transmission channel image feature segmentation data processing operation, the artificial interpretation method is mostly used, which is labor-intensive and low in efficiency, and an intelligent power transmission channel image feature segmentation method is urgently needed to quickly improve the efficiency of power transmission channel image feature segmentation operation and provide reference data for power transmission channel operation and maintenance. At present, in the power transmission channel image feature segmentation data processing operation, the remote sensing image land use classification method is mostly used for processing, such as traditional supervised classification algorithm, object-oriented classification method, random forest machine learning algorithm, CNN / U-Net deep learning algorithm, and multi-source data fusion classification algorithm. The open source or commercial software derived from these algorithms is also commonly used in power transmission channel image feature identification and segmentation operation. The existing image feature segmentation technology has good effect in remote sensing image processing, and the power transmission channel orthophoto data mentioned in the present invention has high spatial resolution, and there are many small fragments in the identification and segmentation process, which is complex for subsequent manual processing. At the same time, since the power transmission channel digital orthophoto data is RGB three-band image data, some high-quality multi-spectral image data feature identification and segmentation methods cannot be applied to the power transmission channel orthophoto feature identification and segmentation processing operation. At present, there is no perfect scheme for power transmission channel orthophoto feature segmentation task, and the effect of application in actual engineering is poor.
[0033] To solve the above problems, the embodiment of the present application provides a power transmission channel image feature intelligent segmentation method, which solves the problem of low precision in power transmission channel image feature segmentation operation by using traditional remote sensing image feature classification method, and constructs a perfect power transmission channel image feature intelligent segmentation process to improve the feature segmentation efficiency, precision and accuracy of power transmission channel image, see Figure 1 The method can include:
[0034] Step 101: performing label rasterization processing, label automatic verification processing and sample enhancement processing on the power transmission channel digital orthophoto data and the corresponding vector label to obtain the power transmission channel digital orthophoto data and the corresponding label mask data;
[0035] Step 102: data slicing processing is performed on the power transmission channel digital orthographic image data and the corresponding label mask data, so as to obtain a power transmission channel ground object segmentation model sample set;
[0036] Step 103: based on the power transmission channel ground object segmentation model sample set, a multi-round dynamic optimization strategy is adopted to train a pre-constructed power transmission channel ground object semantic segmentation model network structure, so as to obtain a trained power transmission channel ground object segmentation model; the power transmission channel ground object semantic segmentation model network structure adopts a sliding window transformer backbone network to extract multi-scale power transmission channel ground object features, optimizes semantic expression through fusion of a feature pyramid and a double attention mechanism, and adaptively generates ground object segmentation kernel parameters by using a dynamic kernel generator; after iterative updating by a dynamic kernel optimizer, segmentation mask data corresponding to the power transmission channel digital orthographic image data is output by a segmentation head;
[0037] Step 104: target segmentation mask data corresponding to target power transmission channel digital orthographic image data is generated based on the power transmission channel ground object segmentation model; raster vectorization conversion is performed on the target segmentation mask data, so as to obtain corresponding divided vector data; topological optimization algorithm is adopted to integrate the divided vector data, so as to obtain continuous vector surface results of the entire power transmission channel.
[0038] The embodiment of the present application solves the problems of low efficiency of traditional artificial domestication sample construction and insufficient model generalization ability caused by uneven sample distribution by performing label rasterization processing on power transmission channel digital orthographic image data and vector labeling, combining with automatic coordinate verification to ensure data space alignment, and sample enhancement processing; when training a dedicated segmentation model based on the sample library, a sliding window transformer backbone network is used to extract multi-scale features in layers, semantic expression is accurately optimized through fusion of a feature pyramid network and a channel space double attention mechanism, ground object dedicated segmentation kernel parameters are adaptively generated by using a dynamic kernel generator, and high-precision segmentation masks are output after iterative updating by a dynamic kernel optimizer, which can improve the intelligent segmentation ability of power transmission channel ground objects and improve the edge segmentation precision of typical ground objects, and solve the adaptability defects of traditional technologies in the power transmission channel scene; raster vectorization conversion is performed on the target segmentation mask data, and topological optimization algorithm is adopted for integration, so that continuous vector surface results can be used for power transmission channel hazard analysis, floating object monitoring and other operation and maintenance businesses, the problem of small broken surfaces caused by elimination of high-resolution images is solved, so that automatic high-precision segmentation of power transmission channel ground objects and intelligent optimization of vector results are realized, and the ground object segmentation efficiency, precision and accuracy of power transmission channel images are improved.
[0039] In specific implementation, first, step 101: label rasterization processing, label automatic verification processing and sample enhancement processing are performed on the power transmission channel digital orthographic image data and the corresponding vector labeling, so as to obtain the power transmission channel digital orthographic image data and the corresponding label mask data.
[0040] In one embodiment, the annotation rasterization processing is performed on the power transmission channel digital orthographic image data and the corresponding vector annotation, including: performing automatic coordinate reference verification on the power transmission channel digital orthographic image data and the vector annotation data; after determining that the coordinate systems of the power transmission channel digital orthographic image data and the vector annotation data are consistent through automatic coordinate reference verification, obtaining feature field information in the vector annotation data; mapping the feature field information into RGB three-band color values; and converting the mapped feature field information into three-band raster annotation mask data based on the digital orthographic image data.
[0041] In the above embodiment, when performing annotation rasterization processing on the power transmission channel digital orthographic image data and the corresponding vector annotation, first, an automatic coordinate reference verification operation is performed. By reading the coordinate system definition information of the power transmission channel digital orthographic image data and the vector annotation data, automatic comparison and analysis is performed. If it is found that the coordinate systems are inconsistent, a coordinate system conversion operation is performed to completely align the spatial reference systems of the two. Then, the feature category field information stored in the vector annotation data is read, including eighteen typical power transmission channel features, such as cultivated land, forest land, scattered trees, shrubs and grasses, grassland, road, highway, railway, bare land surface, hardened ground, water system, building, construction operation area, plastic greenhouse, mulch, dust screen, power transmission tower, and others.
[0042] Each of the above features is mapped to the corresponding RGB three-band color value according to a preset rule. Based on the spatial range and pixel size of the power transmission channel digital orthographic image data, a vector data rasterization processing algorithm is used to convert the mapped vector annotation data into three-band raster annotation mask data that is pixel-level aligned with the digital orthographic image data, thereby realizing accurate spatial matching of the annotation information and the image data.
[0043] In one embodiment, annotation automatic verification processing is performed on the power transmission channel digital orthographic image data and the corresponding vector annotation, including: adding an Alpha band to the digital orthographic image data and the three-band raster annotation mask data to generate four-band data; removing data in the four-band data that exceeds a set intersection over union threshold based on the intersection over union of the four-band data, to obtain annotation automatic verification processing data.
[0044] In the above embodiment, an additional Alpha band is first added to the input digital orthographic image data and the corresponding three-band raster annotation mask data to generate a data format containing RGBA four bands. This operation expands the data dimension to support subsequent invalid area analysis.
[0045] Subsequently, an intersection over union value is calculated based on the generated RGBA four-band data, and a predefined intersection over union threshold is set as a verification standard. When the calculated intersection over union value exceeds the threshold, it is determined that the segmentation data is invalid data, and a rejection operation is performed. This process ensures that only valid data is retained for subsequent processing, improving the overall quality of the training data.
[0046] In one embodiment, sample enhancement processing is performed on the power transmission channel digital orthographic image data and the corresponding vector annotation, including: determining the proportion of pixels of different categories in the annotation automation verification processing data; performing geometric transformation of rotation and cropping, and radiation transformation processing of brightness and contrast disturbance on the category pixels whose proportion is less than a preset proportion threshold, to obtain the power transmission channel digital orthographic image data and the corresponding annotation mask data.
[0047] In the above embodiment, first, the pixel quantity of all ground object categories is counted based on the annotation mask data, and the proportion value of each category pixel in the total pixel quantity is calculated. By quantitatively analyzing the distribution of the proportions of different category pixels, the few-sample categories whose pixel proportions are lower than a preset threshold are identified, such as typical small-sample objects in the floating objects, such as plastic greenhouses, mulching films, and dust screens.
[0048] For these identified few-sample categories, geometric transformation processing of rotation and cropping operations is performed, and radiation transformation processing of brightness adjustment and contrast adjustment is applied.
[0049] The combination of the above-mentioned geometric transformation and radiation transformation is applied to the digital orthographic image data and its corresponding annotation mask data synchronously, ensuring the spatial alignment consistency of the enhanced data. Finally, by re-counting the distribution of the proportions of each category pixel in the enhanced data, the power transmission channel digital orthographic image data and the corresponding annotation mask data set with significantly improved balance are generated.
[0050] In specific implementation, after step 101: performing annotation rasterization processing, annotation automation verification processing, and sample enhancement processing on the power transmission channel digital orthographic image data and the corresponding vector annotation to obtain the power transmission channel digital orthographic image data and the corresponding annotation mask data, step 102: performing data slicing processing on the power transmission channel digital orthographic image data and the corresponding annotation mask data is performed to obtain a power transmission channel ground object segmentation model sample set.
[0051] In the embodiment, the power transmission channel digital orthographic image data and the corresponding label mask data are subjected to data slicing processing to obtain a power transmission channel ground object segmentation model sample set, including: performing secondary division based on the pixel size of the power transmission channel digital orthographic image data and the corresponding label mask data; generating sample slices of equal size through cutting operation on the divided data; and distributing the cut sample slices according to a predetermined proportion to generate a training set, a verification set and a test set of the power transmission channel ground object segmentation model.
[0052] In the above embodiment, the secondary division operation is first performed based on the pixel size specification of the original data.
[0053] The slicing algorithm with a preset size is used to automatically cut the complete divided image data and the corresponding mask data, and through the block-by-block pixel scanning and boundary recognition technology, the large-scale divided data is converted into a plurality of sample slice units of uniform size. The cutting process strictly maintains the spatial synchronization of the digital orthographic image data and the label mask data, and ensures that the image and the label in each sample slice unit are pixel-level aligned.
[0054] The equal-size sample slice units generated by cutting are subjected to standardized grouping processing. The sample slices are divided into three independent data sets according to a predefined allocation proportion: the training set is used for learning and weight adjustment of model parameters, the verification set is used for model performance evaluation and hyperparameter optimization during the training process, and the test set is used for independent performance verification after the model training is completed. Through the proportion allocation mechanism, a complete power transmission channel ground object segmentation model sample set is constructed to provide standardized data support for subsequent model training.
[0055] During the sample set construction process, all slice units maintain the spectral characteristics and spatial accuracy of the original data. The training set drives the model feature extraction ability through a large number of sample slices, the verification set continuously feeds back the confusion matrix and recall rate index to guide the model optimization direction, and the test set finally ensures that the model generalization performance meets the engineering application standard. The entire process realizes the automatic conversion from the original divided data to the standardized training resources, eliminating the efficiency bottleneck caused by manual intervention.
[0056] In actual implementation, after the step 102 of performing data slicing processing on the power transmission channel digital orthographic image data and the corresponding label mask data to obtain a power transmission channel ground object segmentation model sample set, the step 103 of training a pre-constructed power transmission channel ground object semantic segmentation model network structure based on the power transmission channel ground object segmentation model sample set is performed to obtain a trained power transmission channel ground object segmentation model. The power transmission channel ground object semantic segmentation model network structure adopts a sliding window transformer backbone network to extract multi-scale power transmission channel ground object features, optimizes semantic expression by fusing a feature pyramid and a double attention mechanism, and adaptively generates ground object segmentation kernel parameters by using a dynamic kernel generator. After iterative updating by a dynamic kernel optimizer, the segmentation mask data corresponding to the power transmission channel digital orthographic image data is output by a segmentation head.
[0057] In one embodiment, the step of training a pre-constructed power transmission channel ground object semantic segmentation model network structure based on the power transmission channel ground object segmentation model sample set to obtain a trained power transmission channel ground object segmentation model includes the following steps: presetting initial learning rate and loss function weight parameters according to the proportion of each type of sample grid in the sample set; performing multi-round iterative training on the power transmission channel ground object semantic segmentation model network structure based on the sample set; calculating the confusion matrix and recall rate index of the current round after each round of training is completed by using the validation set data; dynamically adjusting the loss function weight and learning rate parameters according to the calculation result; and continuously performing the above multi-round iterative training and loss function weight and learning rate parameter adjustment operations to obtain the trained power transmission channel ground object segmentation model.
[0058] In the above embodiment, the initial learning rate parameter and the loss function weight parameter are preset according to the distribution of the proportion of each type of sample grid in the sample set.
[0059] The multi-round iterative training process of the model is started based on the training set data in the sample set. In each round of training process, multi-scale ground object features are extracted by the sliding window transformer backbone network, different level feature maps are fused by the feature pyramid network, and double optimization processing of channel attention and spatial attention is applied, adaptive ground object segmentation kernel parameters are generated by the dynamic kernel generator in real time, and the prediction result is output by the segmentation head after iterative updating by the dynamic kernel optimizer through back propagation.
[0060] After each round of training is completed, the confusion matrix and the recall rate index of the output of the current model are calculated by using the validation set data. Based on the analysis result of the confusion matrix, the ground object segmentation accuracy defects of specific categories are identified, such as the identification deficiency problem of few sample categories such as floating objects and scattered trees.
[0061] According to the change trend of the recall rate index and the confusion matrix analysis conclusion, the weight parameters of each category and the overall learning rate parameter value in the loss function are dynamically adjusted. This adjustment process specifically strengthens the loss weight of the few-sample category, while optimizing the model convergence speed.
[0062] The above-mentioned multi-round iteration training and parameter dynamic adjustment process is continuously executed in a loop, the model weight state is saved every round, and the verification index change is recorded. When the recall rate index of the verification set remains stable for multiple rounds and the confusion matrix shows that the accuracy of each category is balanced, it is determined that the model training has reached the optimal state.
[0063] Finally, the trained power transmission channel feature segmentation model weight file is output, and the model has high-precision segmentation capability for eighteen types of power transmission channel features such as farmland, forest land, and buildings.
[0064] In one embodiment, it also includes: constructing the power transmission channel feature semantic segmentation model network structure in the following way: extracting hierarchical features through the sliding window transformer backbone network; capturing multi-scale power transmission channel feature of the power transmission channel feature through window attention mechanism at different hierarchical levels; using a feature fusion layer that combines a feature pyramid network and a double attention mechanism to optimize the semantic expression of the network structure; generating feature adaptive segmentation kernel parameters according to the multi-scale power transmission channel feature through a dynamic kernel generator; iteratively updating the segmentation kernel parameters through back propagation via a dynamic kernel optimizer; performing a differentiable convolution operation by a dynamic kernel segmentation head combined with the segmentation kernel parameters to output segmentation mask data in the form of a split frame that is spatially aligned with the input power transmission channel digital orthographic image data.
[0065] In the above embodiment, first, the hierarchical feature extraction operation is implemented through the sliding window transformer backbone network. The backbone network uses a hierarchical window attention mechanism to process the input data at each level. At each level, the local region features are calculated through window self-attention, and then the global feature fusion between different local regions is realized through cross-window connection. This process effectively captures the texture details and shape contour information of vegetation, buildings, roads, and other targets in the power transmission channel, forming multi-scale power transmission channel feature expression. The feature fusion layer that combines a feature pyramid network and a double attention mechanism is used to optimize the semantic expression. The feature pyramid network performs cross-level fusion processing on the multi-scale features output by the backbone network, enhancing the ability to capture high-level semantic information of the target feature. At the same time, channel attention and spatial attention calculations are performed in parallel: the channel attention module optimizes the weight distribution of different feature channels, strengthening the key spectral features; the spatial attention module enhances the spatial position sensitivity of the feature map, improving the positioning accuracy of key locations such as tower edges and road directions. The synergistic effect of the two significantly optimizes the model's semantic understanding ability of complex features.
[0066] The dynamic kernel generator generates adaptive ground feature segmentation kernel parameters in real time based on the input multi-scale feature map. According to the morphological characteristics of different power transmission channel ground objects, the generator dynamically creates exclusive segmentation kernel parameters for cultivated land, forest land, and easily floating objects. The generated parameters are then input into the dynamic kernel optimizer, which iteratively updates the parameters using the backpropagation algorithm and continuously optimizes the discriminative ability of the segmentation kernel using the gradient descent mechanism, effectively solving the class recognition bias problem caused by sample imbalance. Finally, the dynamic kernel segmentation head loads the optimized segmentation kernel parameters to perform differentiable convolution operations. The segmentation head performs convolution calculation on the dynamic kernel parameters and the feature map to generate pixel-level semantic segmentation results. Through real-time adaptive adjustment of the dynamic kernel parameters, the model can accurately segment ground objects of different scales and morphologies in the power transmission channel, and output segmentation mask data in the form of sub-frames that are spatially aligned with the input digital orthographic image data. The entire network structure forms an end-to-end processing chain of "feature extraction-semantic optimization-dynamic kernel generation-parameter iteration-segmentation output". The hierarchical feature extraction mechanism solves the problem of insufficient receptive field of traditional CNN models, the feature pyramid and double attention fusion improve the multi-scale target recognition ability, and the dynamic kernel technology breaks through the generalization limit of fixed convolution kernel, enabling the model to achieve accurate segmentation on 0.1-meter high-resolution RGB images.
[0067] For example: Take a specific example of the workflow of a power transmission channel ground object segmentation model network:
[0068] 1. In the backbone network part, after inputting the original power transmission channel remote sensing image, perform block embedding operation to segment the image into non-overlapping image blocks, each image block is linearly mapped to a fixed-dimensional feature vector and the spatial position information is preserved. Enter the transformer module, divide the feature map into multiple local windows through the hierarchical window attention mechanism, calculate the self-attention in each window to capture the local feature relationship. After completing the first stage calculation, perform block merging operation to merge adjacent feature blocks into one, and through dimension compression, improve the semantic level of the feature map while reducing the spatial resolution. Repeat the above module calculation and block merging process, gradually extract the multi-scale features of the power transmission channel ground object, and implement global feature fusion through cross-window connection at each stage to effectively capture the texture and shape information of the ground object target. Finally, output multiple feature maps of different scales.
[0069] 2、In the neck part, the multi-scale feature maps output by the backbone network are received, and cross-level fusion is performed through a feature fusion path. Doubling upsampling is first performed on the high-level feature maps to match their size with that of the next-high-level feature maps, and then feature integration is performed through cascaded convolution. This process is repeated until fusion with the bottom-level feature maps is completed. The feature maps are respectively sent to a channel attention module and a spatial attention module for parallel calculation. The channel attention module learns the importance weights of different feature channels through global average pooling and a multilayer perceptron, and the spatial attention module generates a spatial attention map through a convolution of a specific size and an activation function. The output features of the two attention modules are weighted and fused to optimize the feature channel weights, improve the spatial position sensitivity, and enhance the semantic expression and positioning accuracy of the target.
[0070] 3、In the dynamic kernel module part, the basic information required for initial kernel update is obtained through convolution, mask prediction and other operations. Based on the basic information, the first kernel update is performed to generate preliminary dynamic kernel related parameters. The dynamic kernel is generated by using category prediction and mask prediction, and the iteration process is repeated multiple times. In each iteration, the dynamic kernel is updated and optimized based on the current features and prediction results, and the kernel parameters are constantly adjusted to better adapt to the morphological diversity of the power transmission channel features. By dynamically adjusting the parameter proportion of different categories of kernels, the class imbalance problem is solved, and the discrimination ability of the kernel is improved.
[0071] 4、In the segmentation head part, the dynamic kernel parameters and related features output by the dynamic kernel module are received, and the dynamic kernel parameters are loaded into the dynamic kernel segmentation head. Convolution kernel and feature map are performed, and each pixel of the feature map is processed to output the pixel-level semantic segmentation result. Through the flexible adjustment of the dynamic kernel, the accuracy of feature segmentation is improved. Finally, after flattening, adjusting the category dimension through a convolution of a specific size and processing through an activation function, the segmentation result map of the power transmission channel features is obtained.
[0072] In the implementation, after step 103: based on the power transmission channel feature segmentation model sample set, a multi-round dynamic optimization strategy is adopted to train the pre-constructed power transmission channel feature semantic segmentation model network structure, and a trained power transmission channel feature segmentation model is obtained, step 104: based on the power transmission channel feature segmentation model, target segmentation mask data corresponding to the target power transmission channel digital orthographic image data is generated; the target segmentation mask data is raster vectorized to obtain corresponding frame vector data; and a topological optimization algorithm is used to integrate the frame vector data to obtain a continuous vector surface result of the entire power transmission channel.
[0073] In an embodiment, a boundary tracking and topology reconstruction algorithm is used to process each of the divided mask data, and by identifying the connected region boundary profile and reconstructing the topological relationship of the surface elements, vector surface data corresponding to the original standard divided DOM spatial coordinates is generated. The conversion process strictly maintains the geometric accuracy of the features, and eliminates the inherent jagged boundary phenomenon of the raster data.
[0074] A topology optimization algorithm is used to integrate and process the divided vector data. First, the area value and the rotation minimum bounding box parameters of all surface elements are calculated, and based on the area and perimeter value, the compactness index of the surface elements is calculated. In combination with the actual business requirements of the power transmission channel, the area threshold and the compactness threshold are set, and the broken surface elements whose area is less than the preset area threshold or which simultaneously meet the compactness less than the preset compactness threshold and the area less than another preset area threshold are identified. Through spatial neighborhood analysis, the adjacent region of the surface element to be optimized is located, and a vector fusion operation is performed on the adjacent maximum area surface element. Finally, a continuous and seamless vector surface data set covering the entire power transmission channel is formed, and the problem of small broken surfaces unique to high-resolution images is completely eliminated.
[0075] In an embodiment, the target segmentation mask data is converted into raster vector data to obtain corresponding divided vector data, including: converting the raster semantic segmentation result of each division in the target segmentation mask data into vector surface element data; and generating divided vector data corresponding to the spatial coordinates of the target power transmission channel digital orthographic image data according to the vector surface element data of different divisions. In the above embodiment, each division of the model output divided segmentation mask data is processed. A boundary tracking and connected domain identification algorithm is used to scan the raster semantic segmentation result of each division, and a contour extraction technique is used to aggregate adjacent pixels of the same category into a continuous closed boundary.
[0076] Based on the extracted boundary profile, topologically complete vector surface element data is constructed, and the spatial form and attribute information of each land feature category such as farmland, forest land, and building are completely preserved. The conversion process strictly maintains the pixel-level accuracy of the original raster data, and ensures that the vector surface boundary is accurately matched with the image features.
[0077] According to the vector surface element data generated by different divisions, a spatial coordinate system matching operation is performed.
[0078] The spatial reference information of the target power transmission channel digital orthographic image data is read, and the vector surface element data of each division is automatically aligned to the same coordinate system. Through node coordinate conversion and projection parameter synchronization, a divided vector data set corresponding to the spatial coordinates of the original digital orthographic image data is generated. The data set retains the division structure of the standard divided tiles, and the coverage range and boundary accuracy of each divided vector data are consistent with the input image, forming a standardized divided vector result that can be directly used for splicing operation.
[0079] In one embodiment, a topological optimization algorithm is used to integrate the divided vector data to obtain a continuous vector surface result of the whole transmission channel, including:
[0080] The area value and the perimeter value of each vector surface element in the divided vector data are calculated.
[0081] For each vector surface element, the length-width parameters of the rotationally minimum bounding box of each vector surface element are extracted; based on the area value and the perimeter value, the compactness of the corresponding vector surface element is calculated; a set area threshold and a compactness threshold are set to determine the vector surface element to be optimized; the area of the vector surface element to be optimized is less than a first preset area threshold, or the compactness of the vector surface element to be optimized is less than a preset compactness threshold and the area is less than a second preset area threshold.
[0082] The adjacent area of the vector surface element to be optimized is located through spatial neighborhood analysis, and a vector merging operation is performed on the vector surface element to be optimized and the vector surface element with the maximum area in the adjacent area to form a continuous vector surface result of the whole transmission channel.
[0083] In the above embodiment, the area value and the perimeter value parameters of all vector surface elements are first calculated. For each surface element, the length-width parameters of its rotationally minimum bounding box are extracted, which accurately quantify the spatial extension characteristics of the surface element. Based on the area value and the perimeter value, the compactness index of the surface element is calculated, which represents the compactness degree of the surface element shape.
[0084] According to the business requirements of the transmission channel, the area threshold and the compactness threshold are set, and the surface elements that meet the area less than the preset area threshold or simultaneously meet the compactness less than the preset compactness threshold and the area less than another preset area threshold are identified as the optimization objects. Such surface elements usually present as small broken surfaces or narrow ground objects unique to high-resolution images.
[0085] The adjacent area range of the surface element to be optimized is located through spatial neighborhood analysis technology. A topological relationship traversal algorithm is used to scan all adjacent surface elements that share boundaries with the surface element to be optimized, and the adjacent surface element with the maximum area is selected from them. A vector fusion operation is performed on the surface element to be optimized and the adjacent surface element with the maximum area, and the broken surface is eliminated through node merging and boundary reconstruction. This process is repeated for all surface elements to be optimized, and finally a seamless continuous vector surface data set covering the whole transmission channel is generated.
[0086] The topology optimization process realizes accurate quantification of the shape based on the rotating minimum bounding box parameters, and identifies the narrow and broken areas to be optimized through the compactness index. The spatial neighborhood analysis ensures that the merging operation conforms to the natural feature distribution law, and the maximum area merging strategy guarantees the spatial consistency of the optimized vector surface. The final result completely solves the problem of small fragments of 0.1-meter high-resolution images, and forms continuous vector surface basic data which can be directly used for channel floating object monitoring, illegal building identification and other operation and maintenance businesses.
[0087] Two specific embodiments are given below to illustrate the specific application of the device of the present application.
[0088] First specific embodiment:
[0089] The embodiment provides a power transmission channel image feature intelligent segmentation method based on deep learning, and the method steps include:
[0090] Step 1, construct a power transmission channel feature segmentation training sample library, adopt a labeled rasterization processing technology, a labeled automatic verification technology, a sample balance quantification analysis and enhancement technology and a training sample library automatic construction technology to realize the construction of the training sample library, Figure 2 is a specific example of a power transmission channel feature segmentation training sample library construction technology flow in the embodiment of the present application, and the specific process is as shown in Figure 2 .
[0091] Figure 2 In the embodiment, the input data includes: standard framing DOM data of the power transmission channel and corresponding labeled vector data;
[0092] and the corresponding labeled vector data is labeled rasterization processing; then the standard framing DOM data is compared and verified; then, based on the IOU labeled result automatic verification, including sample balance analysis and sample enhancement expansion, the sample library is divided and constructed; finally, the power transmission channel image feature segmentation sample library is obtained.
[0093] Step 2, design and train the power transmission channel feature segmentation model structure, adopt a power transmission channel feature segmentation model structure construction technology, a model training technology based on a multi-round dynamic optimization training strategy, and construct a power transmission channel feature segmentation exclusive AI inference model, Figure 3 is a specific example of a power transmission channel feature segmentation model structure design and model training technology flow in the embodiment of the present application, and the specific process is as shown in Figure 3 .
[0094] Figure 3The power transmission channel ground object semantic segmentation model network structure is constructed, including: taking Swin-Transformer as the backbone; taking FNP+double attention mechanism as the Neck; taking a dynamic kernel generator and a dynamic kernel optimizer as the kernel; and taking a dynamic kernel segmentation head as the Head.
[0095] Through the combination of the above-mentioned modules, the power transmission channel ground object semantic segmentation network structure is obtained, so as to perform model training. In the ground object semantic segmentation model training process, the following operations are involved: based on the power transmission channel image ground object segmentation sample library, the model training parameters are preset; the obtained power transmission channel ground object semantic segmentation network structure is used for model training; the validation set index and the model weight and the training state are recorded; then, it is judged whether the validation set index is optimal, if yes, the model weight is determined; if not, the model parameters are adjusted and the model training is performed again. Figure 3 In the embodiment of the present application, the parameters need to be preset before training, and the hyperparameters are initialized. The training process is to use data to make the model learn, and the performance is monitored by recording the validation set index. The final judgment link is the key decision point: if the index is not optimal, adjust the parameters and retrain, form an iterative optimization cycle, and ensure that the model performance continues to improve; until the index is optimal, the cycle is terminated, and the final available model weight is output.
[0096] Step three, after model inference post-processing, the raster vectorization and splicing technology and the vector surface optimization processing are used to realize the optimization processing of the model inference result, Figure 4 is a specific example of a model inference post-processing technical process in the embodiment of the present application, and the specific process is as shown in Figure 4 .
[0097] 1. Process start: segmentation mask result data: this module represents the input data of the process, that is, the segmentation mask result generated by the model inference in the form of a frame.
[0098] 2. Mask raster data vectorization: receiving the segmentation mask data of the last step, converting it into vector surface data through the raster vectorization algorithm, realizing the conversion from pixel data to geometric elements.
[0099] 3. Vector data splicing processing: the vector surface data generated by the frame is integrated into continuous vector data set covering the whole power transmission channel through the topological splicing algorithm, and the frame boundary gap is eliminated.
[0100] 4. Calculate the area and perimeter of the surface element: calculate the geometric parameters of the spliced vector surface element, obtain the area and perimeter value of each surface element, and provide quantitative basis for subsequent optimization.
[0101] 5. Calculate the compactness of the face element: Calculate the compactness index of the face element based on the area and perimeter value (formula C=P24πA), which is used to quantify the regularity of the shape of the face element.
[0102] 6. Is the compactness of the face element less than 0.1? If the compactness is less than 0.1, go to the area secondary judgment; otherwise, jump to the area independent judgment process.
[0103] Is the area of the face element less than 2? For elements with compactness less than 0.1: If the area is also less than 2 square meters, mark it as a normal element; otherwise, mark it as a to-be-optimized element. Is the area of the face element less than 0.5? For elements with compactness not less than 0.1: If the area is less than 0.5 square meters, mark it as a to-be-optimized element; otherwise, mark it as a normal element.
[0104] 7. Mark the face element as a to-be-optimized element: Identify the broken face elements that need to be optimized (including objects with small area or narrow shape).
[0105] 8. Mark the face element as a normal element: Identify the complete face elements that meet the business requirements and directly output to the result data.
[0106] 9. Normal element output: Include the normal elements directly into the final result data set.
[0107] 10. To-be-optimized element output: Pass the to-be-optimized elements to the merging processing module.
[0108] 11. Face element merging: Perform spatial neighborhood analysis on the to-be-optimized elements and perform vector fusion operation with the adjacent maximum area face elements to eliminate broken faces.
[0109] 12. End of the process: Output the continuous and complete vector face data set formed after optimization, which can be directly used for transmission channel operation and maintenance business.
[0110] Further, the labeling rasterization processing technology in step one is based on the standard framed DOM data (.tif) of the transmission channel and the corresponding labeled vector data (.shp), and uses a vector data rasterization processing algorithm to convert the format of the labeled data. The labeling automation verification technology is that due to the strip-shaped distribution of the transmission channel, the standard framed DOM data does not completely fill the entire framed tile, and a large no-data area will reduce the training accuracy, so the data needs to be verified and removed. The sample balance quantification analysis and enhancement technology is based on the labeled mask data after automatic verification, and the number and proportion of pixels of different categories are counted to quantitatively analyze the balance of sample categories and perform data enhancement processing on categories with fewer samples. The training sample library is automatically constructed by slicing and dividing the processed samples to construct the data set.
[0111] Further, the power transmission channel ground object segmentation model structure construction technology in step two constructs an end-to-end series power transmission channel ground object semantic segmentation network structure with Swin-Transformer as the backbone, with the fusion of FNP and double attention mechanism as the Neck, with the dynamic kernel generator and dynamic kernel optimizer as the kernel, and with the dynamic kernel segmentation head as the Head by comparing and analyzing different model modules.
[0112] Further, the grid vectorization and splicing technology in step three converts and splices the segmentation mask data generated by inference into a whole; the vector face optimization processing technology optimizes some small or small and narrow vector objects due to the high resolution of 0.1 m spatial resolution of the power transmission channel orthophoto image data.
[0113] The second specific embodiment provides a power transmission channel image ground object intelligent segmentation method based on deep learning, which is specifically described as follows:
[0114] Step one, construct a power transmission channel ground object segmentation training sample library.
[0115] Substep one, label rasterization processing, based on the standard framed DOM data (.tif) and the corresponding labeled vector data (.shp) of the power transmission channel, use the vector data rasterization processing algorithm to convert the format of the labeled data. The specific process is as follows:
[0116] First, automatically check the coordinate reference information of the standard framed DOM data, use the GDAL module to read the coordinate reference information of the labeled vector data and the standard framed DOM data respectively, compare the coordinate reference information of the standard framed DOM data with that of the labeled vector data, if they are consistent, then pass the verification, otherwise, define the coordinate system and perform coordinate transformation output;
[0117] Second, read the information in the field representing the labeled category of the labeled vector data, including cultivated land, forest land, scattered trees, shrubs and grasses, grassland, road, highway, railway, bare land surface, hardening ground, water system, building, construction operation area, easy floating object-plastic greenhouse, easy floating object-mulch, easy floating object-dustproof net, power transmission tower and other, and map the labeled category to RGB color. Based on the above rules, use the GDAL module to convert the labeled vector data into RGB three-band raster data according to the range of the standard framed DOM data, denoted as labeled mask data.
[0118] Sub-step two, annotation automation verification. Due to the strip distribution of the power transmission channel, the standard framed DOM data will not completely fill the entire framed tile, and the large no-data area will reduce the training accuracy, so the data needs to be checked and removed. The specific process is as follows: first, add an alpha band to the standard framed DOM data and the annotation mask data to generate RGBA four-band standard framed DOM and corresponding annotation mask data; second, calculate the IoU of the four-band standard framed DON and the corresponding annotation mask, define the unverified threshold as IoU less than or equal to 0.6, if the IoU is greater than the threshold, remove the annotation data and the corresponding standard framed DOM data.
[0119] Sub-step three, sample balance quantitative analysis and enhancement. Based on the automatically verified annotation mask data, the pixel number and pixel ratio of different categories are counted, and the balance of the sample categories is quantitatively analyzed, and the data enhancement processing is performed on the categories with less sample number. The specific process is as follows: first, based on the annotation mask data, the total grid number and the grid number of each category are counted, and the grid (pixel) ratio of each category is calculated; second, according to the calculated grid ratio of each type, select the categories with small ratio (less than 1%), such as greenhouse, ground, road, etc., as the sample enhancement object. Geometric transformation (rotation, cropping) and radiation transformation (brightness, contrast disturbance, etc.) are used for sample enhancement; finally, based on the supplemented standard framed DOM data and the corresponding annotation mask data, the grid ratio of each type is recalculated, which provides the basis for the weighted loss function modification of subsequent model training.
[0120] Sub-step four, automatic construction of training sample library. The specific process is as follows: first, based on the above standard framed DOM data and the corresponding annotation mask data, the data (5000x5000 pixels) is sliced again using the slicing algorithm, and the slice size is 2500x2500 pixels; second, the sliced data is divided into training set, validation set and test set according to the ratio of 7:2:1, the training set with a ratio of 70% is used for model learning and parameter adjustment; the validation set with a ratio of 20% is used to evaluate the performance of the model during training, so as to adjust the hyperparameters and select the model; the test set with a ratio of 10% is used to independently and objectively evaluate the performance of the final model after the model training is completed.
[0121] Step two, power transmission channel ground object segmentation model structure design and model training. Sub-step one, construct the power transmission channel ground object semantic segmentation model network structure, construct an end-to-end series of power transmission channel ground object semantic segmentation network structure with Swin-Transformer as the backbone, with FNP and double attention mechanism as the neck, with dynamic kernel generator and dynamic kernel optimizer as the kernel, and with dynamic kernel segmentation head as the head. The specific process is as follows: first, use Swin-Transformer as the backbone network. Through its hierarchical window attention mechanism, the multi-scale features of the power transmission channel ground object are extracted step by step. Each stage calculates the local features through window self-attention, and realizes global feature fusion through cross-window connection, effectively capturing the texture and shape information of vegetation, buildings, roads and other targets; second, use FNP and double attention mechanism as the model's neck. FNP performs cross-level fusion on the multi-scale features output by the backbone, enhancing the semantic expression of the target, using a parallel calculation method of channel attention and spatial attention to optimize the feature channel weight and spatial position sensitivity respectively, and improving the positioning accuracy of the ground object; third, use dynamic kernel generator and dynamic kernel optimizer as the model's kernel. According to the input feature map, dynamic kernel parameters are generated for different ground object categories, which adaptively match the diversity of the power transmission channel ground object form. Through iterative update of kernel parameters by back propagation, the discriminative ability of the kernel is optimized using the gradient descent algorithm to solve the class imbalance problem; finally, use dynamic kernel segmentation head as the model's head. Load the dynamic kernel parameters into the dynamic kernel segmentation head, perform the derivable operation of the convolution kernel and the feature map, and output the pixel-level semantic segmentation result. Through the flexible adjustment of the dynamic kernel, the ground object segmentation accuracy is significantly improved. Sub-step two, model training, based on the power transmission channel ground object semantic segmentation model network structure constructed above, combined with the proportion of each type of sample grid in the training sample library statistics, use the multi-round dynamic optimization training strategy to train the model.
[0122] The specific process is as follows: first, considering the proportion of each type of sample grid in the training sample library statistics, preset the learning rate, loss function weight and other parameters; second, based on the preset parameters, set a certain number of initial training for the training set in the sample library, record the validation set indicators such as each class confusion matrix and recall rate, and save the model weight and training state; third, based on the recorded validation set indicators, adjust the model training parameters such as loss function weight and learning rate; finally, repeat the above two steps until the model is trained to the optimal, and output the model weight. Step three, model inference post-processing. Sub-step one, grid vectorization and splicing technology. The specific process is as follows: first, based on the segmentation mask data in the form of frames generated by the ground object semantic segmentation model inference prediction, use the grid vectorization algorithm to convert the segmentation mask data into vector data, and generate the semantic segmentation vector data corresponding to the standard frame DOM;
[0123] Secondly, the generated frame vector data is spliced into a whole by using a vector splicing algorithm. In step two, the vector surface optimization processing is performed. Since the orthographic image data of the power transmission channel has a high resolution of 0.1 m spatial resolution, there are some small or small and narrow vector objects, which need to be optimized. The specific process is as follows: first, the area, perimeter, and length and width of the rotation minimum bounding box of each surface element in the vector data are calculated; secondly, the aspect ratio and compactness of the rotation minimum bounding box of each surface element are calculated. The compactness is the main basis for representing whether the surface element is narrow. C represents the compactness of the surface element, A represents the area of the surface element, and P represents the perimeter of the surface element. The formula is as follows:
[0124]
[0125] Thirdly, combined with the business requirements and the actual data acquisition range of the power transmission channel, the surface elements with an area A less than 0.5 m 2 or a compactness C less than 0.1 and an area A less than 2 m 2 are set as optimization objects; finally, the largest area of the surface element adjacent to the surface element to be optimized is read, and the surface element to be optimized is combined with it.
[0126] The specific embodiment is specially optimized and innovated for the algorithm and process of the image feature intelligent segmentation processing task of the power transmission channel, constructs a dedicated model for image feature segmentation of the power transmission channel, and realizes efficient and high-precision segmentation of the power transmission channel feature. At the same time, the specific implementation process of the model inference post-processing is specially proposed, so that the generated feature segmentation result has high business application, and the feature segmentation result can also use the training sample construction process proposed in the present application to continuously construct training samples, and realize continuous improvement and optimization of the model.
[0127] In addition, a specific example of a power transmission channel feature segmentation model network structure diagram is provided, Figure 5 which is a specific schematic diagram of a power transmission channel feature segmentation model network structure diagram in the embodiment of the present application, and the model structure is as shown in Figure 5 .
[0128] First, Swin-Transformer is used as the backbone network. After inputting the original transmission channel remote sensing image, the patch embedding operation is first performed to divide the image into non-overlapping image blocks, and each image block is linearly mapped to a fixed-dimensional feature vector while retaining the spatial position information. The first Swin Transformer block is entered, and the hierarchical window attention mechanism is used to divide the feature map into multiple local windows, and self-attention is calculated in each window to capture local feature relationships.
[0129] After completing the first stage of calculation, the patch merging operation is performed to merge 4 adjacent feature blocks into 1, and the semantic level of the feature map is improved through dimension compression while reducing the spatial resolution. The above Swin Transformer block calculation and patch merging process are repeated to extract multi-scale features of transmission channel ground objects step by step, and global feature fusion is achieved through cross-window connection to effectively capture texture and shape information of vegetation, buildings, roads, and other targets. Finally, 4 feature maps of different scales are output.
[0130] Second, the feature fusion path (FNP) and dual attention mechanism are used as the neck of the model. The four multi-scale feature maps output by the backbone network are received and fused across levels through the feature fusion path (FNP): first, perform 2 times upsampling (x2 upsample) on the high-level feature map to match the size of the next high-level feature map, then perform feature integration through cascade convolution (Cascade Conv), and repeat this process until the bottom layer feature map is fused (Merge); respectively sent to the channel attention (CAM) and spatial attention (SAM) modules for parallel calculation: the channel attention module learns the importance weight of different feature channels through global average pooling and multilayer perceptron, and the spatial attention module generates a spatial attention map through 1×1 convolution (1×1 Conv) and sigmoid activation function; the output features of the two attention modules are weighted and fused to optimize the feature channel weight and improve the spatial position sensitivity, enhancing the semantic expression and positioning accuracy of the target.
[0131] Again, the dynamic core generator and the dynamic core optimizer are used as the dynamic core module (Kernel) of the model. Through operations such as convolution (Conv) and mask prediction (Mask Predictions), the initial basic information required for kernel update is obtained. Based on these basic information, the first kernel update (Kernel Update) is performed to generate preliminary dynamic kernel related parameters. Then, using class prediction (Class Predictions) and mask prediction, a dynamic kernel (Dynamic Kernel) is generated, and the S-2 iteration process is repeated. In each iteration, based on the current features and prediction results, the dynamic kernel is updated (Kernel Update) to continuously adjust the kernel parameters to better adapt to the diversity of the transmission channel features (such as vegetation, buildings, etc.), and to solve the class imbalance problem by dynamically adjusting the parameter proportion of different class kernels, thereby improving the discrimination ability of the kernel.
[0132] Finally, the dynamic kernel segmentation head is used as the Head of the model. The dynamic kernel parameters and related features output by the dynamic kernel module are received, and the dynamic kernel parameters are loaded into the dynamic kernel segmentation head. The convolution kernel and the feature map are operated, and the feature map is processed pixel by pixel, and the pixel-level semantic segmentation result is output. Through the flexible adjustment of the dynamic kernel, the accuracy of feature segmentation is significantly improved, and finally after the flattening, 1x1 convolution (adjusting the class dimension) and Softmax function processing, the feature segmentation result image of the transmission channel is obtained.
[0133] Of course, it can be understood that the above detailed process can also have other variations, and the related variations should fall within the protection scope of the present application.
[0134] As described above, the present application proposes a deep learning-based intelligent feature segmentation method for transmission channel images, which uses transmission channel feature segmentation training sample library construction technology to realize the transmission channel image feature segmentation standard sample library. This technology supports the continuous and rapid output of the work results into the sample library, thereby realizing the continuous training of the model. Through in-depth analysis of the deep learning theory, the Swin-Transformer, dynamic kernel mechanism, spatial and channel dual attention mechanism, and other technologies are integrated to build a "Backbone-Neck-Kernel-Head" end-to-end series structure of a transmission channel image feature segmentation exclusive AI model, which improves the intelligent feature segmentation capability of the transmission channel. The model inference post-processing technology is used to realize the segmentation result vectorization and small face removal.
[0135] The embodiment of the present application also provides a power transmission channel image ground object intelligent segmentation device, as described in the following embodiment. Since the principle of solving the problem of the device is similar to the power transmission channel image ground object intelligent segmentation method, the implementation of the device can be referred to the implementation of the power transmission channel image ground object intelligent segmentation method, and the repeated parts will not be described here.
[0136] The embodiment of the present application also provides a power transmission channel image ground object intelligent segmentation device, which can improve the efficiency, precision and accuracy of the ground object segmentation of the power transmission channel image, Figure 6 For the structural example of the power transmission channel image ground object intelligent segmentation device in the embodiment of the present application, as shown in Figure 6 The device comprises:
[0137] The label mask data generation module 601 is configured to perform label rasterization processing, label automatic verification processing and sample enhancement processing on the power transmission channel digital orthographic image data and the corresponding vector label, so as to obtain the power transmission channel digital orthographic image data and the corresponding label mask data.
[0138] The sample set generation module 602 is configured to perform data slicing processing on the power transmission channel digital orthographic image data and the corresponding label mask data, so as to obtain a power transmission channel ground object segmentation model sample set.
[0139] The power transmission channel ground object segmentation model output module 603 is configured to train a pre-constructed power transmission channel ground object semantic segmentation model network structure based on the power transmission channel ground object segmentation model sample set by adopting a multi-round dynamic optimization strategy, so as to obtain a trained power transmission channel ground object segmentation model. The power transmission channel ground object semantic segmentation model network structure adopts a sliding window transformer backbone network to extract multi-scale power transmission channel ground object features, optimizes semantic expression by fusing a feature pyramid and a double attention mechanism, and adaptively generates ground object segmentation kernel parameters by using a dynamic kernel generator. After iterative updating by a dynamic kernel optimizer, the power transmission channel digital orthographic image data corresponding to the segmentation mask data is output by a segmentation head.
[0140] The vector data integration module 604 is configured to generate target segmentation mask data corresponding to target power transmission channel digital orthographic image data based on the power transmission channel ground object segmentation model. The target segmentation mask data is subjected to raster vector conversion to obtain corresponding divided vector data. The divided vector data is integrated by using a topological optimization algorithm to obtain a continuous vector surface result of the entire power transmission channel.
[0141] In one embodiment, the annotation mask data generation module is specifically configured to: perform automatic coordinate reference checking on the power transmission channel digital orthographic image data and the vector annotation data; after determining that the coordinate systems of the power transmission channel digital orthographic image data and the vector annotation data are consistent through the automatic coordinate reference checking, obtaining field information of ground objects in the vector annotation data; mapping the field information of ground objects into RGB three-band color values; and converting the mapped field information of ground objects into three-band raster annotation mask data based on the digital orthographic image data.
[0142] In one embodiment, the annotation mask data generation module is specifically configured to: add an Alpha band to the digital orthographic image data and the three-band raster annotation mask data to generate four-band data; remove data in the four-band data that exceeds a set intersection-over-union threshold based on the intersection-over-union of the four-band data to obtain annotation automatic checking processing data.
[0143] In one embodiment, the annotation mask data generation module is specifically configured to: determine the proportion of pixels of different categories in the annotation automatic checking processing data; and perform geometric transformation of rotation and cropping, and radiation transformation processing of brightness and contrast disturbance on the category pixels whose proportion is less than a preset proportion threshold to obtain the power transmission channel digital orthographic image data and the corresponding annotation mask data.
[0144] In one embodiment, the sample set generation module is specifically configured to: perform secondary division based on the pixel size of the power transmission channel digital orthographic image data and the corresponding annotation mask data; generate sample slices of equal size through cutting operation on the divided data; and distribute the cut sample slices according to a predetermined proportion to generate a training set, a validation set and a test set of the power transmission channel ground object segmentation model.
[0145] In one embodiment, the power transmission channel ground object segmentation model output module is specifically configured to: preset an initial learning rate and a loss function weight parameter according to the proportion of each type of sample grid counted in the sample set; perform multiple rounds of iterative training on the power transmission channel ground object semantic segmentation model network structure based on the sample set; calculate the confusion matrix and recall rate index of each round after each round of training is completed through the validation set data; dynamically adjust the loss function weight and the learning rate parameter according to the calculation result; and continuously perform the above multiple rounds of iterative training and the loss function weight and learning rate parameter adjustment operation to obtain the trained power transmission channel ground object segmentation model.
[0146] In one embodiment, the network structure construction module is further included for constructing the power transmission channel ground object semantic segmentation model network structure in the following manner: extracting hierarchical features through a sliding window transformer backbone network; capturing multi-scale power transmission channel ground object features of the power transmission channel ground object through window attention mechanisms for different hierarchical levels; optimizing semantic expression of the network structure through a feature fusion layer that fuses a feature pyramid network and a double attention mechanism; generating ground object adaptive segmentation kernel parameters according to the multi-scale power transmission channel ground object features through a dynamic kernel generator; iteratively updating the segmentation kernel parameters through a dynamic kernel optimizer through back propagation; and performing a differentiable convolution operation through a dynamic kernel segmentation head combined with the segmentation kernel parameters to output segmentation mask data in a segmented form that is spatially aligned with the input power transmission channel digital orthographic image data.
[0147] In one embodiment, the vector data integration module is specifically configured to: convert the raster semantic segmentation result of each segment of the target segmentation mask data into vector surface feature data; and generate segmented vector data corresponding to the spatial coordinates of the target power transmission channel digital orthographic image data according to the vector surface feature data of different segments.
[0148] In one embodiment, the vector data integration module is specifically configured to: calculate an area value and a perimeter value of each vector surface feature in the segmented vector data; extract a length-width parameter of a rotationally minimal bounding box of each vector surface feature for each vector surface feature; calculate a compactness of the corresponding vector surface feature based on the area value and the perimeter value; determine a vector surface feature to be optimized based on a set area threshold and a compactness threshold; the area of the vector surface feature to be optimized is less than a first preset area threshold, or the compactness of the vector surface feature to be optimized is less than a preset compactness threshold and the area of the vector surface feature to be optimized is less than a second preset area threshold; and perform a vector merging operation on the vector surface feature to be optimized and a vector surface feature with the largest area in a neighboring region of the vector surface feature to be optimized through spatial neighborhood analysis to position the neighboring region, thereby forming a continuous vector surface result of the entire power transmission channel.
[0149] An embodiment of a computer device for implementing all or part of the above-mentioned power transmission channel image ground object intelligent segmentation method is provided, and the computer device specifically includes the following content:
[0150] A processor, a memory, a communications interface and a bus; wherein the processor, the memory, the communications interface complete the communication between each other through the bus; the communications interface is used for realizing the information transmission between the related devices; the computer device can be a desktop computer, a tablet computer and a mobile terminal, etc., and the embodiments are not limited thereto. In the embodiments, the computer device can be implemented by referring to the embodiments for realizing the power transmission channel image ground object intelligent segmentation method and the embodiments for realizing the power transmission channel image ground object intelligent segmentation device, the contents are incorporated herein, and the repeated parts will not be described herein.
[0151] Figure 7 A computer device provided by the embodiments of the present application provides a schematic diagram of a computer device, which discloses a schematic block diagram of the system structure of the computer device 1000 of the embodiments of the present application. As shown in the figure, Figure 7 The computer device 1000 can include a central processor 1001 and a memory 1002; the memory 1002 is coupled to the central processor 1001. It is worth noting that the Figure 7 is exemplary; other types of structures can also be used to supplement or replace the structure to realize the telecommunication function or other functions.
[0152] As shown in the figure, Figure 7 The computer device 1000 can also include a communications module 1003, an input unit 1004, an audio processor 1005, a display 1006, a power supply 1007. It is worth noting that the computer device 1000 does not necessarily include all the components shown in the Figure 7 In addition, the computer device 1000 can also include components not shown in the Figure 7 can refer to the prior art.
[0153] As shown in the figure, Figure 7 The central processor 1001 is also sometimes referred to as a controller or an operation control, which can include a microprocessor or other processor device and / or a logic device, the central processor 1001 receives input and controls the operation of the various components of the computer device 1000.
[0154] The memory 1002, for example, can be one or more of a cache, a flash memory, a hard drive, a removable media, a volatile memory, a non-volatile memory or other suitable device. The above-mentioned information related to the device can be stored, and the program related to the information can also be stored. And the central processor 1001 can execute the program stored in the memory 1002 to realize information storage or processing, etc.
[0155] The input unit 1004 provides input to the central processing unit 1001. The input unit 1004 is, for example, a key or a touch input device. The power supply 1007 is used to supply power to the computer device 1000. The display 1006 is used to display display objects such as images and characters. The display 1006 can be, for example, an LCD display, but is not limited thereto.
[0156] The memory 1002 can be a solid-state memory such as a read-only memory (ROM), a random access memory (RAM), a SIM card, and the like. It can also be a memory that retains information even when power is off, can be selectively erased, and is provided with more data, an example of which is sometimes referred to as an EPROM or the like. The memory 1002 can also be some other type of device. The memory 1002 includes a buffer memory 1021 (sometimes referred to as a buffer). The memory 1002 can include an application / function storage section 1022 for storing application programs and function programs or for storing a flow for executing the operation of the computer device 1000 by the central processing unit 1001.
[0157] The memory 1002 can also include a data storage section 1023 for storing data such as contacts, digital data, pictures, sounds, and / or any other data used by the computer device. A driver storage section 1024 of the memory 1002 can include various drivers of the computer device for a communication function and / or for executing other functions of the computer device such as a messaging application, an address book application, and the like.
[0158] The communication module 1003 is a transmitter / receiver that transmits and receives signals via the antenna 1008. The communication module (transmitter / receiver) 1003 is coupled to the central processing unit 1001 to provide input signals and receive output signals, which can be the same as in the case of a conventional mobile communication terminal.
[0159] Based on different communication technologies, a plurality of communication modules 1003 can be provided in the same computer device, such as a cellular network module, a Bluetooth module, and / or a wireless local area network module, and the like. The communication module (transmitter / receiver) 1003 is also coupled to the speaker 1009 and the microphone 1010 via the audio processor 1005 to provide audio output via the speaker 1009 and receive audio input from the microphone 1010, thereby enabling a conventional telecommunication function. The audio processor 1005 can include any suitable buffer, decoder, amplifier, and the like. In addition, the audio processor 1005 is also coupled to the central processing unit 1001, thereby enabling recording on the local machine through the microphone 1010 and enabling playing of a sound stored on the local machine through the speaker 1009.
[0160] The embodiment of the present application further provides a computer readable storage medium, which stores a computer program, and the computer program is executed by a processor to realize the power transmission channel image ground object intelligent segmentation method.
[0161] The embodiment of the present application further provides a computer program product, which comprises a computer program, and the computer program is executed by a processor to realize the power transmission channel image ground object intelligent segmentation method.
[0162] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can adopt a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can adopt a computer program product in the form of being implemented on one or more computer usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer usable program codes.
[0163] The present application is described with reference to flowcharts and / or block diagrams of the method, device (system), and computer program product according to the embodiments of the present application. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, and the combination of the flows and / or blocks in the flowcharts and / or block diagrams can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate a means for implementing the functions specified in the flowcharts and / or block diagrams. Figure 1 The functions specified in one flow or multiple flows and / or blocks Figure 1 The means for implementing the functions specified in one flow or multiple flows and / or blocks.
[0164] The above-described specific embodiments further explain the purposes, technical solutions, and beneficial effects of the present application. It should be understood that the above-described specific embodiments are only for the specific embodiments of the present application and are not used to limit the protection scope of the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application should be included in the protection scope of the present application.
Claims
1. A method for intelligent segmentation of ground features in power transmission channel images, characterized in that, The method comprises the following steps: annotated raster processing, annotation automatic verification processing and sample enhancement processing are performed on the power transmission channel digital orthographic image data and the corresponding vector annotation to obtain the power transmission channel digital orthographic image data and the corresponding annotation mask data; data slicing processing is performed on the power transmission channel digital orthographic image data and the corresponding annotation mask data to obtain a power transmission channel ground object segmentation model sample set; based on the power transmission channel ground object segmentation model sample set, a multi-round dynamic optimization strategy is adopted to train a pre-constructed power transmission channel ground object semantic segmentation model network structure to obtain a trained power transmission channel ground object segmentation model; the power transmission channel ground object semantic segmentation model network structure adopts a sliding window transformer backbone network to extract multi-scale power transmission channel ground object features, optimizes semantic expression through a fusion feature pyramid and a double attention mechanism, and adaptively generates ground object segmentation kernel parameters by using a dynamic kernel generator; after iterative updating by a dynamic kernel optimizer, the power transmission channel digital orthographic image data corresponding segmentation mask data is output by a segmentation head; based on the power transmission channel ground object segmentation model, target segmentation mask data corresponding to target power transmission channel digital orthographic image data is generated; grid vectorization conversion is performed on the target segmentation mask data to obtain corresponding divided vector data; topological optimization algorithm is used to integrate the divided vector data to obtain continuous vector surface results of the entire power transmission channel; the method further comprises the following steps of constructing a power transmission channel ground object semantic segmentation model network structure: layered features are extracted by a sliding window transformer backbone network; different layers are captured by a window attention mechanism to gradually capture multi-scale power transmission channel ground object features of the power transmission channel ground object; a feature fusion layer adopting a fusion feature pyramid network and a double attention mechanism is used to optimize semantic expression of the network structure; a dynamic kernel generator is used to generate adaptive segmentation kernel parameters of the ground object according to the multi-scale power transmission channel ground object features; the segmentation kernel parameters are iteratively updated by a dynamic kernel optimizer through back propagation; a dynamic kernel segmentation head combines the segmentation kernel parameters to perform a derivable convolution operation, and outputs segmentation mask data in a divided form which is spatially aligned with the input power transmission channel digital orthographic image data.
2. The method of claim 1, wherein, annotated raster processing is performed on the power transmission channel digital orthographic image data and the corresponding vector annotation, including: automatic coordinate reference verification is performed on the power transmission channel digital orthographic image data and the vector annotation data; after the automatic coordinate reference verification determines that the coordinate systems of the power transmission channel digital orthographic image data and the vector annotation data are consistent, ground object field information in the vector annotation data is obtained; the ground object field information is mapped into RGB three-band color values; based on the digital orthographic image data, the mapped ground object field information is converted into three-band raster annotation mask data.
3. The method of claim 2, wherein, annotation automatic verification processing is performed on the power transmission channel digital orthographic image data and the corresponding vector annotation, including: four-band data is generated by adding an Alpha band to the digital orthographic image data and the three-band raster annotation mask data; based on the intersection over union of the four-band data, data exceeding a set intersection over union threshold is removed from the four-band data to obtain annotation automatic verification processing data.
4. The method of claim 1, wherein, Based on the power transmission channel ground object segmentation model sample set, a multi-round dynamic optimization strategy is adopted to train the pre-constructed power transmission channel ground object semantic segmentation model network structure, and a trained power transmission channel ground object segmentation model is obtained, including: According to the proportion of each type of sample grid in the sample set, the initial learning rate and the loss function weight parameter are preset; Based on the sample set, the power transmission channel ground object semantic segmentation model network structure is executed for multiple rounds of iterative training; After each round of training is completed, the confusion matrix and recall rate index of the corresponding round are calculated through the validation set data; According to the calculation result, the loss function weight and the learning rate parameter are dynamically adjusted; The above-mentioned multi-round iterative training and loss function weight and learning rate parameter adjustment operations are continuously executed to obtain the trained power transmission channel ground object segmentation model.
5. The method of claim 1, wherein, The topological optimization algorithm is used to integrate the divided vector data to obtain the continuous vector surface result of the whole power transmission channel, including: Calculate the area value and perimeter value of each vector surface element in the divided vector data; For each vector surface element, extract the length and width parameters of the rotation minimum bounding box of each vector surface element; based on the area value and perimeter value, calculate the compactness of the corresponding vector surface element; set the area threshold and compactness threshold to determine the vector surface element to be optimized; the area of the vector surface element to be optimized is less than the first preset area threshold, or the compactness of the vector surface element to be optimized is less than the preset compactness threshold and the area is less than the second preset area threshold; Through spatial neighborhood analysis, the adjacent area of the vector surface element to be optimized is located, and the vector merging operation is performed on the vector surface to be optimized and the vector surface element with the maximum area in the adjacent area to form the continuous vector surface result of the whole power transmission channel.
6. An electric transmission channel image ground object intelligent segmentation device, characterized in that, It includes: The labeled mask data generation module is used for labeling rasterization processing, labeling automatic verification processing and sample enhancement processing on the power transmission channel digital orthographic image data and the corresponding vector labeling to obtain the power transmission channel digital orthographic image data and the corresponding labeled mask data; The sample set generation module is used for data slicing processing on the power transmission channel digital orthographic image data and the corresponding labeled mask data to obtain the power transmission channel ground object segmentation model sample set; The power transmission channel ground object segmentation model output module is used for training the pre-constructed power transmission channel ground object semantic segmentation model network structure based on the power transmission channel ground object segmentation model sample set, and obtaining the trained power transmission channel ground object segmentation model; the power transmission channel ground object semantic segmentation model network structure adopts a sliding window transformer backbone network to extract multi-scale power transmission channel ground object features, optimizes semantic expression through a fusion feature pyramid and a double attention mechanism, and uses a dynamic kernel generator to adaptively generate ground object segmentation kernel parameters; After iteration and update by the dynamic kernel optimizer, the segmentation mask data corresponding to the power transmission channel digital orthographic image data is output by the segmentation head; The vector data integration module is used for generating target segmentation mask data corresponding to the target power transmission channel digital orthographic image data based on the power transmission channel ground object segmentation model; The target segmentation mask data is converted into raster vector data. The topological optimization algorithm is used to integrate the divided vector data, and a continuous vector surface result of the whole transmission channel is obtained. The transmission channel object segmentation model output module is further configured to The transmission channel object semantic segmentation model network structure is constructed in the following manner: The hierarchical feature is extracted through the sliding window transformer backbone network; The multi-scale transmission channel object features of the transmission channel object are captured step by step through the window attention mechanism for different hierarchical layers; The feature fusion layer fusing the feature pyramid network and the double attention mechanism is used to optimize the semantic expression of the network structure; the dynamic kernel generator is used to generate the adaptive segmentation kernel parameters of the object according to the multi-scale transmission channel object features; The segmentation kernel parameters are iteratively updated through the dynamic kernel optimizer through back propagation; The dynamic kernel segmentation head combines the segmentation kernel parameters to perform the differentiable convolution operation, and outputs the segmentation mask data in the divided form which is spatially aligned with the input transmission channel digital orthographic image data.
7. A computer device comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, The processor executes the computer program to implement the method of any one of claims 1 to 5.
8. A computer-readable storage medium, characterized in that, The computer readable storage medium stores a computer program, and the computer program is executed by the processor to implement the method of any one of claims 1 to 5.
9. A computer program product, characterised in that, The computer program product comprises a computer program, and the computer program is executed by the processor to implement the method of any one of claims 1 to 5.
Citation Information
Patent Citations
Semantic segmentation method based on dynamic kernel and Gaussian kernel fusion strategy
CN118314354A
Language-guided video target anaphora segmentation method
CN118658091A