Visual distribution line conductor segmentation method based on line attention mechanism and related system
Through the visual distribution line conductor segmentation method based on the line attention mechanism, the image information is processed using the row self-attention, column self-attention and mutual attention mechanism, which solves the problem of more error detection in the distribution line conductor segmentation by the general visual model, and achieves the segmentation effect of high precision and high generalization ability.
Patent Information
- Application Number
- CN202510357977.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-25
- Publication Date
- 2025-07-11
AI Technical Summary
The general visual model has many errors in the distribution line conductor body and defect segmentation, and has insufficient generalization ability and cannot reach the practical level.
The visual distribution line conductor segmentation method based on the line attention mechanism is adopted, and the image information is processed through the row attention, column attention, self-attention and mutual attention mechanism, and the prompt word information is encoded and decoded, and the features of different scales are gradually captured.
It improves the accuracy and generalization ability of wire segmentation of distribution lines, reduces false detection and missed detection, and meets the needs of efficient and robust segmentation in complex environments.
Smart Images

Figure CN120298426A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of distribution line segmentation, and specifically relates to a visual distribution line conductor segmentation method and related system based on a line attention mechanism. Background Technique
[0002] A distribution line refers to a line that delivers electricity from a step-down substation to a distribution transformer or from a distribution substation to a power-consuming unit. Basic requirements for the operation and maintenance of distribution lines are to be safe and reliable, maintain power supply continuity, reduce line losses, improve transmission efficiency, and ensure good power quality. Due to being exposed to the natural environment for a long time, distribution lines are prone to defects such as conductor corrosion and tree obstacles. If not discovered and processed in time, it will affect the stable operation of the distribution line. Regular inspection of the distribution line can master the operation status of the line, timely discover defects and potential hazards threatening the safe operation of the line along the line, thereby improving power supply reliability and reducing the occurrence of line accidents.
[0003] Currently, the identification of the distribution line conductor body and related defects mainly uses methods such as object detection and semantic segmentation based on deep neural networks to automatically identify the images collected by distribution network unmanned aerial vehicles. Due to relatively few labeled samples of conductor equipment bodies and defects and relatively special morphological features, the above methods have low accuracy in conductor equipment bodies and defects and cannot reach the practical level. At present, with the rise of large model technology, general vision large models show strong generalization ability in fields such as classification, detection, and segmentation due to their advantages in the number of parameters and data volume. However, due to the particularity of the power scenario and the complexity of conductor bodies and defects, general vision large models still have many misdetections and insufficient generalization ability in such scenarios.
[0004] How to modify the structure of the general vision large model to adapt to the characteristics of distribution network conductor body and defect data, combined with the original generalization ability of the general vision large model, to improve the segmentation accuracy of the general vision large model for distribution network conductor bodies and related defects has important practical value. Summary of the Invention
[0005] The purpose of the present invention is to overcome the problem that the general vision large model still has many misdetections and insufficient generalization ability in such scenarios, and provide a visual distribution line conductor segmentation method and related system based on a line attention mechanism.
[0006] To achieve the above purpose, the present invention adopts the following technical solutions: In the first aspect, the present invention provides a visual distribution line conductor segmentation method based on a line attention mechanism, including the following steps: Obtain the image information and prompt word information of the distribution line conductor; Process the image information of the distribution line conductor to obtain a sequence of image patches; The information in each row of image patches in the image patch sequence is updated row by row from top to bottom using the row self-attention formula to obtain an updated row image patch sequence; the information in each column of image patches in the image patch sequence is updated column by column from left to right using the column self-attention formula to obtain an updated column image patch sequence; the updated row image patch sequence and the updated column image patch sequence are combined to obtain an updated three-dimensional image patch sequence; The updated three-dimensional image patch sequence is recombined into a two-dimensional image patch sequence, and the two-dimensional image patch sequence is updated using the self-attention formula and an image encoder to obtain an updated two-dimensional image patch sequence, and the updated two-dimensional image patch sequence is recombined into a three-dimensional image patch sequence; Taking the image patches within a row in the recombined three-dimensional image patch sequence from top to bottom row by row as the first key vector and the first value vector, the prompt word information is encoded row by row to obtain a row query vector; taking the image patches within a column in the recombined three-dimensional image patch sequence from left to right column by column as the second key vector and the second value vector, the prompt word information is encoded column by column to obtain a column query vector; the row query vector is updated using the row cross-attention formula in combination with the first key vector and the first value vector to obtain an updated row query vector; the column query vector is updated using the column cross-attention formula in combination with the second key vector and the second value vector to obtain an updated column query vector; The updated row query vector and the updated column query vector are decoded to obtain the power distribution line conductor segmentation result.
[0007] A further improvement of the present invention lies in that the specific method for processing the image information of the power distribution line conductor to obtain an image patch sequence is as follows: The image information of the power distribution line conductor is normalized to obtain an image ; The image is subjected to convolution processing to obtain an image patch sequence ; where is the number of channels of the image patch sequence, is the height of the image, is the width of the image, is the downsampling factor.
[0008] A further improvement of the present invention lies in that the specific method for updating the information in each row of image patches in the image patch sequence row by row from top to bottom using the row self-attention formula to obtain an updated row image patch sequence is as follows: Obtain the feature vectors of all the image patches in the th row in the image patch sequence ; Use the row self-attention formula to update the feature vectors of the image patch sequence row by row from top to bottom for all rows in the
[0009] obtain the updated row image patch sequence , where is the dimension of the image patch feature vector.
[0010] A further improvement of the present invention is that, column by column, use the column self-attention formula to update the information in each column of the image patch sequence from left to right to obtain the updated column image patch sequence. The specific method is as follows: Obtain the image patch sequence in the column, and obtain the feature vectors of all image patches ; Use the column self-attention formula to update the feature vectors of the image patches in all columns of the image patch sequence from left to right column by column. The column self-attention formula is:
[0011] obtain the updated three-dimensional image patch sequence , where is the dimension of the image patch feature vector.
[0012] A further improvement of the present invention is that, the updated three-dimensional image patch sequence is reorganized into a two-dimensional image patch sequence, and the self-attention formula and the image encoder are used to update the two-dimensional image patch sequence to obtain the updated two-dimensional image patch sequence, and the specific method of reorganizing the updated two-dimensional image patch sequence into a three-dimensional image patch sequence is as follows: Reorganize the three-dimensional image patch sequence into a two-dimensional image patch sequence , with a dimension of , where , use the self-attention formula and the image encoder to update the two-dimensional image patch sequence to obtain the updated image patch sequence ; the self-attention formula is: .
[0013] A further improvement of the present invention is that, row by row from top to bottom, take out the image patches in one row of the reorganized three-dimensional image patch sequence as the first key vector and the first value vector, and perform row encoding on the prompt information to obtain the row query vector. The specific method is as follows: Obtain the updated image patch sequence The feature vectors of all image patches in the th row and use them as the first key vector and the first value vector ; Use the prompt encoder to encode the prompt information to obtain a feature sequence , and use it as the row query vector ; Combine the first key vector and the first value vector and use the row cross-attention formula to update the row query vector to obtain the updated query vector ; The row cross-attention formula is:
[0014] where is the number of rows of the three-dimensional image patch, is the dimension of the image patch feature vector.
[0015] A further improvement of the present invention is that, taking columns as units, the image patches in a column of the reorganized three-dimensional image patch sequence are taken out from left to right in sequence as the second key vector and the second value vector, and the specific method for column-encoding the prompt information to obtain the column query vector is as follows: Obtain the updated image patch sequence The feature vectors of all image patches in the kth column The second key vector and the second value vector ; Use the prompt encoder to encode the prompt information to obtain a feature sequence , and use it as the column query vector ; Combine the second key vector and the second value vector and use the column cross-attention formula to update the column query vector ; The column cross-attention formula is:
[0016] where is the number of columns of the three-dimensional image patch, is the dimension of the image patch feature vector.
[0017] In a second aspect, the present invention provides a visual power distribution line conductor segmentation system based on a line attention mechanism, including: A data acquisition module, configured to acquire image information and prompt information of a power distribution line conductor; An image processing module for processing the image information of the power distribution line conductors to obtain a sequence of image blocks; A row-column self-attention module for updating the information in each row of image blocks in the sequence of image blocks row by row from top to bottom using the row self-attention formula to obtain an updated sequence of row image blocks; updating the information in each column of image blocks in the sequence of image blocks column by column from left to right using the column self-attention formula to obtain an updated sequence of column image blocks; combining the updated sequence of row image blocks and the updated sequence of column image blocks to obtain an updated three-dimensional sequence of image blocks; A two-dimensional image block sequence updating module for reorganizing the updated three-dimensional sequence of image blocks into a two-dimensional sequence of image blocks, updating the two-dimensional sequence of image blocks using the self-attention formula and an image encoder to obtain an updated two-dimensional sequence of image blocks, and reorganizing the updated two-dimensional sequence of image blocks into a three-dimensional sequence of image blocks; A row-column cross-attention module for sequentially taking out the image blocks in a row of the reorganized three-dimensional sequence of image blocks as the first key vector and the first value vector row by row from top to bottom to perform row encoding on the prompt word information to obtain a row query vector; sequentially taking out the image blocks in a column of the reorganized three-dimensional sequence of image blocks as the second key vector and the second value vector column by column from left to right to perform column encoding on the prompt word information to obtain a column query vector; updating the row query vector using the row cross-attention formula in combination with the first key vector and the first value vector to obtain an updated row query vector; updating the column query vector using the column cross-attention formula in combination with the second key vector and the second value vector to obtain an updated column query vector; A decoding module for decoding the updated row query vector and the updated column query vector to obtain the power distribution line conductor segmentation result.
[0018] A further improvement of the present invention lies in that the function of the row-column self-attention module is implemented by the following method: Obtain the sequence of image blocks the feature vectors of all the image blocks in the row; Use the row self-attention formula to update the feature vectors of the image blocks in all rows of the sequence of image blocks row by row from top to bottom. The row self-attention formula is:
[0019] to obtain an updated sequence of row image blocks ; Obtain the sequence of image blocks the feature vectors of all the image blocks in the column; The column self-attention formula is used to update the feature vectors of the image patches in each column of the image patch sequence from left to right in sequence. Among them, the column self-attention formula is:
[0020] The updated three-dimensional image patch sequence is obtained. .
[0021] A further improvement of the present invention is that the function of the two-dimensional image patch sequence update module is realized by the following method: The three-dimensional image patch sequence is reorganized into a two-dimensional image patch sequence , The dimension is where , and the self-attention formula and the image encoder are used to update the two-dimensional image patch sequence to obtain the updated image patch sequence ; where the self-attention formula is: .
[0022] A further improvement of the present invention is that the function of the row-column mutual attention module is realized by the following method: Obtain the feature vectors of all the image patches in the th row of the updated image patch sequence and use them as the first key vector and the first value vector ; ; The prompt encoder is used to encode the prompt information to obtain the feature sequence , and use it as the row query vector ; Combine the first key vector and the first value vector and use the row mutual attention formula to update the row query vector to obtain the updated query vector ; the row mutual attention formula is: ; Obtain the feature vectors of all the image patches in the th column of the updated image patch sequence as the second key vector and the second value vector ; The prompt encoder is used to encode the prompt information to obtain the feature sequence , and use it as the column query vector ; Combine the second key vector and the second value vector The column cross-attention formula updates the column query vector as follows; the column cross-attention formula is: .
[0023] In a third aspect, the present invention provides an electronic device, including a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the steps of the visual distribution line conductor segmentation method based on the line attention mechanism are implemented.
[0024] In a fourth aspect, the present invention provides a storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the visual distribution line conductor segmentation method based on the line attention mechanism are implemented.
[0025] Compared with the prior art, the present invention has the following beneficial effects: By designing the line attention mechanism, the present invention can better capture the structured features of the distribution line conductors, reducing false detections and missed detections. The present invention captures features at different scales step by step from rows to columns and then to the global, making the model robust to distribution line conductors in different scenarios. The present invention guides the attention calculation through prompt information encoding, enabling the model to adapt to different input conditions and enhancing the generalization ability in diverse scenarios. The present invention decomposes the image into a sequence of image patches, reducing the computational dimension while retaining local context information, taking into account both computational efficiency and information expression ability. The step-by-step feature update (row update, column update, self-attention, cross-attention) of the multi-level attention mechanism of the present invention can fully explore the feature correlation of the image patches, reducing information loss. The design of the row, column self-attention and cross-attention calculations of the present invention is more focused on the feature update of specific regions, reducing unnecessary computational redundancy compared to global attention. Through the guidance of prompt information and step-by-step feature processing, the model of the present invention can flexibly adapt to distribution line images under different conditions. In summary, the false detection rate and missed detection rate of the present invention in complex environments are significantly reduced, and at the same time, it has efficient and robust segmentation capabilities, meeting the high requirements for accuracy and generalization ability in the distribution line conductor segmentation task. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] Figure 1 is a flowchart of the present invention; Figure 2 is a system diagram of the present invention; Figure 3 is a flowchart of Embodiment 1 and Embodiment 2; Figure 4 is a control diagram of Embodiment 1 and Embodiment 2; Figure 5 is a system diagram of Embodiment 4. Detailed implementation manners
[0027] To further understand the content of the present invention, the present invention will be described in detail below with reference to the accompanying drawings and specific embodiments. It should be understood that the embodiments are only for explaining the present invention rather than limiting it.
[0028] See Figure 1 , a visual distribution line conductor segmentation method based on a line attention mechanism, comprising the following steps: S1. Obtain the image information and prompt word information of the distribution line conductor.
[0029] S2. Process the image information of the distribution line conductor to obtain a sequence of image patches.
[0030] S3. Update the information in each row of image patches in the sequence of image patches row by row from top to bottom using the row self-attention formula to obtain an updated sequence of row image patches; update the information in each column of image patches in the sequence of image patches column by column from left to right using the column self-attention formula to obtain an updated sequence of column image patches; combine the updated sequence of row image patches and the updated sequence of column image patches to obtain an updated three-dimensional sequence of image patches.
[0031] S4. Recombine the updated three-dimensional sequence of image patches into a two-dimensional sequence of image patches, update the two-dimensional sequence of image patches using the self-attention formula and an image encoder to obtain an updated two-dimensional sequence of image patches, and recombine the updated two-dimensional sequence of image patches into a three-dimensional sequence of image patches.
[0032] S5. Take the image patches in one row of the recombined three-dimensional sequence of image patches row by row from top to bottom as the first key vector and the first value vector, perform row encoding on the prompt word information to obtain a row query vector; take the image patches in one column of the recombined three-dimensional sequence of image patches column by column from left to right as the second key vector and the second value vector, perform column encoding on the prompt word information to obtain a column query vector; update the row query vector using the row cross-attention formula in combination with the first key vector and the first value vector to obtain an updated row query vector; update the column query vector using the column cross-attention formula in combination with the second key vector and the second value vector to obtain an updated column query vector.
[0033] S6. Decode the updated row query vector and the updated column query vector to obtain the distribution line conductor segmentation result.
[0034] See Figure 2 , a visual distribution line conductor segmentation system based on a line attention mechanism, comprising: A data acquisition module, configured to acquire the image information and prompt word information of the distribution line conductor; An image processing module for processing the image information of the distribution line conductor to obtain a sequence of image blocks; A row-column self-attention module for updating the information in each row of image blocks in the sequence of image blocks row by row from top to bottom using the row self-attention formula to obtain an updated sequence of row image blocks; updating the information in each column of image blocks in the sequence of image blocks column by column from left to right using the column self-attention formula to obtain an updated sequence of column image blocks; combining the updated sequence of row image blocks and the updated sequence of column image blocks to obtain an updated three-dimensional sequence of image blocks; A two-dimensional image block sequence update module for reorganizing the updated three-dimensional sequence of image blocks into a two-dimensional sequence of image blocks, updating the two-dimensional sequence of image blocks using the self-attention formula and an image encoder to obtain an updated two-dimensional sequence of image blocks, and reorganizing the updated two-dimensional sequence of image blocks into a three-dimensional sequence of image blocks; A row-column cross-attention module for taking out the image blocks in a row within the reorganized three-dimensional sequence of image blocks row by row from top to bottom as the first key vector and the first value vector, performing row encoding on the prompt word information to obtain a row query vector; taking out the image blocks in a column within the reorganized three-dimensional sequence of image blocks column by column from left to right as the second key vector and the second value vector, performing column encoding on the prompt word information to obtain a column query vector; updating the row query vector using the row cross-attention formula in combination with the first key vector and the first value vector to obtain an updated row query vector; updating the column query vector using the column cross-attention formula in combination with the second key vector and the second value vector to obtain an updated column query vector; A decoding module for decoding the updated row query vector and the updated column query vector to obtain the segmentation result of the distribution line conductor.
[0035] Embodiment 1: See Figure 3 and Figure 4 , steps S2 - S5 together form a line self-attention wire shape prior modeling module, which includes 1 convolutional layer, 1 row self-attention module, 1 column self-attention module, and several conventional self-attention modules. The line self-attention wire shape prior modeling module acts before the image encoder of the general vision large model, and the specific method is as follows: Normalize the image information of the distribution line conductor to obtain an image ; Perform convolutional processing on the image to obtain a sequence of image blocks ; where is the number of channels of the sequence of image blocks is the height of the image, is the width of the image, is the downsampling factor.
[0036] Obtain the sequence of image patches the feature vectors of all image patches in the th row; Update the feature vectors of the image patch sequence from top to bottom in turn to obtain the updated image patch sequence ; The row self-attention formula is:
[0037] Obtain the updated row image patch sequence , is the dimension of the feature vector of the image patch.
[0038] Obtain the sequence of image patches the feature vectors of all image patches in the th column; Update the feature vectors of the image patch sequence from left to right in turn to obtain the updated image patch sequence ; The column self-attention formula is:
[0039] Obtain the updated three-dimensional image patch sequence .
[0040] Recombine the three-dimensional image patch sequence into a two-dimensional image patch sequence , with the dimension of , where , and use the traditional self-attention formula and image encoder to update the two-dimensional image patch sequence to obtain the updated image patch sequence The self-attention formula is:
[0041] This embodiment realizes the integration and transmission of the information between rows and columns in the image patch sequence, making the features extracted by the network match the unique slender shape of the wire, which is beneficial to the accurate segmentation of the wire in the subsequent process.
[0042] Embodiment 2: See Figure 3 and Figure 4, Steps S6 and S7 together constitute the line mutual attention wire shape prior modeling module. The line mutual attention wire shape prior modeling module includes 1 row mutual attention module and 1 column mutual attention module. The line mutual attention wire shape prior modeling module is used before the decoder of the general vision large model. The specific method is as follows: The updated image patch sequence in the feature vectors of all image patches in the first key vector and the first value vector ; Use the prompt encoder to encode the prompt information to obtain a feature sequence and use it as the row query vector
[0043] Combine the first key vector and the first value vector Update the row query vector using the row mutual attention formula to obtain the updated query vector ; The row mutual attention formula is:
[0044] where is the number of rows of the three-dimensional image patch.
[0045] The updated image patch sequence in the feature vectors of all image patches in the k-th column second key vector and the second value vector Use the prompt encoder to encode the prompt information to obtain a feature sequence and use it as the column query vector ; Combine the second key vector and the second value vector and use the mutual attention formula to update the column query vector ; The column mutual attention formula is:
[0046] where is the number of columns of the three-dimensional image patch.
[0047] This embodiment realizes the fusion of the prompt query vector with the inter-row and inter-column information in the image patch feature sequence, making the prompt query vector pay more attention to the inter-row and inter-column information of the image, which is beneficial to the accurate segmentation of the subsequent wire.
[0048] Example 3: S1: Obtain the image information and prompt information of the distribution line conductor; Assume that the size of the input image is H = 1024, W = 1024, and the number of channels is 3 (RGB image). The prompt information consists of a sequence of length n' = 64, and each word is encoded in a d' = 256 - dimensional representation.
[0049] S2: Normalize and process the image; S21, Normalize the input image to obtain the normalized image .
[0050] S22, Convolutional layer Use convolution kernels, with a size of 32×32 and a stride s = 32. The size of the output feature map is:
[0051]
[0052] Obtain the feature map .
[0053] S23, The feature vectors of all the image patches in the th row of the image patch sequence are represented as , and the formula:
[0054] is used to perform in - row self - attention update on the 32 - row image patch feature vectors to obtain .
[0055] S24, The feature vectors of all the image patches in the k - th column of the image patch sequence are represented as , and the formula:
[0056] is used to perform in - column self - attention update on the 32 - column image patch feature vectors to obtain .
[0057] S25, Recombine into a two - dimensional image patch sequence , and use the self - attention formula:
[0058] to update the two - dimensional image patch sequence , and then use the image encoder to further process the updated image patch sequence to obtain the image patch sequence .
[0059] S3. Obtain the prompt information and perform cross-attention update; S31. Obtain the feature vectors of all image patches in the first key vector and the first value vector ; Encode the prompt information using the prompt encoder to obtain a feature sequence and use it as the row query vector
[0060] Combine the first key vector and the first value vector and use the row cross-attention formula to update the row query vector to obtain the updated query vector ; The row cross-attention formula is:
[0061] S32. Obtain the feature vectors of all image patches in the k-th column of the second key vector and the second value vector ; Use the updated query vector as the column query vector ; Combine the second key vector and the second value vector and use the column cross-attention formula to update the column query vector to obtain the updated query vector ; The column cross-attention formula is:
[0062] S4. Decode the updated query vector to obtain the segmentation result of the power distribution line conductor.
[0063] Example 4: Please refer to Figure 5 as shown. The present invention also provides an electronic device 100 for a visual power distribution line conductor segmentation method based on a line attention mechanism; the electronic device 100 includes a memory 101, at least one processor 102, a computer program 103 stored in the memory 101 and executable on the at least one processor 102, and at least one communication bus 104.
[0064] The memory 101 can be used to store the computer program 103. By running or executing the computer program stored in the memory 101 and calling the data stored in the memory 101, the processor 102 implements the steps of the visual distribution line conductor segmentation method based on the line attention mechanism described in Embodiment 3. The memory 101 may mainly include a program storage area and a data storage area. Among them, the program storage area can store an operating system, application programs required for at least one function (such as a sound playback function, an image playback function, etc.); the data storage area can store data created according to the use of the electronic device 100 (such as audio data, etc.). In addition, the memory 101 may include non-volatile memory, such as a hard disk, a memory, a plug-in hard disk, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, at least one magnetic disk storage device, a flash memory device, or other non-volatile solid-state storage devices.
[0065] The at least one processor 102 can be a Central Processing Unit (CPU), or can also be other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The processor 102 can be a microprocessor or the processor 102 can also be any conventional processor, etc. The processor 102 is the control center of the electronic device 100, and connects various parts of the entire electronic device 100 through various interfaces and lines.
[0066] The memory 101 in the electronic device 100 stores multiple instructions to implement the visual distribution line conductor segmentation method based on the line attention mechanism. The processor 102 can execute the multiple instructions to thereby implement: Obtain the image information and prompt word information of the distribution line conductor; Process the image information of the distribution line conductor to obtain a sequence of image blocks; Obtain the image information and prompt word information of the distribution line conductor; Process the image information of the distribution line conductor to obtain a sequence of image blocks; Update the information in each row of image patches in the image patch sequence row by row from top to bottom using the row self-attention formula to obtain the updated row image patch sequence; update the information in each column of image patches in the image patch sequence column by column from left to right using the column self-attention formula to obtain the updated column image patch sequence; combine the updated row image patch sequence and the updated column image patch sequence to obtain the updated three-dimensional image patch sequence; Recombine the updated three-dimensional image patch sequence into a two-dimensional image patch sequence, and update the two-dimensional image patch sequence using the self-attention formula and the image encoder to obtain the updated two-dimensional image patch sequence, and then recombine the updated two-dimensional image patch sequence into a three-dimensional image patch sequence; Take the image patches within a row in the recombined three-dimensional image patch sequence row by row from top to bottom as the first key vector and the first value vector, and perform row encoding on the prompt word information according to the first key vector and the first value vector to obtain the row-encoded prompt word information as the row query vector; take the image patches within a column in the recombined three-dimensional image patch sequence column by column from left to right as the second key vector and the second value vector, and perform column encoding on the prompt word information according to the second key vector and the second value vector to obtain the column-encoded prompt word information as the column query vector; update the row query vector using the row cross-attention formula to obtain the updated row query vector; update the column query vector using the column cross-attention formula to obtain the updated column query vector; Decode the updated row query vector and the updated column query vector to obtain the distribution line conductor segmentation result.
[0067] Embodiment 5: If the modules / units integrated in the electronic device 100 are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, to implement all or part of the processes in the above method embodiments of the present invention, it can also be completed by a computer program instructing the relevant hardware. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, the steps of the above method embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file or some intermediate form, etc. The computer-readable medium can include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory and read-only memory (ROM, Read-Only Memory).
[0068] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memory, CD-ROM, optical memory, etc.) that contain computer-usable program code.
[0069] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to the embodiments of the present invention. It should be understood that each flow and / or block in the flowchart and / or block diagram, and the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the functions specified in one Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0070] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including instruction means that implement the functions specified in one Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0071] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process. Thus, the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0072] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the above embodiments, those of ordinary skill in the art should understand that: still can modify the specific implementation manners of the present invention or make equivalent replacements. Any modification or equivalent replacement that does not depart from the spirit and scope of the present invention shall be covered by the protection scope of the claims of the present invention.
Claims
1. A visual distribution line conductor segmentation method based on a line attention mechanism, characterized in that Including the following steps: Obtain the image information and prompt information of the distribution line conductor; Process the image information of the distribution line conductor to obtain an image block sequence; Update the information in each row of image blocks in the image block sequence row by row from top to bottom using the row self-attention formula to obtain an updated row image block sequence; update the information in each column of image blocks in the image block sequence column by column from left to right using the column self-attention formula to obtain an updated column image block sequence; combine the updated row image block sequence and the updated column image block sequence to obtain an updated three-dimensional image block sequence; Recombine the updated three-dimensional image block sequence into a two-dimensional image block sequence, update the two-dimensional image block sequence using the self-attention formula and an image encoder to obtain an updated two-dimensional image block sequence, and recombine the updated two-dimensional image block sequence into a three-dimensional image block sequence; Take the image blocks within a row in the recombined three-dimensional image block sequence row by row from top to bottom as the first key vector and the first value vector, perform row encoding on the prompt information to obtain a row query vector; take the image blocks within a column in the recombined three-dimensional image block sequence column by column from left to right as the second key vector and the second value vector, perform column encoding on the prompt information to obtain a column query vector; combine the first key vector and the first value vector and use the row cross-attention formula to update the row query vector to obtain an updated row query vector; combine the second key vector and the second value vector and use the column cross-attention formula to update the column query vector to obtain an updated column query vector; Decode the updated row query vector and the updated column query vector to obtain the segmentation result of the distribution line conductor.
2. The visual distribution line conductor segmentation method based on the line attention mechanism according to claim 1, wherein The specific method for processing the image information of the distribution line conductor to obtain an image block sequence is as follows: Standardize the image information of the distribution line conductor to obtain an image ; Perform convolution on the image to obtain a sequence of image patches ; where is the number of channels of the sequence of image patches, is the height of the image, is the width of the image, is the downsampling factor.
3. The visual distribution line conductor segmentation method based on line attention mechanism according to claim 1, characterized in that The specific method for updating the information in each row of image blocks in the image block sequence row by row from top to bottom using the row self-attention formula to obtain an updated row image block sequence is as follows: Obtain a sequence of image patches the feature vectors of all the image patches in the ; Use the row self-attention formula to update the feature vectors of image patches in all rows of the image patch sequence from top to bottom row by row in sequence, and the row self-attention formula is as follows: Obtain the updated sequence of line image blocks , is the dimension of the image block feature vector.
4. The visual distribution line conductor segmentation method based on line attention mechanism according to claim 1, characterized in that, The specific method for updating the information in each column of image blocks in the image block sequence column by column from left to right using the column self-attention formula to obtain an updated column image block sequence is as follows: Obtain a sequence of image patches the feature vectors of all image patches in the column; Use the column self-attention formula to update the image patch feature vectors of all columns in the image patch sequence from left to right column by column in turn, where the column self-attention formula is: Obtain the updated three-dimensional image patch sequence , is the dimension of the image patch feature vector.
5. The method for visually segmenting conductors of a distribution line based on a line attention mechanism according to claim 1, wherein The specific method for recombining the updated three-dimensional image block sequence into a two-dimensional image block sequence, updating the two-dimensional image block sequence using the self-attention formula and an image encoder to obtain an updated two-dimensional image block sequence, and recombining the updated two-dimensional image block sequence into a three-dimensional image block sequence is as follows: Recombine a three-dimensional image patch sequence into a two-dimensional image patch sequence , with dimensions , where , and use the self-attention formula and an image encoder to update the two-dimensional image patch sequence to obtain an updated image patch sequence ; where the self-attention formula is: 。 6. The visual distribution line conductor segmentation method based on a line attention mechanism according to claim 1, characterized in that The specific method for taking the image blocks within a row in the recombined three-dimensional image block sequence row by row from top to bottom as the first key vector and the first value vector, and performing row encoding on the prompt information to obtain a row query vector is as follows: Obtain the updated image patch sequence in the feature vectors of all image patches in the row and use them as the first key vector and the first value vector ; Encode the prompt information using the prompt encoder to obtain a feature sequence and use it as the row query vector ; Combine the first key vector and the first value vector Use the row mutual attention formula to update the row query vector to obtain the updated query vector ; The row mutual attention formula is: Among them, is the number of rows of the three-dimensional image block, is the dimension of the image block feature vector.
7. The method for visually segmenting conductors of a distribution line based on a line attention mechanism according to claim 1, characterized in that, The specific method for taking the image blocks within a column in the recombined three-dimensional image block sequence column by column from left to right as the second key vector and the second value vector, and performing column encoding on the prompt information to obtain a column query vector is as follows: Obtain the updated image patch sequence The feature vectors of all image patches in the k-th column The second key vector and the second value vector ; Encode the prompt information using the prompt encoder to obtain a feature sequence and use it as the row query vector ; Combine the second key vector and the second value vector to update the column query vector using the column mutual attention formula; the column mutual attention formula is: Among them, is the number of columns of the three-dimensional image block, is the dimension of the image block feature vector.
8. A visual distribution line conductor segmentation system based on a line attention mechanism, characterized in that, Including: A data acquisition module for obtaining the image information and prompt information of the distribution line conductor; An image processing module for processing the image information of the power distribution line conductor to obtain a sequence of image blocks; A row-column self-attention module for updating the information in each row of image blocks in the sequence of image blocks row by row from top to bottom using the row self-attention formula to obtain an updated sequence of row image blocks; updating the information in each column of image blocks in the sequence of image blocks column by column from left to right using the column self-attention formula to obtain an updated sequence of column image blocks; combining the updated sequence of row image blocks and the updated sequence of column image blocks to obtain an updated three-dimensional sequence of image blocks; A two-dimensional image block sequence updating module for reorganizing the updated three-dimensional sequence of image blocks into a two-dimensional sequence of image blocks, updating the two-dimensional sequence of image blocks using the self-attention formula and an image encoder to obtain an updated two-dimensional sequence of image blocks, and reorganizing the updated two-dimensional sequence of image blocks into a three-dimensional sequence of image blocks; A row-column cross-attention module for sequentially taking out the image blocks within a row in the reorganized three-dimensional sequence of image blocks from top to bottom as the first key vector and the first value vector to perform row encoding on the prompt word information to obtain a row query vector; sequentially taking out the image blocks within a column in the reorganized three-dimensional sequence of image blocks from left to right as the second key vector and the second value vector to perform column encoding on the prompt word information to obtain a column query vector; updating the row query vector using the row cross-attention formula by combining the first key vector and the first value vector to obtain an updated row query vector; updating the column query vector using the column cross-attention formula by combining the second key vector and the second value vector to obtain an updated column query vector; A decoding module for decoding the updated row query vector and the updated column query vector to obtain the segmentation result of the power distribution line conductor.
9. The visual distribution line conductor segmentation system based on the line attention mechanism according to claim 8, characterized in that The function of the row-column self-attention module is implemented by the following method: Obtain a sequence of image patches the feature vectors of all the image patches in the ; The row self-attention formula is used to update the feature vectors of the image patches in each row of the image patch sequence from top to bottom in units of rows in the sequence one by one. The row self-attention formula is as follows: Obtain the updated sequence of line image blocks ; Obtain a sequence of image patches the feature vectors of all image patches in the column; Adopt the column self-attention formula to update the feature vectors of the image patches in all columns of the image patch sequence from left to right column by column in sequence, where the column self-attention formula is as follows: Obtain the updated three-dimensional image block sequence .
10. The visual distribution line conductor segmentation system based on the line attention mechanism according to claim 8, characterized in that, The function of the two-dimensional image block sequence updating module is implemented by the following method: Recombine a three-dimensional image patch sequence into a two-dimensional image patch sequence , with dimensions , where , and update the two-dimensional image patch sequence using the self-attention formula and an image encoder to obtain an updated image patch sequence ; where the self-attention formula is: 。 11. The visual distribution line conductor segmentation system based on the line attention mechanism according to claim 8, characterized in that, The function of the row-column cross-attention module is implemented by the following method: Obtain the updated image patch sequence in the feature vectors of all image patches in the row and use them as the first key vector and the first value vector ; Encode the prompt information using the prompt encoder to obtain a feature sequence and use it as the row query vector ; Combine the first key vector and the first value vector Use the row mutual attention formula to update the row query vector to obtain the updated query vector ; The row mutual attention formula is: ; Obtain the updated image patch sequence The feature vectors of all image patches in the k-th column The second key vector and the second value vector ; Encode the prompt information using the prompt encoder to obtain a feature sequence and use it as the row query vector ; Combined with the adoption of the second key vector and the second value vector The column cross-attention formula is used to update the column query vector The column cross-attention formula is as follows: 。 12. An electronic device, comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the visual power distribution line conductor segmentation method based on the line attention mechanism according to any one of claims 1 to 7.
13. A storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of the visual power distribution line conductor segmentation method based on the line attention mechanism according to any one of claims 1 to 7.