Optical cable drawing analysis method based on dynamic pre-cutting and task decoupling alignment
Through the methods of dynamic pre-cutting and task decoupling and alignment, the problem of insufficient detection performance in optical cable drawing analysis is solved, efficient target area extraction and text recognition are achieved, and detection accuracy and calculation efficiency are improved.
Patent Information
- Application Number
- CN202510766718.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-10
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2045-06-10
AI Technical Summary
Existing object detection algorithms are difficult to meet high-precision requirements in optical cable drawing analysis tasks, especially when dealing with specific tasks in complex scenarios, detection performance and memory usage are insufficient, and there is a lack of explicit task decoupling strategies. The differences between classification and positioning tasks are not fully paid attention to.
The optical cable drawing analysis method based on dynamic pre-cutting and task decoupling and alignment is adopted to remove interference information through dynamic cropping, and combine K-Means weak edge perception and Fourier transform decoupling detection head to design a hybrid frequency domain decoupling structure and task alignment structure to optimize detection accuracy.
It improves the calculation efficiency and detection accuracy of optical cable drawing analysis, automatically filters background content, enhances the feature learning ability of classification and positioning tasks, alleviates the contradiction between classification and positioning, and improves the final detection accuracy.
Smart Images

Figure CN120279343A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer vision, and specifically relates to an image preprocessing method, a design of an object detection head, and a cable drawing parsing solution. Background Art
[0002] In the field of computer vision, object detection has always been a key and difficult point of research. With the rapid development of deep learning technology, object detection algorithms based on convolutional neural networks have made remarkable progress in terms of accuracy and efficiency. However, in practical industrial applications, especially in specific tasks in complex scenarios, there are still many challenges. As a typical application scenario in industrial inspection, the core task of cable drawing parsing is to automatically extract key information, such as text information and equipment positions, from cable design drawings, and combine text recognition technology to automatically extract the content of the drawings, reducing the manual burden.
[0003] Existing object detection algorithms, such as Faster R-CNN, YOLO series, etc., although performing excellently in general object detection tasks, often have difficulty meeting actual requirements when dealing with scenarios such as cable drawings with specific structures and semantics. Cable drawings usually contain a large number of lines, symbols, and text information. There are complex spatial relationships and semantic associations among these elements. However, most of this information does not need to be parsed into a standard data format, and the rich element content may instead interfere with the final parsing result. In addition, cable drawings have high resolution, large scale, and dense detection targets, which pose higher requirements for the detection performance and memory occupancy of the algorithm.
[0004] In terms of image preprocessing, traditional preprocessing methods mainly focus on operations such as image enhancement, denoising, and normalization to improve the accuracy of subsequent object detection. In the cable drawing parsing task, considering the design and characteristics of the drawings comprehensively, it is also necessary to clearly cut out the parsing target area and eliminate redundant interference information in the preprocessing stage.
[0005] In the task of optical cable drawing parsing, we have relatively high requirements for the accuracy of object detection because the detection quality directly determines the subsequent text recognition and parsing matching effects. The design of the object detection head is one of the cores in the object detection algorithm, and its performance directly affects the accuracy and efficiency of the detection results. Existing object detection heads usually use decoupled heads to predict objects of different sizes and shapes. However, in mainstream detection models such as the YOLO series, there is a lack of a more explicit task decoupling strategy in detection, and the differences between the classification and localization tasks have not been fully considered. On the other hand, when the model output results are post-processed, the predicted confidence only considers the classification score and lacks consideration of the localization quality. Some detection boxes with high localization quality may be discarded during NMS. The task-aligned detection head aligns the classification and localization information, which can alleviate the contradiction between classification and localization, is conducive to obtaining high-quality detection boxes, and improves the drawing detection accuracy. Therefore, an efficient object detection head alignment structure is exactly what the task needs. Summary of the Invention
[0006] The present invention provides an optical cable drawing parsing method based on dynamic pre-clipping and task decoupling alignment, aiming to efficiently obtain the target areas for detection and recognition in the optical cable drawing, eliminate interference content, and at the same time propose a new decoupling-alignment detection head to improve the detection and parsing accuracy.
[0007] To achieve the above object, the technical solution adopted by the present invention is as follows: An optical cable drawing parsing method based on dynamic pre-clipping and task decoupling alignment includes the following steps: Step 1, obtain the original substation drawing data. The drawing data is specifically divided into two categories: optical cable drawings and terminal block drawing slices. The additional introduction of terminal block drawing slices is to supplement data and improve the generalization ability of the training model in text detection.
[0008] Step 2, analyze the characteristics of the optical cable drawing data, and perform preprocessing according to the prior information of the drawing, including prior clipping of interference information, and using dynamic clipping processing based on K-Means weak edge perception to obtain the sheared target area.
[0009] Step 2.1, prior clipping of interference information. The truly effective information area in the optical cable drawing is a rectangular table, which is called the target area; the entire contour composed of reference marks around the drawing is called the reference mark area; the remaining area is the redundant area. Except for the rectangle formed by the reference mark area, the target area has the largest outer contour area on the drawing. The structure of the reference mark area is complex, and the contour detection effect is difficult to predict. To exclude the interference brought by the complex structure, the reference mark area is dynamically clipped according to the horizontal auxiliary line drawn from the midpoints of the left and right edges of the drawing. Assume the original image is , with the size of , and the search step , pixel threshold , the upper limit of the auxiliary line length is 500. The horizontal position of the left-side cutting is determined by the following formula (1): (1); where is the horizontal coordinate of the left-side cutting, is the number of search steps, represents the pixel value of the image at the height and the horizontal position . represents finding the minimum that meets the conditions from all possible sets. The horizontal position of the right-side cutting is determined by the following formula (2): (2); where is the horizontal coordinate of the right-side cutting, represents finding the maximum that meets the conditions from all possible sets. Cut the drawing and retain the part with the horizontal coordinate from to to obtain .
[0010] Step 2.2, under the condition of many weak edges, a dynamic cutting method based on K-Means weak edge perception is designed. By analyzing the gradient amplitude, the Canny edge detection threshold is dynamically adjusted, and the target area is selected and cut according to the edge area. To prevent abnormal breakage of the target area contour, the image is smoothed with a large Gaussian kernel before contour detection. The calculation of the Gaussian kernel is shown in formulas (3), (4), and (5): (3); (4); (5); where k = 11, is the normalized Gaussian convolution kernel. Gaussian smoothing is shown in formula (6): (6); Use Canny detection to cut the target area. To improve the calculation efficiency and dynamically set the double threshold of Canny detection, the following method is adopted: First, sample the image pixels, then consider the weak edge characteristics, use the Mini Batch K-Means algorithm to cluster the sampled gradients, and finally dynamically set the high and low thresholds of Canny detection according to the clustering results. Assume that the total number of image pixels is , the number of sampled pixels is , calculate the pixel gradient amplitude , and sample obtained from a data . Two Sobel operators used to calculate the gradient are shown in formulas (7) and (8): (7); (8); After convolving the smoothed image with the operator, the gradient and direction are further obtained, as shown in formulas (9) and (10): (9); (10); where and are the gradient amplitudes in the x - direction and y - direction obtained after convolving the image with and respectively. Next, Mini Batch K - Means algorithm is used to cluster , and the loss function to be optimized is shown in formula (11): (11); where represents the number of clusters, is the center of the th cluster, and represents the Euclidean distance. Randomly select a batch from , batch = 500, and assign each gradient amplitude to the nearest cluster center, as shown in formula (12): (12); Then update the cluster center, as shown in formula (13): (13); where is the set of pixel gradients assigned to the th cluster. Continuously select and repeat the above steps until the maximum number of iterations is reached or the convergence condition is met, as shown in formula (14): (14); where is the maximum number of iterations, is the minimum change in cluster center, and are the cluster centers of the th iteration and the The cluster centers of the current iteration. Considering that the convolution of large Gaussian kernels will greatly blur the drawing content and increase the number of weak edges, sufficient attention should be paid to it. Therefore, set k = 3 to divide the non-edge gradient centers , weak edge gradient centers and strong edge gradient centers . Set the low gradient threshold low_threshold of Canny detection to the right boundary of the corresponding cluster, and the high gradient threshold high_threshold to the right boundary of the corresponding cluster. According to the gradient magnitude and direction obtained from formulas (9) and (10), perform non-maximum suppression on the gradient magnitude, and then judge the gradient magnitude information around the weak edges according to the double gradient thresholds to divide them into non-edges or edges. After completing all the above processes, a binary contour image is obtained.
[0011] Step 2.3, crop the target area. Use the cv2.findContours function in the python library to process the binary image. For all the contours, only keep the contour endpoints and detect the outer contour. Next, calculate the area of each contour, and the target area is the area of the contour with the largest area. The area is calculated as shown in formula (15): (15); where are the coordinates of consecutive points on the contour, is the number of vertices on the contour.
[0012] Step 3, taking the YOLOv11 model as the baseline, modify the object detection head and introduce a hybrid frequency domain decoupling structure.
[0013] Step 3.1, use Fourier transform to decouple the input features of the detection head. Separate the high and low frequency information through Fourier transform: the high frequency information is used to enhance the classification ability and capture the detailed features of the target; the low frequency information is used to improve the localization accuracy and capture the overall structure of the target. Define as the training batch, as the number of channels, and perform a two-dimensional discrete Fourier transform on the input features of the detection head. The two-dimensional discrete Fourier transform is shown in formula (16): (16); where are the spatial domain coordinates, are the frequency domain coordinates, is the feature value of the input feature map at batch , channel , position , and is the result in the frequency domain after Fourier transform at batch 、 Channel 、 Frequency The complex value at. After processing the frequency domain information using the frequency shift function torch.fft.fftshift in the python library, we get . Then create a mask , as shown in formula (17): (17); Where , , and control the size of the low-frequency region. After processing the frequency domain information using , shift the frequency again to restore the frequency domain order, and use the inverse Fourier transform to restore the high and low frequency features in the spatial domain, as shown in formulas (18), (19), (20): (18); (19); (20); In the formula is the low-frequency feature, is the high-frequency feature, is the torch.fft.fftshift method in python, is the inverse Fourier transform.
[0014] Step 3.2, fuse the decoupled features and the original features. The method of fusing the decoupled features and the original features is shown in formulas (21) (22): (21); (22); In the formula is the feature of the localization branch after fusion, is the feature of the classification branch after fusion, and are learnable parameters, learning the mixing ratio of the frequency domain features and the original features from the model.
[0015] Step 4, propose a task alignment structure based on the collection and distribution mechanism. The present invention designs a task alignment structure to ensure the unity of the objectives of the classification and localization tasks, rather than being relatively isolated. First, restore the number of channels of the classification and localization tasks, then fuse the feature information of different tasks through the collection and distribution of task features, and finally re-inject the information into the two task branches.
[0016] Step 4.1, restore the task channels. The classification and localization features have different numbers of channels. Use a 1*1 convolutional kernel to change the feature channels to obtain and 。 The number of channels of is , where is the number of channels of the input feature map of the detection head, and is the detection category.
[0017] Step 4.2, collect and distribute task features. First, initially fuse the features of the two tasks to obtain , as shown in Equation (23): (23); where is a 1×1 convolution, and (24); where is the new feature extraction information, is a depthwise separable convolution of 3×3, with the input channels being , and the output channels being D. Its specific implementation is as shown in Equation (25): (25); where is the value of the output feature map of the depthwise separable convolution at position and channel , is the pointwise convolution weight (equivalent to a 1×1 convolution kernel with one input channel and output channel D), is the weight of the depth convolution, . After that, perform information interaction between channels to obtain , and split it into the localization distribution feature and the classification distribution feature , as shown in Equations (26), (27), and (28): (26); (27); (28); Step 4.3, re-inject the information into the two task branches to obtain and . The injection process is as shown in Equations (29) and (30): (29); (30); where is the Sigmoid function.
[0018] Thus, the construction of the new detection head is completed.
[0019] Step 5: For the text targets in the object detection results, use text recognition technology to parse the content of the text string.
[0020] Step 6: Combine the graphic element position information and text position information in object detection, as well as the text recognition content, parse the drawing semantics according to electrical engineering expertise, and reject recognition for possible abnormal information.
[0021] Step 6.1: Drawing semantics parsing under the background of electrical engineering expertise. Although we have determined the target area, there are still interfering texts within the area. In object detection, the detected graphic elements include texts, device locators, and optical cable numbers. The text contains all the character information to be parsed, and the device locator and optical cable number assist in obtaining the spatial position of the parsed text. The semantics parsing strategy is as follows: Determine all text search intervals through the device locator and optical cable number graphic elements. Then, according to the corresponding relationship between the starting device and the ending device in the device locator, match the text content within the starting search interval and the ending search interval. Finally, organize the number information within the starting device name, ending device name, and optical cable number, and output to excel.
[0022] Step 6.2: Reject abnormal information. There are naturally differences and design anomalies in drawing design, and the parsing results cannot be guaranteed to be perfect. The present invention designs a rejection function, which can give a rejection warning for possible problems such as text overlap and information omission.
[0023] In summary, through the collaborative design of dynamic pre-cutting and task decoupling-alignment detection head, the present invention has achieved significant optimization of computational efficiency and detection accuracy in the optical cable drawing parsing task. The pre-cutting mechanism adopts a dynamic strategy to automatically filter out irrelevant background content in the image preprocessing stage, and automatically adjusts the contour detection parameters according to the image feature information to cut out the key target area. This technology solves the problem of manually adjusting parameters in traditional contour detection methods, optimizes the memory occupancy during training, and provides higher-quality input data for subsequent processes. The decoupling-alignment detection head proposed by the present invention adopts a dual-branch structure, uses the frequency-domain decomposition method to independently optimize the classification and localization tasks, enhances the feature learning ability of each task, and avoids feature interference. At the same time, the task alignment strategy based on the collection and distribution mechanism effectively alleviates the contradiction between classification and localization, unifies the prediction targets of tasks, and improves the final detection accuracy. Finally, the present invention combines the background of electrical professional knowledge, automatically matches the text content according to the specified parsing strategy and adds a rejection function, which has high reliability and versatility. Description of the Drawings
[0024] Figure 1 is the flowchart of the optical cable drawing recognition and parsing method of the present invention.
[0025] Figure 2 is the flowchart of the dynamic pre-cutting of the present invention.
[0026] Figure 3 is the model structure diagram of the decoupling-alignment detection head of the present invention.
[0027] Figure 4 is the target area diagram of the substation optical cable drawing given by the present invention.
[0028] Figure 5 is the present invention for Figure 4 The schematic diagram of the parsing result of the target area of the shown drawing. Detailed Embodiment
[0029] Next, the feasible embodiments of the present invention will be described in conjunction with the drawings.
[0030] As Figure 1 shown, the optical cable drawing parsing method based on dynamic pre-cutting and task decoupling alignment of the present invention includes the following steps: Step 1, obtain substation drawing data, and the drawing data includes optical cable drawing and terminal block drawing slices; Step 2, analyze the characteristics of the optical cable drawing data, perform preprocessing according to the prior information of the drawing, perform prior cutting on the interference information, and use dynamic cutting based on K-Means weak edge perception to obtain the target area; Step 3: When using the detector to locate the target area information, a hybrid frequency-domain decoupling structure is introduced at the head of the detector to fully consider the feature requirements of different tasks and enhance their respective task branches. Step 4: For different task branches, a task alignment structure based on a collection and distribution mechanism is designed to ensure the unity of the objectives of the classification and localization tasks. Step 5: Apply the improved detector to the primitive detection of the optical cable drawings and identify the text content. Step 6: Perform semantic parsing of the drawings under the background of electrical professional knowledge on the text recognition results, reject possible abnormal information, and output the drawing parsing results.
[0031] The specific implementation steps are as follows: S1. Obtain the substation drawing dataset. On the one hand, collect the optical cable drawing data in all substation drawings, which will be used as the targets for detection and parsing after preprocessing. On the other hand, considering that the vast majority of the text content in the target area of the optical cable drawings is single and repetitive and not the content that needs to be parsed; while the form of the target text content to be parsed is relatively complex but the quantity is small, so consider using the text data in other types of drawings to supplement the dataset to improve the generalization ability of text detection. The specific supplementation method is to select the terminal block drawings similar in form to the optical cable drawings and use the method of manual screening to cut out the drawing slices that conform to the text features in the target area of the optical cable drawings. The slice size is fixed at 1760*1760.
[0032] S2, as Figure 2 shown, preprocess the optical cable drawings based on prior knowledge.
[0033] Specifically, it is divided into the following steps: S2.1. Perform prior cropping on the interference information. Analyze all the optical cable drawing data. The real effective information area is a rectangular table, which is called the target area; the entire outline composed of reference marks around the drawing is called the reference mark area; the rest of the area is the redundant area. Except for the rectangle formed by the reference mark area, the target area has the largest outer contour area on the drawing. The structure of the reference mark area is complex and the contour detection effect is difficult to predict. Therefore, crop the reference mark areas on the left and right sides of the drawing so that the target area must have the largest outer contour area on the cropped drawing. The cropping method dynamically searches for the cropping edge according to the horizontal auxiliary line drawn from the midpoints of the left and right edges of the drawing. We set the original image as , with the size of , the search step , the pixel value threshold , and the upper limit of the auxiliary line length is 500. The horizontal position of the left cropping is determined by the following formula (1): (1); where is the horizontal coordinate of the left cropping, is the number of search steps, represents the pixel value of the image at height and horizontal position . represents finding the minimum satisfying condition from all possible sets. The horizontal position of the right cropping is determined by the following formula (2): (2); where is the horizontal coordinate of the right cropping, represents finding the maximum satisfying condition from all possible sets. Finally, crop the drawing, only retaining the part of the image with horizontal coordinates from to to obtain .
[0034] S2.2. A dynamic cropping method based on K-Means weak edge perception is designed. By analyzing the gradient magnitude, the Canny edge detection threshold is dynamically adjusted, and the target area is selected and cropped according to the edge area. A large Gaussian kernel is used to smooth the image to prevent abnormal breaks in the contour of the target area. The calculation of the Gaussian kernel is shown in formulas (3), (4), and (5): (3); (4); (5); where k = 11, is the normalized Gaussian convolution kernel. Gaussian smoothing is shown in formula (6): (6); Use Canny contour detection to crop the target area. To improve the calculation efficiency and dynamically set the double thresholds of Canny detection, the following method is adopted: sample the image gradient magnitude, then consider the weak edge characteristics, use the Mini Batch K-Means algorithm to cluster the gradient magnitude, and finally dynamically set the high and low thresholds of Canny detection according to the clustering results. Assume that the total number of pixels in the image is , the number of sampled pixels is , calculate the pixel gradient magnitude , and sample data to obtain . The two Sobel operators used to calculate the gradient are shown in formulas (7) and (8): (7); (8); After performing a convolution operation on the smoothed image with the operator, the gradient magnitude is further obtained and direction , as shown in formulas (9) and (10): (9); (10); where and are the x-direction gradient and y-direction gradient magnitudes obtained after convolving the image separately with and respectively. Next, perform clustering on using the Mini Batch K-Means algorithm. The loss function to be optimized is as shown in formula (11): (11); where represents the number of clusters, is the center of the th cluster, represents the Euclidean distance. Randomly select a batch from , batch = 500, and assign each gradient magnitude to the nearest cluster center, as shown in formula (12): (12); Then update the cluster center, as shown in formula (13): (13); where is the set of pixel gradients assigned to the th cluster. Continuously select and repeat the above steps until the maximum number of iterations is reached or the convergence condition is satisfied, as shown in formula (14): (14); where is the maximum number of iterations, is the minimum change in the cluster center, and are the cluster centers of the th iteration and the cluster centers of the th iteration respectively. Considering that convolving with a large Gaussian kernel will greatly blur the drawing content and increase the weak edges, sufficient attention should be paid to it. Therefore, set k = 3 to divide the non-edge gradient centers , weak-edge gradient centers and the center of the strong edge gradient . Set the low gradient threshold low_threshold of Canny detection to the right boundary of the corresponding clustering cluster, and the high gradient threshold high_threshold to the right boundary of the corresponding clustering cluster. According to the gradient magnitude and direction obtained by formulas (9) and (10), perform non-maximum suppression on the gradient magnitude to initially obtain contour information. Then, according to the double gradient threshold, set the pixels lower than low_threshold as non-edge pixels, the pixels higher than high_threshold as edge pixels, and judge the gradient magnitude of the surrounding area of the pixels whose gradient magnitude is between low_threshold and high_threshold. If there is at least one strong edge with a gradient magnitude higher than high_threshold in its surroundings, retain the pixel as an edge pixel; otherwise, mark it as a non-edge pixel. After completing all the above processes, a binary contour image is obtained.
[0035] S2.3. Cut the target area according to the contour detection result. For the binary contour image, use the cv2.findContours function in the python library to process. For all contours, only retain the endpoints and only detect the outer contours. Calculate the area of each contour, and according to the prior knowledge, take the contour with the largest area as the target area contour. Calculate the area contour Adopt the Shoelace theorem, as shown in formula (15): (15); where are the coordinates of consecutive points on the contour, is the number of vertices on the contour.
[0036] S3. Modify the YOLOv11 object detection head, introduce a hybrid frequency-domain decoupling structure, introduce the feature requirements of different tasks, and enhance their respective task branches. The model structure diagram is as Figure 3 shown.
[0037] S3.1. Use Fourier transform to decouple the input features of the detection head. The YOLOv11 detection head decouples the shared features of the localization and classification tasks, but does not fully consider the differences between the classification and localization tasks. The classification task needs to capture the representative features of the target, such as texture, edges, and contours, and this information is usually concentrated in the high-frequency components. The localization task depends on detecting the external contour and overall structure of the target, and generally this information is usually concentrated in the low-frequency components. By decomposing the input features into high-frequency and low-frequency components through Fourier transform, they can be used for the classification and localization tasks respectively: the high-frequency information is used to enhance the classification ability and capture the detailed features of the target; the low-frequency information is used to improve the localization accuracy and capture the overall structure of the target. For the input features of the detection head , where is the training batch, is the number of channels, and a two-dimensional discrete Fourier transform is performed on its last two dimensions. The two-dimensional discrete Fourier transform is shown in Equation (16): (16); where are the spatial coordinates, are the frequency domain coordinates, is the eigenvalue of the input feature map at batch , channel , position . is the complex value at batch , channel , frequency in the frequency domain after Fourier transform. After shifting the frequency domain information using torch.fft.fftshift in the python library, is obtained. Then, a mask is created to separate the low-frequency and high-frequency information. The creation of the mask is shown in Equation (17): (17); where , , and control the size of the low-frequency region. After multiplying and with the frequency domain information of respectively, and then shifting the frequency back to restore the frequency domain order, the high-frequency and low-frequency features are obtained using the inverse Fourier transform, as shown in Equations (18), (19), and (20): (18); (19); (20); where is the low-frequency feature, is the high-frequency feature, is the torch.fft.fftshift method in python, is the inverse Fourier transform.
[0038] S3.2, Fusing the decoupled features and the original features. The method of fusing the decoupled features and the original features is shown in Equations (21) and (22): (21); (22); In the formula is the localization branch feature after fusion, is the classification branch feature after fusion, and are learnable parameters used to automatically adjust the mixing ratio of frequency domain features and original features.
[0039] S4. Propose a task alignment structure based on the collection and distribution mechanism. The alignment structure ensures the information interaction between the classification and localization tasks, guarantees the consistency of classification and localization information, avoids the problem of mismatch between the target location and class prediction, and thus improves the detection accuracy. First, restore the number of channels of the localization feature and the classification feature, retain the task differences, then collect and distribute the task features, fuse the feature information of different tasks, and finally re-inject the information into the two task branches to complete the task alignment.
[0040] S4.1, Restore the task channels. The classification and localization features have different numbers of channels. Use two 1×1 convolutional kernels to convolve and respectively to change the feature channels and obtain and . Among them, has the number of channels , is the number of channels of the input feature map of the detection head, is the discretization parameter for the localization output; has the number of channels , is the detected category.
[0041] S4.2, Collect and distribute the task features. Refer to the neck structure of the GOLD-YOLO model and introduce the collection and distribution mechanism into the alignment of the classification and localization tasks. First, preliminarily fuse the features of the two tasks to obtain , as shown in formula (23): (23); Among them, is a 1×1 convolution with the output channel being fusion_size, concatenates the feature maps by channel. Next, further feature extraction is performed, as shown in formula (24): (24); Among them, is the result of further extracting information from the fused features, is a 3×3 depthwise separable convolution with the input channel number and the output channel number D being fusion_size. Its specific implementation is shown in formula (25): (25); where is the value of the output feature map of depthwise separable convolution at position and channel , is the pointwise convolution weight (equivalent to a 1×1 convolution kernel with one input channel and output channel D), is the weight of depthwise convolution, . After that, the information between channels is interacted to obtain , which is split into the localization distribution feature and the classification distribution feature , as shown in formulas (26), (27), and (28): (26); (27); (28); where the number of channels is fusion_size, and have the same number of channels as and respectively.
[0042] S4.3. Re-inject the information into the two task branches to obtain and . The injection process is shown in formulas (29) and (30): (29); (30); where is the Sigmoid function.
[0043] Thus, the feature extraction for the classification and localization tasks is completed, and the alignment of the two is effectively considered. After the detection head is constructed, the object detection model can be trained according to the dataset.
[0044] S5. For the text objects in the object detection task, use the CRNN text recognition technology to perform end-to-end string content parsing on the detected text regions.
[0045] S6. According to the graphic element position information and text position information in the object detection, perform drawing semantic parsing and reject recognition for possible abnormal information.
[0046] S6.1, Semantic parsing of drawings in the context of electrical engineering expertise. Although we have identified the target area, there are still interfering texts within the area. In object detection, the detected primitive elements include texts, device locators, and optical cable numbers. Texts contain all the character information to be parsed, and device locators and optical cable numbers assist in obtaining the spatial positions of the parsed texts. Next, the search and matching strategy for parsing is given: (1) Determine the horizontal coordinate range of all search targets according to the coordinates of the device locators.
[0047] (2) Determine the vertical coordinate range of all search targets according to the coordinates of the optical cable numbers.
[0048] (3) Assume there are 2 device locators (there are only two device locators on the optical cable drawing: the starting device locator and the ending device locator) and N optical cable numbers. Since the horizontal and vertical ranges have been determined, it is possible to divide search intervals, among which there are N starting search intervals and N ending search intervals. All parsing targets are located within these search intervals.
[0049] (4) Search for all texts within each search interval and identify which texts belong to the starting search interval and which texts belong to the ending search interval.
[0050] (5) According to the coordinates of each starting search interval, match the coordinates of the uniquely corresponding ending search interval.
[0051] (6) Match the text contents within the corresponding starting search interval and ending search interval. By using the characteristics of the horizontal and vertical coordinates of these texts, the name of the ending device corresponding to the starting device can be determined.
[0052] (7) Organize the name of the starting device, the name of the ending device, and the number information within the optical cable number primitive elements, standardize the data format, and obtain the output excel table.
[0053] S6.2, Rejection of abnormal information. There are naturally differences and design anomalies in the drawing design, and it is difficult to ensure that the parsing results are completely correct. The present invention designs a rejection function and gives a rejection warning for problems such as possible text overlap and information omission. The rejection scheme is as follows: (1) If there are parentheses in the name of the starting device or the ending device, ensure the integrity of the parentheses pair. If the parentheses are incomplete, there may be a problem of text adhesion to symbols, and a warning is marked for the parsing result of this excel item.
[0054] (2)According to the statistics of the drawing information, the end device name generally consists of two strings. If the end device name in the parsing result consists of only one string, it indicates that text may have been missed during inspection, or the drawing design method is different. Mark a warning for the parsing result of this excel entry.
[0055] (3)There may be a situation where the optical cable number graphic elements are missed during inspection. The rejection strategy is to check whether all the optical cable number texts and the optical cable number graphic elements match one by one.
[0056] For Figure 4 the final parsing result of the target area of the optical cable drawing is as Figure 5 shown. The present invention can automatically crop the target area and perform target detection on the text and graphic elements within the target area. According to the corresponding relationship between the start device, the end device, and the optical cable number, match the text recognition content and standardize and output it to an excel table. If the above abnormal parsing result appears, use a special color to fill the problem content, and the rejected abnormal information needs to be manually reviewed.
[0057] Those skilled in the art can clearly understand that the embodiments of the present invention can be implemented through computer programs and corresponding general hardware platforms. From this understanding, the key technical part of the embodiments of the present invention, in essence or what can be said to contribute to the prior art, can be embodied in the form of a computer program, that is, a software product. This computer program or software product can be stored in a storage medium and includes multiple instructions to drive a device (such as a personal computer, a server, a single-chip microcomputer, an embedded microcontroller (MCU), or a network device, etc.) containing a data processing unit to execute the methods described in different embodiments or certain parts of the embodiments of the present invention.
[0058] The present invention provides an optical cable drawing parsing method based on dynamic pre-cropping and task decoupling alignment. There are various methods and ways to implement this technical solution, and the above is only one of the specific implementation manners of the present invention. It should be noted that for professionals in this technical field, without departing from the principle of the present invention, various improvements and optimizations can be made, and these improvements and optimizations should also be regarded as the protection scope of the present invention. Each component not clearly described in this embodiment can be implemented using existing technologies.
Claims
1. A method for parsing optical cable drawings based on dynamic pre-cutting and task decoupling alignment, characterized in that It includes the following steps: Step 1: Obtain the substation drawing data, where the drawing data includes the optical cable drawing and the terminal block drawing slices; Step 2: Analyze the characteristics of the optical cable drawing data, perform preprocessing according to the prior information of the drawing, perform prior cropping on the interference information, and use dynamic cropping based on K-Means weak edge perception to obtain the target area; Step 3: When using the detector to locate the target area information, introduce a hybrid frequency domain decoupling structure at the head of the detector, fully consider the characteristic requirements of different tasks, and enhance their respective task branches; Step 4: For different task branches, design a task alignment structure based on the collection and distribution mechanism to ensure the unity of the goals of the classification and positioning tasks; Step 5: Apply the improved detector to the primitive detection of the optical cable drawing and identify the text content; Step 6: Perform drawing semantic parsing on the text recognition results under the background of electrical professional knowledge, reject abnormal information, and output the drawing parsing results.
2. The optical cable drawing analysis method based on dynamic pre-cutting and task decoupling alignment according to claim 1, wherein Step 2 includes the following steps: Step 2.1: Prior cropping of interference information; The effective information area in the optical cable drawing is a rectangular table, which is called the target area; the entire outline composed of reference marks around the drawing is called the reference mark area; the remaining area is the redundant area; Except for the rectangle formed by the reference mark area, the target area has the largest outer contour area on the drawing; To exclude the interference caused by complex structures, the reference mark area is dynamically cropped according to the horizontal auxiliary lines drawn from the midpoints of the left and right edges of the drawing; the original image is set as , with a size of , a search step , a pixel threshold , and the upper limit of the auxiliary line length is 500; the horizontal position of the left-side crop is determined by the following formula (1): (1); Among them is the horizontal coordinate of the left-side cutting is the number of search steps represents the pixel value of the image at height and horizontal position ; represents finding the smallest one that meets the conditions from all sets ; The horizontal position of the right-side cropping is determined by the following formula (2): (2); Among them is the horizontal coordinate of the right-side cutting represents finding the largest one that meets the conditions from all sets; cutting the drawing and retaining the part with the horizontal coordinate from to to to obtain ; Step 2.2: Under the condition of increasing weak edges, design a dynamic cropping method based on K-Means weak edge perception, dynamically adjust the Canny edge detection threshold by analyzing the gradient amplitude, and select and crop the target area according to the edge area; Step 2.3, cut the target area; use the cv2.findContours function in the python library to process the binary image. For all contours, only keep the contour endpoints and detect the outer contour; next, calculate the area of each contour, and the target area is the area of the contour with the largest area. The area is calculated as shown in formula (15): (15); wherein are the coordinates of consecutive points on the contour, is the number of vertices on the contour.
3. The optical cable drawing parsing method based on dynamic pre-cutting and task decoupling alignment according to claim 2, wherein, Step 2.2 specifically includes: Smooth the image with a large Gaussian kernel before contour detection; The calculation of the Gaussian kernel is shown in formulas (3), (4), and (5): (3); (4); (5); where k = 11, is a standardized Gaussian convolution kernel; Gaussian smoothing is shown in formula (6): (6); Next, sample the image pixels. Then, considering the weak edge characteristics, use the Mini Batch K-Means algorithm to cluster the sampled gradients. Finally, dynamically set the high and low thresholds of the Canny edge detection according to the clustering results. Assume that the total number of image pixels is , and the number of sampled pixels is . Calculate the pixel gradient magnitude , and sample data to obtain . The two Sobel operators used to calculate the gradient are shown in formulas (7) and (8): (7); (8); After convolving the smoothed image with the operator, the amplitude is further obtained and the direction , as shown in formulas (9) and (10): (9); (10); Among them and are the x-direction gradient and the y-direction gradient magnitude obtained after the image I is convolved with and respectively; for using the Mini Batch K-Means algorithm for clustering, the loss function to be optimized is shown in formula (11): (11); Among them represents the number of clusters, is the center of the th cluster, represents the Euclidean distance; randomly select a batch from , and assign each gradient magnitude to the nearest cluster center, as shown in formula (12): (12); Then update the cluster center, as shown in formula (13): (13); where is the set of pixel gradients assigned to the th cluster; continuously select and repeat the above steps until the maximum number of iterations is reached or the convergence condition is satisfied, as shown in Equation (14): (14); Among them is the maximum number of iterations, is the minimum change in the cluster center, and are respectively the cluster center of the th iteration and the cluster center of the th iteration; Set the number of clusters k = 3 to divide non-edge gradient centers , weak-edge gradient centers and strong-edge gradient centers ; Set the low gradient threshold low_threshold for Canny detection to be the right boundary of the corresponding cluster, and the high gradient threshold high_threshold to be the right boundary of the corresponding cluster; Based on the gradient magnitude and direction obtained from formulas (9) and (10), perform non-maximum suppression on the gradient magnitude to initially obtain contour information, and then, according to the double gradient thresholds, judge the gradient magnitude information around the weak edges and classify them as non-edges or edges.
4. A method for parsing an optical cable drawing based on dynamic pre-cutting and task decoupling alignment according to claim 3, characterized in that Step 3 includes the following steps: Step 3.1: Use Fourier transform to decouple the input features of the detection head; separate high-frequency and low-frequency information through Fourier transform: high-frequency information is used to enhance the classification ability and capture the detailed features of the target; low-frequency information is used to improve the positioning accuracy and capture the overall structure of the target; Definition is the training batch, is the number of channels, and for the input feature of the detection head a two-dimensional discrete Fourier transform is performed. The two-dimensional discrete Fourier transform is shown in formula (16): (16); Among them is the spatial domain coordinate, is the frequency domain coordinate, is the eigenvalue of the input feature map at batch , channel , position ; is the complex value at batch , channel , frequency after Fourier transform in the frequency domain; is obtained by processing the frequency domain information using the frequency shift function torch.fft.fftshift in the python library ; Then create a mask , as shown in formula (17): (17); Among them , , and control the size of the low-frequency region; After using to process the frequency-domain information, shift the frequency again to restore the frequency-domain order, and use the inverse Fourier transform to restore the high and low frequency characteristics in the spatial domain, as shown in formulas (18), (19), and (20): (18); (19); (20); where is the low-frequency feature, is the high-frequency feature, is the torch.fft.fftshift method in Python, is the inverse Fourier transform; Step 3.2, fuse the decoupled features and the original features; through learnable parameters and , automatically learn the feature ratio relationship, as shown in formulas (21) and (22): (21); (22); wherein is the fused localization branch feature, is the fused classification branch feature.
5. A method for parsing optical cable drawings based on dynamic pre-cutting and task decoupling alignment according to claim 4, characterized in that Step 4 includes the following steps: Step 4.1, restore the task channels; the classification and localization features have different numbers of channels. Use two 1*1 convolutional kernels to perform convolution on and respectively to change the feature channels, obtaining and ; Step 4.2, collect and distribute task features; First, initially fuse the features of the two tasks to obtain , as shown in formula (23): (23); Among them is a 1*1 convolution concatenates the feature maps by channel Next, further feature extraction, as shown in formula (24): (24); Among them is the new feature extraction information is a 3*3 depthwise separable convolution, and the input channels are , the number of output channels is D, and its specific implementation is as shown in formula (25): (25); Among them is the value of the output feature map of depthwise separable convolution at position and channel . is the pointwise convolution weight is the weight of depthwise convolution ; then, for inter-channel information interaction is performed to obtain , which is split into the localization distribution feature and the classification distribution feature as shown in formulas (26), (27), and (28): (26); (27); (28); Step 4.3, re-inject the information into the two task branches to obtain and ; The injection process is shown in formulas (29) and (30) as follows: (29); (30); Among them is the Sigmoid function.
6. A method for parsing optical cable drawings based on dynamic pre-cutting and task decoupling alignment according to claim 5, characterized in that, Step 5 includes the following steps: Step 5.1: Drawing semantic parsing; According to electrical professional knowledge, the defined detection primitives include text, device locators, and optical cable numbers; Text recognition obtains all the parsed text content, and the device locators and optical cable numbers assist in obtaining the spatial positions of the parsed text; The search strategy is: Determine all text search ranges according to the device locators and optical cable number primitives; Then, according to the matching relationship between the starting device and the ending device in the device locator, match the text content within the corresponding starting search range and ending search range; Finally, sort out the number information in the starting device name, ending device name, and optical cable number, and output to excel; Step 5.2: Reject abnormal information; For existing abnormal problems, give a rejection warning.
7. A method for parsing an optical cable drawing based on dynamic pre-cutting and task decoupling alignment according to claim 6, characterized in that The abnormal information rejected in Step 5.2 includes the following steps: (1) If there are parentheses in the names of the starting device and the ending device, ensure the integrity of the parenthesis pairs; if the parentheses are incomplete, there will be a problem of text-symbols adhesion, and mark a warning for the corresponding excel parsing result; (2) According to the statistics of the drawing information, the name of the ending device generally consists of two strings; if the name of the ending device in the parsing result consists of only one string, it indicates that text has been missed during inspection, or the drawing design method is different, and mark a warning for the excel parsing result; (3) There is a situation where the optical cable numbering graphic elements are missed during inspection; the rejection strategy is to check whether all the optical cable numbering texts and optical cable numbering graphic elements match one by one.
Citation Information
Patent Citations
Method for detecting weak edges of images on basis of discharge information of multilayer neuron groups
CN103679710A
Transformer substation optical cable diagram intelligent analysis method and device based on diffusion model, and storage medium
CN118037710A
Image instance detection segmentation model construction method based on edge information enhancement
CN118762042A