Terminal block drawing detection method based on two-stage optimization and multi-stage feature enhancement

Through the methods of dual-stage optimization and multi-level feature enhancement, the problems of category imbalance and difficulty in identifying small targets in terminal strip drawing detection are solved, the detection accuracy and efficiency are improved, and the efficient identification of complex drawings is achieved.

CN120260069AActive Publication Date: 2025-07-04NANJING UNIV OF POSTS & TELECOMM
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
CN202510750636.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-06
Publication Date
2025-07-04
Estimated Expiration
2045-06-06

AI Technical Summary

Technical Problem

The existing terminal strip drawing detection methods have problems such as imbalance in the category of the primitives, difficulty in identifying small objects, serious background interference, and loss of information during feature extraction, resulting in inaccuracy and inefficiency of detection.

Method used

A method based on dual-stage optimization and multi-level feature enhancement is adopted, including sliding window slicing, interactive data synthesis, dual-focus loss function, improved YOLOv1 model and feature pyramid network, to improve the accuracy and efficiency of drawing detection.

Benefits of technology

It significantly improves the detection ability of small and medium-sized targets in the terminal strip drawing, alleviates the problem of category imbalance, enhances feature extraction ability, and improves the overall detection efficiency and accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120260069A_ABST
    Figure CN120260069A_ABST
Patent Text Reader

Abstract

The invention discloses a terminal block drawing detection method based on dual-stage optimization and multi-stage feature enhancement, which comprises the following steps: slicing a drawing by utilizing a slicing technology, and carrying out secondary processing on a sliced image through a rule-based interactive synthetic data enhancement method to obtain an enhanced sliced image data set; inputting the processed image data set into a multi-level feature enhanced target detection model for detection to obtain a model prediction category and a coordinate result; the prediction result and the real result are input into a bifocus loss function for calculation, a loss value is obtained, and a final detection model is trained through multiple rounds of iteration through a back propagation optimization model; inputting the sliced image of the detection drawing into the trained model, and fusing slicing results to obtain a final detection result; according to the method, the problem of class imbalance in the drawing is effectively solved, and through the primitive detection model, the recognition accuracy of small targets in the drawing is remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of object detection, and particularly relates to a terminal block drawing detection method based on two-stage optimization and multi-level feature enhancement. Background Art

[0002] Substations have become important hubs of intelligent and information-based power grids. As a key technical document describing the equipment connection method, line distribution, and system interaction relationship, the accuracy of drawing interpretation directly affects the operation and maintenance efficiency of equipment, the accuracy of fault troubleshooting, and the feasibility of future expansion and transformation. However, currently, it mainly relies on manual drawing interpretation, which is not only cumbersome and time-consuming but also requires high labor costs. Due to the dense drawing information, complex symbols, and intricate cable interaction relationships, manual interpretation is easily affected by subjective factors and misinterpreted, which will not only lead to high error correction costs during the construction stage but also may affect the project progress and equipment safety. With the rapid development of artificial intelligence, computer-aided drawing interpretation has been able to effectively replace manual drawing interpretation.

[0003] The core of computer-aided drawing interpretation lies in whether it can accurately locate key graphic elements in the drawing and perform precise classification to support subsequent graphic element matching operations. With the development of deep learning, existing engineering object detection models have achieved good accuracy in detecting graphic elements. However, directly using these models still cannot fully meet the requirements of drawing object detection, and there are some technical challenges, which are specifically manifested as follows: 1. The distribution of each category of graphic elements in the drawing is unbalanced in terms of the number of samples and the number of target instances, showing significant long-tail characteristics. Due to the direction sensitivity and context information continuity of the drawing, common data augmentation methods (such as copying, flipping, etc.) are difficult to apply effectively. At the same time, the terminal block drawing has strict geometric rule constraints. It is difficult to accurately ensure rule consistency when using a generative model to create graphic elements. Coupled with the scarcity of data for few-sample categories, it further increases the difficulty of training the generative model.

[0004] 2. Few-sample targets not only have a huge quantity gap with other categories but are generally small targets that are difficult to identify. Existing loss functions for difficult samples often significantly reduce the recognition accuracy of other graphic elements while improving the recognition effect of difficult targets. Therefore, it is difficult to effectively cope with this extremely unbalanced difficult target recognition challenge.

[0005] 3. Some graphic elements in the drawing are similar to the background features (such as the corners of cables and the corners of tables), and are easily interfered during detection, resulting in a decrease in recognition accuracy.

[0006] 4. In the drawing, small target primitives and other primitives each account for 50%, but there are significant size differences between the two types of primitives. The model needs to detect these targets with extremely large differences simultaneously, which not only increases the complexity of recognition but also reduces the overall detection efficiency and accuracy.

[0007] 5. After the small target primitives in the drawing undergo multiple layers of downsampling processing, their effective information is easily lost during the feature extraction process, resulting in the model being unable to accurately identify and locate small targets. Summary of the Invention

[0008] In order to avoid and overcome the technical problems existing in the prior art, the present invention provides a terminal block drawing detection method based on two-stage optimization and multi-level feature enhancement. The present invention can improve the accuracy of target detection in terminal block drawings and more accurately determine the presence and location of targets.

[0009] To achieve the above object, the present invention provides the following technical solutions: In a first aspect, a substation terminal block drawing detection method based on two-stage collaborative optimization is provided, including: Using a sliding window to perform overlapping slicing on the input substation terminal block drawing to obtain sliced images; Using a rule-based interactive data synthesis method to perform data enhancement on the sliced images, and taking the enhanced sliced images and the original sliced images together as the training data set; Inputting the class and coordinate results obtained by the model prediction of the training data set into a dual-focus loss function for calculation to obtain a loss value, optimizing the model through backpropagation, and obtaining the final optimized model through multiple rounds of iteration; Directly performing overlapping slicing on the test drawing and inputting it into the final optimized model to obtain the sliced prediction result corresponding to each slice, and obtaining the final result through a slice fusion algorithm.

[0010] Optionally, in the step of using a sliding window to perform overlapping slicing on the input substation terminal block drawing to obtain sliced images, the horizontal and vertical movement steps are calculated according to preset slice size parameters and overlapping ratios, and adjacent slices are ensured to retain a specified overlapping area by means of a sliding window; when the sliding window reaches the edge position of the image, zero-padding technology is used to complement the part exceeding the original image boundary to generate complete sliced images with uniform sizes.

[0011] Optionally, in the step of using a rule-based interactive data synthesis method to perform data enhancement on the sliced images and taking the enhanced sliced images and the original sliced images together as the training data set, the steps of the rule-based interactive data synthesis method include: Identifying the position and size of the table area in the input sliced image, as well as parameters such as the starting position of the table rows, row height, and number of rows. Create a visualization window and fine-tune the table and orientation parameters through a sliding control to ensure that the generated effect meets the actual requirements; According to the set parameters, generate primitive elements with corner structures at random positions in the specified rows of the table according to the preset ratio and geometric rules.

[0012] Optionally, input the category and coordinate results obtained by predicting the training dataset through the model into the bifocal loss function for calculation to obtain a loss value, optimize the model through backpropagation, and obtain the final optimized model through multiple rounds of iteration. In the bifocal loss function is:

[0013] where is the bounding box regression loss function of the primitive element detection model, is the classification loss function of the primitive element detection model, is the regression loss weight, is the classification loss weight.

[0014] The bounding box regression loss function of the primitive element detection model is:

[0015] where is the intersection over union (IoU) after specific adjustment, representing the degree of overlap between the predicted box and the ground truth box; The intersection over union (IoU) after specific adjustment is:

[0016] where is the adjusted IoU metric, and are distance metrics used to measure the matching degree between two rectangular boxes, is the width of the input image, is the height of the input image; The adjusted IoU metric is:

[0017] where is the central region interaction ratio, is the first hyperparameter, is the second hyperparameter; The distance metrics and are:

[0018]

[0019] Among them, and are the upper left coordinates of the ground truth box, and are the lower right coordinates of the ground truth box, and are the upper left coordinates of the predicted box, and are the lower right coordinates of the predicted box; The classification loss function of the primitive detection model is:

[0020] Among them, is the predicted probability that the model assigns the sample to the category , is the adjustment factor, is the weight coefficient of the category to which the sample n belongs.

[0021] The weight coefficient of the category to which the sample n belongs

[0022] Among them, is the total number of categories, is the number of valid samples of the category ; The number of valid samples of the category is:

[0023] Among them, is the number of samples of the category , is the decay coefficient.

[0024] After directly overlapping and slicing the test drawing and inputting it into the final optimized model to obtain the slice prediction result corresponding to each slice, the slice fusion algorithm for obtaining the final result is as follows: First, map the coordinates of the detected target bounding boxes in each slice back to the original image coordinate system, while retaining the category and confidence information, and merge the converted detection results in a unified set. At this time, the set contains the detection results from all slices. In particular, there may be a large number of duplicate detections in the slice overlap area; Next, apply the Non-Maximum Suppression (NMS) algorithm. Sort the detection boxes in descending order of confidence. Select the box with the highest confidence and add it to the result set. Calculate the Intersection over Union (IoU) between this box and the remaining detection boxes, and remove other boxes whose IoU is greater than the threshold. Repeat this process until all detection boxes are processed. Then, apply a confidence threshold filter to the results after NMS processing to remove low-confidence detections.

[0025] Finally, fine-tune the bounding box positions of the same target detected multiple times to obtain the final detection results.

[0026] In a second aspect, an improved method for a YOLOv11 object detection model based on multi-level feature enhancement is provided, including: Design a C3K2_KStar module in the shallow layers (P2 - P3 layers) of the model backbone network to obtain key details for enhancement and discrimination in the early stage of feature extraction; Design a C3K2_FKConv module in the deep layers (P4 - P5 layers) of the model backbone network to capture more subtle non-linear features; Design a DB-HSFPN module in the model neck network, use DySample to replace static upsampling and build a two-way information interaction mechanism to enhance the feature expression ability.

[0027] Optionally, in the design of the C3K2_KStar module in the shallow layers (P2 - P3 layers) of the model backbone network to obtain key details for enhancement and discrimination in the early stage of feature extraction, the C3K2_KStar module includes a C3K_KStar module and a KStarBlock module. When the c3k parameter is set to True, the C3K_KStar module is used; when the c3k parameter is set to False, the KStarBlock module is used; the C3K_KStar module is composed of multiple KStarBlock modules; In the KStarBlock, the output feature can be expressed as:

[0028] where represents the input feature, represents the depth convolution operation, represents batch normalization, represents two feature transformations.

[0029] In the KStarBlock module, the input feature X is first divided into two feature branches, which are processed by a linear transformation FC and a non-linear transformation GR-KAN respectively to obtain different feature representations, generating two feature branches F1 and F2.

[0030] Next, these two feature branches are fused through element-wise multiplication to implicitly map the input features into a high-dimensional non-linear feature space, and the output features are expressed as:

[0031] where, “ ” represents element-wise multiplication.

[0032] Finally, the output features extract spatial information through depth convolution, and the extracted features are processed through batch normalization to stabilize the feature distribution and accelerate model convergence. The final output features are passed to the next layer of the network or module.

[0033] The GR-KAN combines rational functions and group parameter sharing mechanisms and is a special type of multi-layer perceptron (MLP). It improves the model's expressive power while maintaining computational efficiency through group rational functions.

[0034] The operation of GR-KAN on the input vector x can be expressed as:

[0035] where, represents the entire GR-KAN transformation operation, represents the function composition operation, represents the number of input channels, represents the number of output channels, represents the i-th element of the input vector x, represents the number of channels in each group, calculated as , represents the number of groups represents the floor operation to determine the group index to which the input channel i belongs, is the weight connecting the i-th input channel to the j-th output channel, and F represents the rational function of the group; The rational function F of each group is defined as:

[0036] where, , , …, and , …, are learnable parameters.

[0037] In GR-KAN, the channels within the same group share the same rational function parameters and , but each input-output connection still has unique weights . Specifically, for the input channel index , its group index is , and the rational function of this group is used. It can be represented in matrix form as:

[0038] where, W represents the weight matrix, is the group rational function applied to the input; In actual implementation, GR-KAN can be represented as two consecutive operations:

[0039] where, is to apply the group rational function to the input, represents the standard linear layer. This implementation makes GR-KAN can be regarded as a special MLP, where the activation function is located before the linear layer, and the activation function is a learnable and group-shared rational function.

[0040] Optionally, the C3K2_FKConv module is designed in the deep layer (P4~P5 layers) of the model backbone network to capture more subtle non-linear features. The C3K2_FKConv module includes a C3K2_FKConv module and a Bottleneck_FKConv module. When the c3k parameter in C3K2_FKConv is set to True, the C3K2_FKConv module is used. When the c3k parameter is set to False, the Bottleneck_FKConv module is used; the C3K2_FKConv module is composed of multiple Bottleneck_FKConv modules; In the Bottleneck_FKConv module, the output feature can be represented as:

[0041] where, represents the input feature, represents the residual convolutional layer using FastKAN.

[0042] The operation of the residual convolutional layer on the input vector x can be represented as:

[0043] where, represents the basic convolutional operation, represents layer normalization, represents the radial basis function transformation non-linear transformation, represents the activation function In FKConv, each group of input features first undergoes a non - linear transformation through the SiLU activation function to enhance the expressive ability of the input features. Then, through a basic convolution operation, local features are extracted and a preliminary output is generated; Next, the preliminary output result passes through a normalization layer to stabilize the feature distribution and reduce the problem of unstable gradients during training. The normalized features are mapped into a high - dimensional spline space through the radial basis function (RBF) to generate the basis function representation required for spline convolution. At this time, the number of channels of the features is expanded to gride_size times the original; Finally, the generated basis function representation is further input into the spline convolution layer to perform a non - linear transformation on the high - dimensional features. The output of the spline convolution and the output of the basic convolution are fused by element - wise addition to form a residual link, thereby retaining the original information of the input features and enhancing the non - linear expressive ability at the same time.

[0044] The role of the RBF (radial basis function) is to map the normalized input features into a high - dimensional non - linear space, providing a basis function representation for subsequent convolution operations. The introduction of RBF significantly enhances the non - linear expressive ability of the module and expands the feature dimension, providing a basis for capturing complex feature patterns. The mathematical definition of RBF is:

[0045] where, represents the normalized input features, c represents the center point of the RBF, represents the width parameter of the RBF.

[0046] In FKConv, the specific process of RBF is as follows: First, the input is the normalized features, denoted as , where B is the batch size, is the number of channels per group, H and W are the height and width of the feature map respectively. For each position in the feature map, its feature vector is . The module generates a set of RBF center point sets , where represents the number of center points; Next, for each input feature vector x and all center points, its RBF mapping value is calculated. Through this mapping, the input features are extended to a high - dimensional space, and the number of channels of the features is expanded from to , and the specific formula is:

[0047] Finally, after the feature vectors at all positions pass through the RBF mapping, the expanded high - dimensional feature map is obtained, denoted as The feature vector of each position is represented as:

[0048] where is the RBF mapping value at the position (i, j) with the Kth center point.

[0049] Optionally, in the DB-HSFPN module designed in the neck network of the model, using DySample to replace static upsampling and constructing a two-way information interaction mechanism to enhance the feature expression ability, the DB-HSFPN module is a feature fusion pyramid network, including a channel attention module (CA), a dimension matching module (DM), a dynamic downsampling module (DySample) and a two-way feature fusion module.

[0050] In the DB-HSFPN module, first, multi-scale feature maps are extracted from the backbone network, denoted as , , , representing low-level, middle-level, and high-level features respectively. These feature maps are used as the input of the network and enter the channel attention module for screening respectively; Secondly, each input feature map passes through the channel attention module and the dimension matching module in turn to generate the screened feature , which is represented as:

[0051] The channel attention module (CA) extracts global context information through max-pooling operation and average-pooling operation, and generates channel weights through a fully connected layer and a Sigmoid function. The calculation formula is:

[0052] where is the Sigmoid activation function, is the ReLU activation function, and represent two kinds of convolution operations, represents the average-pooling operation, represents the max-pooling operation The dimension matching module (DM) reduces the channels of the features screened by the CA module to 256 through a 1×1 convolution operation to ensure that features of different scales can be effectively fused. The calculation formula is:

[0053] Then, the high-level features pass through the dynamic upsampling module , the semantic information is transmitted downward step by step and fused with the middle-layer and low-layer information, specifically manifested as follows: High-level features After the upsampling operation, it is fused with the middle-level features to obtain the middle-level features , and the specific formula is:

[0054] Middle-level features After the upsampling operation, it is fused with the low-level features to obtain the low-level features , and the specific formula is:

[0055] Through the top-down path, the high-level semantic information is transmitted to the low-level features step by step, enhancing the category recognition ability of the low-level features; Next, the low-level features pass through the downsampling operation, and the spatial and detail information is transmitted upward step by step and fused with the middle-level and high-level features, specifically manifested as follows: Low-level features After passing through the CA module, it is fused with the middle-level features to obtain the final middle-level features , and the specific formula is:

[0056] Middle-level features After passing through the CA module, it is fused with the high-level features to obtain the final high-level features , and the specific formula is:

[0057] Through the bottom-up path, the low-level spatial and edge information is transmitted to the high-level features step by step, enhancing the position information of the high-level features; Finally, after the bidirectional fusion, the multi-scale final output features can be expressed as:

[0058] In the dynamic upsampling module (DySample), the goal is to generate an upsampled feature map after the upsampling operation on the input feature map , where s is the upsampling factor, and are the height and width of the input feature map respectively. The specific process of DySample is as follows: First, construct the initial sampling grid , used to represent the coordinates of the standard sampling points. The grid positions are defined by bilinear interpolation rules to ensure uniform distribution of the sampling points. This grid is repeated in the channel dimension times and finally adjusted to the shape for use in combination with the offset, where represents the number of groups of the input features in the channel dimension; Next, the input feature map X undergoes linear projection to generate the original offset and is adjusted to the target spatial size through a pixel shuffle operation to obtain the new offset , expressed as:

[0059] where is the first weight matrix, is the first offset vector; Then, a dynamic range factor is used to control the offset, and the dynamic range factor is generated through an additional linear layer, expressed as:

[0060] where is the second weight matrix, is the second offset vector; By modulating the dynamic range factor , the final offset is obtained,

[0061] The initial sampling grid is added to the final offset to obtain the set of sampling points :

[0062] Finally, the coordinates in the set of sampling points are normalized. For each group of input feature maps , the offset and the set of sampling points are independently generated. Using the grid sampling function , values are extracted from the input feature map according to the set of sampling points to generate each group of independently upsampled feature maps , and the results of each group are concatenated to obtain the final upsampled feature map .

[0063] Compared with the prior art, the beneficial effects of the present invention are: (1) The present invention generates scarce graphic elements in the drawing by introducing a rule-based interactive data synthesis method, and proposes a dual-focus loss function, which improves the small target detection ability and alleviates the dominant position of the majority classes in the loss. Category prior information is introduced in the classification to make it closer to the current data distribution. The combination of these two stages effectively solves the problem of serious imbalance in the sample categories existing in the drawing.

[0064] (2) The present invention makes different improvements to the backbone network of the YOLOv11 graphic element detection model in the deep layer and the shallow layer, enabling the model to obtain key details in the early stage of feature extraction and capture more subtle non-linear features in the later stage. This improvement enables accurate positioning and identification of small target graphic elements in complex drawings, greatly improves the feature extraction ability, and enhances the overall detection efficiency and accuracy.

[0065] (3) The present invention proposes a new feature pyramid DB-HSFPN, which removes redundant information through high-level feature screening and retains semantic information crucial for downstream tasks. Utilizing the learnable mechanism of DySample, it enhances the ability to retain small target features. Meanwhile, the designed bidirectional fusion mechanism enables the high-level and low-level features to complement each other, improving the detail capture ability. DB-HSFPN can help the model accurately identify and locate small targets, effectively alleviating the scale difference problem. Brief Description of the Drawings

[0066] Figure 1 is the overall structural diagram of the method of the present invention; Figure 2 is the flow chart of the operation steps of the present invention; Figure 3 is the sliced example diagram of the substation terminal block drawing given by the present invention; Figure 4 is given by the present invention Figure 3 is the schematic diagram of the synthetic data enhancement result of the shown drawing Figure 5 is the schematic diagram of the network structure of the graphic element detection model of the present invention; Figure 6 is the schematic diagram of the module framework structure of C3K2_KStar of the present invention Figure 7 is the schematic diagram of the module framework structure of C3K2_FKConv of the present invention Figure 8 is the schematic diagram of the module framework structure of DB-HSFPN of the present invention Figure 9 is for the present invention Figure 4 is the schematic diagram of the recognition result of the sliced drawing shown; Detailed Embodiments The present invention will be further described below in conjunction with the accompanying drawings. The following embodiments are only used to more clearly illustrate the technical solutions of the present invention and cannot be used to limit the protection scope of the present invention.

[0067] Embodiment 1 As Figure 1 and 2 shown, a terminal block drawing detection method based on two-stage optimization and multi-level feature enhancement includes the following steps: S1: Use a sliding window to perform overlapping slicing on the input substation terminal block drawing to obtain sliced images.

[0068] Specifically, in step S1, calculate the horizontal and vertical movement steps according to the preset slicing size parameters and overlapping ratio, and ensure that adjacent slices retain a specified overlapping area through the sliding window method; when the sliding window reaches the edge position of the image, use zero-padding technology to complete the part exceeding the original image boundary, and generate complete sliced images with uniform sizes.

[0069] More specifically, the present invention sets the slicing size parameter to 1760 × 1760 pixels, define the overlapping rate as 0.3. The system first calculates the movement step as 1232 pixels according to the formula "step = slicing size × (1 - overlapping rate)"; then start from the upper left corner of the image and move horizontally and vertically in a grid pattern, moving 1232 pixels each time. For the edge area of the original image, when the remaining space is less than a complete slice, use zero-padding technology to complete the part exceeding the original image boundary. When used for detection, each slice will record its exact position in the original image according to the upper left corner coordinates for subsequent mapping of the detection results back to the original coordinate system.

[0070] S2: Use a rule-based interactive data synthesis method to perform data enhancement on the sliced images, and use the enhanced sliced images and the original sliced images together as the training dataset.

[0071] As Figure 3 and Figure 4 shown, specifically, step S2 is as follows: S21: Identify the position and size of the table area in the input sliced image, as well as parameters such as the starting position of the table rows, row height, and number of rows. More specifically, step S21 is as follows: First, convert the input image to a grayscale image, apply inverse binaryzation to make the table lines white and the background black. Then use dilation processing to enhance the connectivity of the table lines through a 3×3 rectangular structuring element, and the system detects all the contours in the image and selects the contour with the largest area as the table area. Calculate the bounding rectangle of this contour to obtain the starting coordinates (x, y) of the table, as well as the width and height.

[0072] Then, extract the region of interest (ROI) according to the detected table area, convert it to a grayscale image and perform binary processing.

[0073] Finally, calculate the horizontal projection, that is, accumulate the pixel values along the horizontal direction to obtain the pixel density distribution of each row. By analyzing this distribution, identify the boundary positions between rows in the table. When the pixel value is lower than the threshold (indicating there is content) and no row has been detected before, record it as the starting position of a row. The system calculates the difference between the starting positions of adjacent rows, takes the median as the standard row height, and calculates the total number of rows.

[0074] S22: Create a calibration window, provide multiple sliders for adjusting key parameters: the starting coordinates of the table, row height, number of rows, and the direction of adding corner lines (left or right), to ensure that the generated effect meets the actual requirements. The window is added with an image vertical cropping function to focus on a specific area of the large image for calibration. During the calibration process, draw the table row reference lines in real time, and use different colors to identify valid and invalid (beyond the image boundary) rows.

[0075] S23: According to the set parameters, generate primitive elements with corner structures at random positions in the specified rows of the table according to the preset ratio and geometric rules. For each row, the system calculates the horizontal line segment length according to the row index, so that the corner lines of different rows show an increasing relationship.

[0076] S24: Label the enhanced slice image and the original slice image, and use the two images together as the training data set.

[0077] S3: Input the processed image data set into the object detection model with multi-level feature enhancement for detection, and obtain the predicted category and coordinate results of the model.

[0078] Furthermore, based on the terminal block drawing, there are ten types of primitive elements to be recognized in the object detection process, specifically unit nameplate, end mark, upper left mark, upper right mark, upper and lower left mark, upper and lower right mark, lower right mark, lower left mark, left cable, and right cable.

[0079] As Figure 5 shown, specifically, the object detection model is improved and constructed based on the YOLOv11 model, and the specific improvements are as follows: S31: In the backbone network of the primitive element detection model, design a C3K2_Kstar module in the shallow network (P2 - P3 layers) to replace the C3K2 module in the shallow network of the original YOLOv11 model, so that the model can enhance and distinguish key details in the early stage of feature extraction, and lay a good foundation for extracting stronger features in the deep layer.

[0080] As Figure 6As shown, specifically, the C3K2_KStar module includes a C3K_KStar module and a KStarBlock module. When the c3k parameter is set to True, the C3K_KStar module is used; when the c3k parameter is set to False, the KStarBlock module is used. The C3K_KStar module is composed of multiple KStarBlock modules. In KStarBlock, the output features are described as follows:

[0081] Among them, represents the input features, represents the depth convolution operation, represents batch normalization, represents two types of feature transformations.

[0082] More specifically, the steps of the KStarBlock module are as follows: D1: In the KStarBlock module, the input feature X is first divided into two feature branches, which are processed by linear transformation FC and non-linear transformation GR-KAN respectively to obtain different feature representations, generating two feature branches F1 and F2.

[0083] Specifically, GR-KAN combines rational functions and group parameter sharing mechanisms and is a special type of multi-layer perceptron (MLP). It improves the model's expressive power while maintaining computational efficiency through group rational functions.

[0084] The operation of GR-KAN on the input vector x is described as follows:

[0085] Among them, represents the entire GR-KAN transformation operation, represents the function composition operation, represents the number of input channels, represents the number of output channels, represents the i-th element of the input vector x, represents the number of channels in each group, calculated as , represents the number of groups represents the floor operation to determine the group index to which the input channel i belongs, the weight connecting the i-th input channel to the j-th output channel, and F represents the rational function of the group; The rational function F of each group is defined as:

[0086] Among them, , , …, and , …, are learnable parameters.

[0087] In GR-KAN, channels within the same group share the same rational function parameters and , but each input-output connection still has unique weights . Specifically, for the input channel index , the group index it belongs to is , and the rational function of this group is used. Here it can be represented in matrix form as:

[0088] where W represents the weight matrix, is the group rational function applied to the input.

[0089] In actual implementation, GR-KAN can be represented as two consecutive operations:

[0090] Among them, is to apply the group rational function to the input, represents a standard linear layer. This implementation makes GR-KAN can be regarded as a special MLP, where the activation function is located before the linear layer, and the activation function is a learnable and group-shared rational function.

[0091] D2: Fuse the two feature branches and through element-wise multiplication, implicitly map the input features to a high-dimensional non-linear feature space, and the output feature is represented as:

[0092] where, " " represents element-wise multiplication. Through the fusion of the two branches, it can better suppress background noise and irrelevant information, and reduce the possibility of false detection and missed detection.

[0093] D3: The output feature extracts spatial information through depth convolution, and processes the extracted features through batch normalization to stabilize the feature distribution and accelerate the model convergence. The final output feature is passed to the next layer of the network or module.

[0094] S32: In the backbone network of the primitive detection model, in the deep network (P4 - P5 layers), the C3K2_FKConv module is designed to replace the C3K2 module in the deep network of the original YOLOv11 model. The C3K2_FKConv can improve the quality and robustness of feature extraction while retaining the efficient structure of the network, enabling the model to capture more subtle non - linear features and accurately locate and identify small targets in complex drawings.

[0095] As Figure 7 shown, specifically, the C3K2_FKConv module includes the C3K2_FKConv module and the Bottleneck_FKConv module. When the c3k parameter in the C3K2_FKConv is set to True, the C3K2_FKConv module is used; when the c3k parameter is set to False, the Bottleneck_FKConv module is used; the C3K2_FKConv module is composed of multiple Bottleneck_FKConv modules; In the Bottleneck_FKConv module, the output feature can be expressed as:

[0096] Among them, represents the input feature, represents the residual convolutional layer using FastKAN.

[0097] The said residual convolutional layer The operation on the input vector x can be expressed as:

[0098] Among them, represents the basic convolutional operation, represents layer normalization, represents the radial basis function transformation non - linear transformation, represents the activation function More specifically, the steps of the FKConv module are as follows: D1: Each group of input features first undergoes non - linear transformation through the SiLU activation function to enhance the expression ability of the input features. Then, through a basic convolutional operation, local features are extracted and preliminary outputs are generated.

[0099] D2: The preliminary output results pass through the normalization layer to stabilize the feature distribution and reduce the problem of gradient instability during training. The normalized features are mapped to a high - dimensional spline space through the radial basis function RBF to generate the basis function representation required for spline convolution. At this time, the number of channels of the features is expanded to gride_size times the original.

[0100] The role of the RBF (Radial Basis Function) is to map the normalized input features to a high-dimensional non-linear space, providing a basis function representation for subsequent convolutional operations. The introduction of RBF significantly enhances the non-linear expression ability of the module, while expanding the feature dimension, providing a basis for capturing complex feature patterns. The mathematical definition of RBF is as follows:

[0101] Among them, represents the normalized input features, c represents the center point of the RBF, represents the width parameter of the RBF.

[0102] More specifically, in FKConv, the specific steps of RBF are as follows: D21: The input is the normalized features, denoted as , where B is the batch size, is the number of channels per group, H and W are the height and width of the feature map respectively. For each position in the feature map, its feature vector is . The module will generate a set of RBF center point sets , where represents the number of center points.

[0103] D22: For each input feature vector x and all center points, calculate its RBF mapping value. Through this mapping, the input features are extended to a high-dimensional space, and the number of channels of the features is expanded from to . The specific formula is:

[0104] D23: After the feature vectors at all positions are mapped by RBF, an extended high-dimensional feature map is obtained, denoted as . The feature vector at each position is represented as:

[0105] Among them, is the RBF mapping value at the position (i,j) and the Kth center point.

[0106] D3: The generated basis function representation is further input into the spline convolutional layer to perform non-linear transformation on the high-dimensional features. The output of the spline convolution and the output of the basic convolution are fused by element-wise addition to form a residual link, thereby retaining the original information of the input features and making the extracted features more expressive.

[0107] S33: In the neck network of the primitive detection model, a new feature fusion pyramid network, the DB-HSFPN module, is designed to replace the feature fusion pyramid network in the neck of the original YOLOv11 model. DySample is used to replace the static upsampling in the original pyramid and a two-way information interaction mechanism is constructed to enhance the feature expression ability.

[0108] As Figure 8 shown, specifically, the DB-HSFPN module includes a channel attention module (CA), a dimension matching module (DM), a dynamic downsampling module (DySample), and a two-way feature fusion module, which can not only improve the small target detection performance but also alleviate the problem of detection target scale difference.

[0109] The specific steps of the DB-HSFPN module are as follows: D1: Extract multi-scale feature maps from the backbone network, denoted as , , , representing low-level, middle-level, and high-level features respectively. These feature maps are used as the input of the network and enter the channel attention module for screening respectively.

[0110] D2: Each input feature map successively passes through the channel attention module and the dimension matching module to generate the screened feature , expressed as:

[0111] More specifically, the channel attention module (CA) extracts global context information through max pooling operation and average pooling operation, and generates channel weights through a fully connected layer and a Sigmoid function. The calculation formula is:

[0112] Among them, is the Sigmoid activation function, is the ReLU activation function, and represent two kinds of convolution operations, represents the average pooling operation, represents the max pooling operation.

[0113] The dimension matching module (DM) reduces the channels of the features screened by the CA module to 256 through a 1×1 convolution operation to ensure the effective fusion of features at different scales. The calculation formula is:

[0114] D3: The high-level features pass through the dynamic upsampling module , and gradually transmit semantic information downward, fusing with the middle-level and low-level information. Specifically, it is manifested as follows: High-level features After upsampling operation, they are fused with the middle-level features to obtain the middle-level features , and the specific formula is:

[0115] Middle-level features After upsampling operation, they are fused with the low-level features to obtain the low-level features , and the specific formula is:

[0116] Through the top-down path, the high-level semantic information is gradually transmitted into the low-level features, enhancing the category recognition ability of the low-level features.

[0117] In the dynamic upsampling module (DySample), the goal is to generate an upsampled feature map after the input feature map passes through the upsampling operation, where s is the upsampling factor, and are the height and width of the input feature map respectively. DySample can enhance the retention of small target features and improve the detection recall rate through a learnable mechanism.

[0118] The specific process of DySample is as follows: D31: Construct an initial sampling grid , which is used to represent the coordinates of the standard sampling points. The grid positions are defined by the bilinear interpolation rule to ensure that the sampling points are evenly distributed. This grid is repeated times in the channel dimension and finally adjusted to the shape for use in combination with the offset, where, represents the number of groups of the input features in the channel dimension.

[0119] D32: The input feature map X generates the original offset through linear projection and is adjusted to the target spatial size through the pixel rearrangement operation (Pixel Shuffle) to obtain the new offset , which is expressed as:

[0120] where, is the first weight matrix, is the first offset vector.

[0121] D33: The offset is controlled using a dynamic range factor, and the dynamic range factor is generated through an additional linear layer , expressed as:

[0122] where is the second weight matrix, is the second offset vector

[0123] By modulating the dynamic range factor , the final offset is obtained

[0124] The initial sampling grid is added to the final offset to obtain the set of sampling points :

[0125] D34: Normalize the coordinates in the set of sampling points . For each group of input feature maps , an offset and a set of sampling points are independently generated. Using the grid sampling function , values are extracted from the input feature map according to the set of sampling points to generate each group of independently upsampled feature maps , and the results of each group are concatenated to obtain the final upsampled feature map .

[0126] D4: The low-level features pass through the downsampling operation and gradually transmit spatial and detail information upward to be fused with the middle-level and high-level features, specifically manifested as: The low-level features are fused with the middle-level features after passing through the CA module to obtain the final middle-level features , and the specific formula is:

[0127] The middle-level features are fused with the high-level features after passing through the CA module to obtain the final high-level features , and the specific formula is:

[0128] Through the bottom-up path, the low-level spatial and edge information is gradually transmitted to the high-level features, enhancing the location information of the high-level features.

[0129] D5: The multi-scale final output feature obtained after bidirectional fusion can be expressed as:

[0130] Bidirectional fusion enables the high-level and low-level features to complement each other, which can not only maintain the global semantic information but also retain local details, helping to improve the model's ability to capture details.

[0131] The primitive detection network structure of the present invention consists of three parts: the backbone network, the neck network, and the head network. In the backbone network, the first two C3k2 layers are improved to C3k2_KStar layers, and the last two C3k2 layers are improved to C3k2_FKConv layers. In the neck network, the DB-HSFPN feature pyramid network is used to form a new YOLOv8 network structure, as Figure 5 shown, the Head network structure of the present invention is exactly the same as YOLOv11.

[0132] More specifically, the primitive detection model of the present invention modifies the first two C3k2 layers in the backbone network to C3k2_KStar layers, and the last two C3k2 layers to C3k2_FKConv layers, and uses the DB-HSFPN network in the neck network, while other processing units and network structures remain unchanged; its structure mainly includes: Backbone network: This stage is responsible for extracting multi-level feature representations of the image; it consists of 5 Conv layers, 2 C3k2_KStar layers, 2 C3k2_FKConv layers, one SPPF layer, and one C2SPA layer; Neck network: This stage adopts an innovative dynamic bidirectional HSFPN structure and uses the Dysample dynamic upsampling algorithm to effectively retain detail information; this module is located between the backbone network and the head network and serves as a bridge for feature fusion and enhancement, capable of realizing bidirectional information fusion of multi-scale feature maps and enhancing the information of targets at different scales; Head network: Responsible for the final object detection and classification tasks, including a detection head for generating bounding boxes and a classification head for object recognition.

[0133] To verify the improvement effect of the module improvements proposed in the present invention on the drawing detection performance, the experiment was evaluated on a drawing slice dataset that was not involved in training. All models were trained with the same hyperparameter settings, using AdamW as the optimizer for 300 iterations. The experimental platform was an NVIDIA RTX 4090 GPU (CUDA 11.8), and the baseline model was YOLOv11.

[0134] Table 1

[0135] As shown in Table 1, all the proposed improved modules and their combinations significantly improved the drawing detection performance. Among them, the combination of three modules outperformed the baseline model in terms of accuracy, recall, mAP and other indicators, verifying the improvement effect of multi-module fusion on the detection effect.

[0136] S4: Input the class and coordinate results obtained by predicting the training dataset through the model into the bi-focal loss function for calculation to obtain the loss value, and optimize the model through backpropagation. After multiple iterations, the final optimized model is obtained.

[0137] Although the data synthesis method proposed in S2 can expand the categories with extremely small sample sizes, in order to maintain the integrity of the feature distribution of the real dataset, the data imbalance problem still objectively exists. In addition, there are generally a large number of small targets in the drawings. Due to their tiny physical sizes, it is difficult to maintain a consistent standard during the annotation process. At the same time, small targets are more sensitive to size changes, resulting in obvious instability during the bounding box regression process, significantly increasing the difficulty and complexity of model training.

[0138] To solve these problems, we propose the Bi-Focal Loss strategy to improve the model performance through two aspects of optimization: on the one hand, replace the CIoU Loss (Complete Intersection over Union Loss) in the original model with the Focaler_MPDIoU Loss (Focused Minimum Point Distance Intersection over Union Loss) to enhance the model's detection and precise positioning capabilities for small targets; on the other hand, replace the Binary Cross-Entropy Loss in the original model with the Class-Prior Guided Focal Loss to make the loss function more suitable for the current data distribution characteristics, effectively alleviating the class imbalance problem. The synergistic effect of these two loss functions not only optimizes the regression accuracy of the target bounding box but also balances the detection performance of each category, forming a more comprehensive target detection optimization framework.

[0139] Specifically, the bi-focal loss function is as follows:

[0140] Among them, is the bounding box regression loss function of the primitive detection model, is the classification loss function of the primitive detection model, is the regression loss weight, is the classification loss weight.

[0141] The bounding box regression loss function of the primitive detection model is:

[0142] Among them, is the intersection over union (IoU) after specific adjustment, representing the overlapping degree between the predicted box and the ground truth box; The intersection over union (IoU) after specific adjustment is:

[0143] Among them is the adjusted IoU metric, and is the distance metric, used to measure the matching degree between two rectangular boxes, is the width of the input image, is the height of the input image; The adjusted IoU metric is:

[0144] Among them, is the central region interaction ratio, is the first hyperparameter, is the second hyperparameter The distance metric and is:

[0145]

[0146] Among them, and are the upper left coordinates of the ground truth box, and are the lower right coordinates of the ground truth box, and are the upper left coordinates of the predicted box, and are the lower right coordinates of the predicted box; The classification loss function of the primitive detection model is as follows:

[0147] where is the predicted probability of the model for the sample belonging to the category , is the adjustment factor, is the weight coefficient of the category to which the sample n belongs.

[0148] The weight coefficient of the category to which the sample n belongs is as follows:

[0149] where is the total number of categories, is the category ; The valid sample number of the category is as follows:

[0150] where is the category , is the attenuation coefficient.

[0151] The category prediction result obtained in step S3 and the target box coordinate information are input into the Bi-Focal Loss function for loss calculation. Specifically, the category prediction is evaluated by the Class-Prior Guided Focal Loss, which dynamically adjusts the weights of each category according to the category distribution characteristics of the dataset; at the same time, the target box coordinates calculate the positioning error through the Focaler_MPDIoU Loss to improve the detection accuracy of small targets. After the system combines the two parts of the loss values, it calculates the gradient through the backpropagation algorithm and updates the model parameters. The entire training process requires multiple rounds of iteration, during which a learning rate decay strategy is adopted, and finally a target detection model with optimized performance is converged. In addition, parameters need to be set when the improved YOLOv11 model is trained. The slice detection results are as Figure 9 shown.

[0152] To verify the effectiveness of the collaborative optimization method of drawing synthesis data augmentation and dual focal loss function, experiments were conducted to compare four settings: no optimization, only using drawing synthesis data augmentation, only using dual focal loss function optimization, and the collaborative optimization of both. All experiments were based on the unoptimized YOLOv11 model and kept the parameters consistent.

[0153] Table 2

[0154] As shown in Table 2, using either synthetic data augmentation or dual focal loss function optimization alone can bring certain performance improvements. When the two are combined, the model achieves the best performance in all indicators, effectively alleviating the performance bottleneck caused by the imbalance of drawing categories.

[0155] S5: Directly perform overlapping slicing on the test drawing and input it into the final optimized model to obtain the slice prediction results corresponding to each slice, and obtain the final result through the slice fusion algorithm.

[0156] Specifically, step S5 is described as follows: S51: Map the coordinates of the detected target bounding boxes in each slice back to the original image coordinate system, while retaining the corresponding class labels and confidence scores. Through the coordinate transformation algorithm, considering the position and size ratio of the slice in the original image, accurately calculate the absolute position of each detected box in the original image. All the transformed detection results are merged into a unified set, which now contains detection results from all slices. In particular, there may be a large number of duplicate detected targets in the slice overlap area; S52: Apply the non-maximum suppression (NMS) algorithm to eliminate redundancy in the merged detection results. First, sort all the detected boxes according to the confidence from high to low, then select the box with the highest confidence to add to the final result set, and calculate its intersection over union (IoU) with the remaining detected boxes, and remove other boxes whose IoU with it is greater than the preset threshold, indicating that they may be duplicate detections of the same target. Repeat this process until all detected boxes are processed. Subsequently, apply a confidence threshold filtering mechanism to the results after NMS processing to remove detection results with a confidence lower than the set threshold, further improving the detection accuracy; S53: Fine-tune the position of the bounding box for the same target detected multiple times. For the same target that is detected multiple times in different slices and retained by NMS, based on its detection confidence and position information, use algorithms such as weighted average to finely adjust the bounding box coordinates to improve the positioning accuracy, and finally obtain a more accurate primitive detection result; On the other hand, the present invention also discloses a computer device, including a memory and a processor. The memory stores a computer program, and when the computer program is executed by the processor, the processor executes the steps of the above method.

[0157] In another embodiment provided by the present application, a computer program product including instructions is further provided. When it runs on a computer, it causes the computer to execute any one of the mobile source emission prediction methods based on temporal feature migration in the above embodiments.

[0158] It can be understood that the systems, devices, and storage media provided in the embodiments of the present invention correspond to the methods provided in the embodiments of the present invention. The explanations, examples, and beneficial effects of the relevant content can refer to the corresponding parts in the above methods.

[0159] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from a website, computer, server, or data center to another website, computer, server, or data center by wire (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that the computer can access or a data storage device such as a server or data center that includes one or more integrated available media. The available medium can be a magnetic medium (for example, a floppy disk, a hard disk, a magnetic tape), an optical medium (for example, a DVD), or a semiconductor medium (for example, a solid state disk (SSD)).

[0160] It should be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including", or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article, or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article, or device. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of additional identical elements in the process, method, article, or device including the element.

[0161] Each embodiment in this specification is described in a related manner. For the same or similar parts among the embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments. In particular, for system embodiments, since they are basically similar to method embodiments, the description is relatively simple, and reference can be made to the relevant parts of the method embodiments for the relevant content.

[0162] The above embodiments are only used to illustrate the technical solutions of the present invention, not to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A terminal block drawing detection method based on two-stage optimization and multi-level feature enhancement, characterized in that, Including: Overlappingly slice the input substation terminal strip drawing using a sliding window to obtain sliced images; Use a rule-based interactive data synthesis method to perform data augmentation on the sliced images, and use the augmented sliced images and the original sliced images together as the training dataset; Input the processed image dataset into a multi-level feature enhanced object detection model for detection to obtain the predicted class and coordinate results of the model; Input the class and coordinate results obtained by model prediction of the training dataset into a dual-focus loss function for calculation to obtain a loss value, optimize the model through backpropagation, and obtain the final optimized model through multiple rounds of iteration; Directly perform overlapping slicing on the test drawing and input it into the final optimized model to obtain the sliced prediction results corresponding to each slice, and obtain the final result through a slice fusion algorithm.

2. The terminal block drawing detection method based on two-stage optimization and multi-level feature enhancement according to claim 1, wherein The step of overlappingly slicing the input substation terminal strip drawing using a sliding window to obtain sliced images includes calculating the horizontal and vertical movement steps according to preset slice size parameters and overlapping ratios, and ensuring that adjacent slices retain a specified overlapping area through the sliding window method; When the sliding window reaches the edge position of the image, use zero-padding technology to complement the part exceeding the original image boundary to generate complete sliced images with uniform sizes.

3. The terminal block drawing detection method based on two-stage optimization and multi-level feature enhancement according to claim 1, wherein, The step of using a rule-based interactive data synthesis method to perform data augmentation on the sliced images and using the augmented sliced images and the original sliced images together as the training dataset. The steps of the rule-based interactive data synthesis method include: Identify the position and size of the table area in the input sliced image, as well as parameters such as the starting position of the table row, row height, and number of rows; Create a visualization window, and fine-tune the table and direction parameters through a sliding control to ensure that the generated effect meets the actual requirements; Generate primitives with corner structures at random positions in the specified rows of the table according to the set parameters according to a preset ratio and geometric rules.

4. The method for detecting terminal block drawings based on two-stage optimization and multi-level feature enhancement according to claim 1, characterized in that, In the step of inputting the processed image dataset into a multi-level feature enhanced object detection model for detection to obtain the predicted class and coordinate results of the model, the multi-level feature enhanced object detection model is constructed based on the YOLOv11 algorithm, including: Design a C3K2_KStar module in the shallow layer of the model backbone network, namely P2~P3 layers; Design a C3K2_FKConv module in the deep layer of the model backbone network, namely P4~P5 layers; Design a DB-HSFPN module in the model neck network, use DySample to replace static upsampling and construct a two-way information interaction mechanism; The C3K2_KStar module includes a C3K_KStar module and a KStarBlock module. When the c3k parameter is set to True, use the C3K_KStar module. When the c3k parameter is set to False, use the KStarBlock module; the C3K_KStar module is composed of multiple KStarBlock modules; In the KStarBlock, the output feature is expressed as: Among them, represents the input feature, represents the depth convolution operation, represents batch normalization, represents two kinds of feature transformations; In the KStarBlock module, the input feature X is first divided into two feature branches, which are processed by a linear transformation FC and a non-linear transformation GR-KAN respectively to obtain different feature representations, generating two feature branches F1 and F2; Next, these two feature branches are fused through element-wise multiplication to implicitly map the input features into a high-dimensional non-linear feature space, and the output features are expressed as: Among them, " " represents element-wise multiplication; Finally, output features Spatial information is extracted through depth convolution, and the extracted features are processed through batch normalization to stabilize the feature distribution and accelerate model convergence. The final output features are passed to the next network layer or module.

5. The method for detecting a terminal block drawing based on two-stage optimization and multi-level feature enhancement according to claim 4, wherein The operation of the GR-KAN on the input vector x is expressed as: Among them, represents the entire GR-KAN transformation operation, represents the function composition operation, represents the number of input channels, represents the number of output channels, represents the i-th element of the input vector x, represents the number of channels per group, calculated as , represents the number of groups represents the floor operation to determine the group index to which the input channel i belongs, the weight connecting the i-th input channel to the j-th output channel, where F represents the rational function of the group; Specifically, the GR-KAN is represented as two consecutive operations: Among them, applies a group rational function to the input, represents a standard linear layer.

6. The terminal block drawing detection method based on two-stage optimization and multi-level feature enhancement according to claim 3, characterized in that, The C3K2_FKConv module includes a C3K2_FKConv module and a Bottleneck_FKConv module. When the c3k parameter in C3K2_FKConv is set to True, the C3K2_FKConv module is used. When the c3k parameter is set to False, the Bottleneck_FKConv module is used; the C3K2_FKConv module is composed of multiple Bottleneck_FKConv modules; In the Bottleneck_FKConv module, the output features are expressed as: Among them, represents the input feature, represents the residual convolutional layer using FastKAN; The residual convolutional layer The operation on the input vector x is expressed as: Among them, represents the basic convolution operation, represents layer normalization, represents the radial basis function transformation non-linear transformation, represents the activation function; The mathematical definition of the RBF is: Among them, represents the normalized input feature, c represents the center point of the RBF, represents the width parameter of the RBF.

7. The terminal block drawing detection method based on two-stage optimization and multi-level feature enhancement according to claim 3, wherein The DB-HSFPN module is a feature fusion pyramid network, including a channel attention module CA, a dimension matching module DM, a dynamic downsampling module DySample and a bidirectional feature fusion module; The specific process of the DB-HSFPN is as follows: First, extract multi-scale feature maps from the backbone network, denoted as , , , representing low-level, middle-level, and high-level features respectively; these feature maps are used as the input of the network and enter the channel attention module for screening respectively; Secondly, each input feature map successively passes through the channel attention module and the dimension matching module to generate the filtered features , which is expressed as: The channel attention module CA extracts global context information through max pooling operation and average pooling operation, and generates channel weights through a fully connected layer and a Sigmoid function. The calculation formula is: Among them, is the Sigmoid activation function, is the ReLU activation function, and represent two kinds of convolution operations, represents the average pooling operation, represents the max pooling operation; The dimension matching module DM reduces the channels of the features filtered by the CA module to 256 through a 1×1 convolution operation. The calculation formula is: Then, the high-level features pass through the dynamic upsampling module , and the semantic information is transmitted downward level by level to fuse with the middle-level and low-level information, specifically manifested as follows: High-level features After the upsampling operation, it is combined with the middle-level features to obtain the middle-level features , and the specific formula is as follows: Middle-level features After the upsampling operation, it is combined with the low-level features to obtain the low-level features , and the specific formula is as follows: Through the top-down path, high-level semantic information is gradually transmitted to low-level features, enhancing the category recognition ability of low-level features; Next, the low-level features pass through the downsampling operation and are gradually transmitted upward with spatial and detail information to fuse with the middle-level and high-level features. Specifically, it is manifested as: Low-level features After passing through the CA module and the middle-level features Are fused to obtain the final middle-level features , and the specific formula is: Middle-level features After passing through the CA module, it is fused with the high-level features to obtain the final high-level features , and the specific formula is: Through the bottom-up path, low-level spatial and edge information is gradually transmitted to high-level features, enhancing the position information of high-level features; Finally, after two-way fusion, the multi-scale final output features obtained are expressed as: 。 8. The terminal block drawing detection method based on two-stage optimization and multi-level feature enhancement according to claim 7, wherein, The specific process of the DySample is as follows: First, construct an initial sampling grid , which is used to represent the coordinates of standard sampling points. The grid positions are defined by bilinear interpolation rules to ensure uniform distribution of sampling points. This grid is repeated times and finally adjusted to the shape for use in combination with the offsets, where represents the number of groups of the input features in the channel dimension; Next, the input feature map X generates the original offset through linear projection , and is adjusted to the target spatial dimension through the pixel rearrangement operation PixelShuffle to obtain a new offset , which is expressed as: is the first weight matrix, is the first offset vector, is the upsampling factor, and are the height and width of the input feature map respectively; Then, a dynamic range factor is used to control the offset, and the dynamic range factor is generated through an additional linear layer , expressed as: is the second weight matrix, is the second offset vector; By modulating the dynamic range factor the final offset is obtained , Add the initial sampling grid to the final offset to obtain the set of sampling points : Finally, normalize the coordinates in the set of sampling points For each group of input feature maps , independently generate an offset and a set of sampling points , use the grid sampling function , and according to the set of sampling points , extract values from the input feature map to generate each group of independently upsampled feature maps , and splice the results of each group to obtain the final upsampled feature map .

9. The terminal block drawing detection method based on two-stage optimization and multi-level feature enhancement according to claim 1, wherein, The categories and coordinate results obtained by predicting the training data set through the model are input into the bifocal loss function for calculation to obtain a loss value, and the model is optimized through backpropagation, and the final optimized model is obtained through multiple rounds of iteration. In the above process, the bifocal loss function is as follows: Among them, is the bounding box regression loss function of the primitive detection model, is the classification loss function of the primitive detection model, is the regression loss weight, is the classification loss weight; The bounding box regression loss function of the primitive detection model is as follows: Among them, is the intersection over union after specific adjustment, representing the degree of overlap between the predicted bounding box and the ground truth bounding box; The specific adjusted intersection over union is as follows: where is the adjusted IoU metric, and is the distance metric used to measure the matching degree between two bounding boxes, is the width of the input image, is the height of the input image; The adjusted IoU metric is as follows: Among them, is the central region interaction ratio, is the first hyperparameter, is the second hyperparameter; The distance metric and is as follows: Among them, and are the upper left coordinates of the ground truth box, and are the lower right coordinates of the ground truth box, and are the upper left coordinates of the predicted box, and are the lower right coordinates of the predicted box; The classification loss function of the primitive detection model is as follows: Among them, is the predicted probability of the model for the sample belonging to the category , is the adjustment factor, is the weight coefficient of the category to which the sample n belongs; The class to which the sample n belongs The weight coefficient is as follows: wherein, is the total number of categories, is the category of the number of valid samples; The said category The number of valid samples is: Among them, is the number of samples of the category , is the attenuation coefficient.

10. The terminal block drawing detection method based on two-stage optimization and multi-level feature enhancement according to claim 1, characterized in that After directly overlapping and slicing the test drawing and inputting it into the final optimization model to obtain the slice prediction result corresponding to each slice, the slice fusion algorithm is used to obtain the final result. The steps of the slice fusion algorithm are as follows: First, map the coordinates of the target bounding boxes detected in each slice back to the original image coordinate system, while retaining the category and confidence information, and merge the converted detection results in a unified set. At this time, the set contains detection results from all slices. There may be a large number of duplicate detections, especially in the slice overlap area; Next, apply the non-maximum suppression NMS algorithm to sort the detection boxes from high to low according to the confidence level, select the box with the highest confidence level to add to the result set, calculate the intersection over union IoU with the remaining detection boxes, and remove other boxes with an IoU greater than the threshold. Repeat this process until all detection boxes are processed; then apply a confidence threshold filter to the results after NMS processing to remove low-confidence detections; Finally, fine-tune the position of the bounding box for the same target detected multiple times to obtain the final detection result.

Citation Information

Patent Citations

  • Identification method and device for terminal block drawing short connecting piece primitives and storage medium

    CN115995086A

  • Spatial data integration method and system applied to digital twin cities

    CN116843845A

  • Transformer substation terminal block drawing identification method and system based on computer vision technology

    CN118038481A

  • Electronic component defect detection method based on improved YOLOv5 model

    CN118172318A

  • Sandstone microscopic image classification method and system based on improved Swin Transform

    CN118570797A