Weak supervision target detection method and system for engineering drawings
By optimizing model parameters through weakly supervised learning and the cIOU evaluation indicator, the problems of lack of professional knowledge and high annotation costs in engineering drawing inspection are solved, and the inspection accuracy and efficiency are improved.
Patent Information
- Application Number
- CN202510526458.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-25
- Publication Date
- 2025-09-19
AI Technical Summary
Existing deep learning-based object detection technology in the field of engineering drawings suffers from a lack of professional knowledge and high annotation costs, resulting in low detection accuracy and difficulty in achieving optimal performance.
A weakly supervised learning method is adopted. By selecting the ABCCAD dataset for annotation, image channel transformation and splitting are performed, a saliency prediction module and edge detection network are constructed, and the cIOU evaluation indicator is introduced to optimize the model parameters and reduce the dependence on accurately labeled data.
The accuracy and efficiency of target detection in engineering drawings are improved, the model's ability to parse engineering drawings is enhanced, and model performance is optimized.
Smart Images

Figure CN120673020A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer vision technology, and in particular to a weakly supervised target detection method and system for engineering drawings. Background Art
[0002] Artificial intelligence and deep learning, the driving forces of this era, are capable of deeply modeling real-world problems by leveraging vast amounts of data. They have already demonstrated remarkable potential across a wide range of fields. These technologies are not only flourishing in the scientific and technological fields but have also gradually permeated every aspect of our daily lives. This is particularly true in the field of computer vision, where intuitive display methods enable computers to easily recognize images, videos, and other content, much like the human visual system. This technology has widespread applications in medical diagnosis, security monitoring, intelligent transportation, and other fields, making it a current research hotspot in the fields of artificial intelligence and deep learning.
[0003] However, existing deep learning-based target detection technologies still have some shortcomings. First, most methods require a large amount of accurately labeled data, which is particularly difficult in the field of engineering drawings because the lack of professional knowledge and the high cost of labeling limit the scale and quality of labeled data; second, existing target detection models often ignore the use of structural information and detail features unique to engineering drawings, resulting in low detection accuracy in complex backgrounds; in addition, existing methods lack effective evaluation indicators when optimizing model parameters, making it difficult to achieve optimal performance in specific application scenarios. To address these problems, the present invention proposes a weakly supervised target detection method for engineering drawings, which reduces dependence on accurately labeled data through weakly supervised learning, and optimizes model parameters by introducing the cIOU evaluation indicator to improve the accuracy and efficiency of target detection in engineering drawings. Summary of the Invention
[0004] In view of the above-mentioned existing problems, the present invention provides a weakly supervised target detection method and system for engineering drawings, which is used to solve the problems of lack of professional knowledge and high cost of labeling work in the existing technology, low detection accuracy and difficulty in achieving optimal performance.
[0005] To solve the above technical problems, a weakly supervised object detection method for engineering drawings is proposed, including: Select an engineering drawing dataset, annotate the drawing data, preprocess the annotated dataset, and decompose the drawings into small blocks; build a training and testing model, divide the training network into blocks, modify the network structure and group convolution, perform supervised training, and optimize the parameter model parameters; comprehensively process the engineering drawing images to be tested and input them into the trained network for prediction, verify the changes in model performance under different parameters, and further improve and optimize the algorithm.
[0006] As a preferred solution of the weakly supervised target detection method for engineering drawings described in the present invention, the labeling of the drawing data includes selecting an ABCCAD dataset from the engineering drawing dataset, and using a graffiti tool to label each engineering drawing in the ABCCAD dataset to mark the target area to be detected.
[0007] As a preferred solution of the weakly supervised target detection method for engineering drawings described in the present invention, the preprocessing includes performing image channel transformation and splitting processing on the labeled data set, adjusting the quality and clarity of the image, verifying whether the preprocessed image retains key features and structural information, and grouping the preprocessed sub-images into a data set.
[0008] The image channel conversion includes reading the red, green and blue channel values of the color image, calculating the grayscale value of the image using a weighted average formula, assigning the calculated grayscale value to each pixel to generate a grayscale image, and adjusting the contrast and brightness of the converted grayscale image.
[0009] The weighted average formula is: , Among them, R is the red channel value, G is the green channel value, B is the blue channel value, and Gray is the converted grayscale value.
[0010] The splitting process includes analyzing the structure of the drawing, determining a cutting plan, cutting along the determined cutting line using a cutting tool, dividing the drawing into sub-images, and saving each sub-image as an independent file.
[0011] As a preferred solution of the weakly supervised target detection method for engineering drawings described in the present invention, the construction of the training and testing model includes constructing a saliency prediction module, constructing an edge detection network, and constructing a saliency refinement module.
[0012] The construction of the saliency prediction module includes initializing the VGG16 network structure, deleting the fifth pooling layer and subsequent layers, building a front-end saliency prediction network, and grouping the remaining convolutional layers. The ASPP module is connected after the fifth convolutional layer to capture contextual information. The output of the ASPP is feature-fused through two consecutive 1×1 convolutions to generate a preliminary saliency map.
[0013] The edge detection network is constructed by using a feature map generated by an intermediate layer of a saliency prediction module, calculating gradients in the horizontal and vertical directions using a Sobel operator, obtaining gradient amplitudes using the gradient calculations in the horizontal and vertical directions, constructing a channel edge map, training the edge detection network using a cross entropy loss function, and minimizing the loss function.
[0014] The formula for obtaining the gradient amplitude is: , in, is the center pixel coordinate, is the horizontal gradient, is the vertical gradient.
[0015] The cross entropy loss function is expressed as: , in, is the cross entropy loss function, M is the number of edge pixels, is the real edge annotation, is the edge graph predicted by the network, and s is the variable index.
[0016] As a preferred solution of the weakly supervised object detection method for engineering drawings described in the present invention, the construction of the training and testing model also includes, during the training process, using the cross-entropy loss of the part with graffiti annotations as a supervision signal, introducing an edge-aware loss function within the salient area, and suppressing pixels outside the salient area.
[0017] The supervisory signal is expressed as: , Among them, K is the set of pixels in the graffiti annotation area, A real salient map of graffiti annotations, To minimize the partial cross entropy loss function, is a rough saliency map, and k is the variable index.
[0018] The edge-aware loss function is expressed as: , in, is the edge weight matrix, is the final saliency map, is the initial saliency map, is the adjustment coefficient, a and b are variable indices, is the edge-aware loss function.
[0019] As a preferred solution of the weakly supervised target detection method for engineering drawings described in the present invention, the optimization parameter model parameters include defining the objective function, optimizing the edge detection network performance, and evaluating the indicator cIOU.
[0020] Defining the objective function includes initializing the weights and biases of the edge detection network, calculating and integrating the loss function, calculating the gradient according to the objective function, and updating the network parameters using gradient descent until the network converges to a preset number of training iterations.
[0021] The objective function is expressed as: , in, To minimize the partial cross entropy loss function, is the edge-aware loss function, is the cross entropy loss function, is the objective function, To minimize the weight coefficient of the partial cross entropy loss function, is the weight coefficient of the edge-aware loss function, is the weight coefficient of the cross entropy loss function.
[0022] The evaluation indicator cIOU includes calculating the intersection area and union area of the predicted saliency map and the true saliency map to obtain the intersection-union ratio, calculating the distance between the center point coordinates of the predicted box and the true box, converting the distance between the center point coordinates into a penalty term, calculating the aspect ratio penalty term, and combining the IOU, center distance penalty term, and aspect ratio penalty term to calculate cIOU.
[0023] The formula for obtaining the intersection-over-union ratio is: , Among them, I is the intersection area of the predicted box and the real box, U is the union area of the predicted box and the real box, It is the intersection and comparison.
[0024] The distance formula for calculating the center point coordinates is: , in, is the center coordinate of the prediction box, is the center coordinate of the real box, and d is the Euclidean distance between the center point of the predicted box and the real box.
[0025] The formula for converting to penalty term is: , in, is the center distance penalty term, c is the normalization factor related to the size of the predicted box and the true box, and d is the Euclidean distance between the center point of the predicted box and the true box.
[0026] The formula for calculating the aspect ratio penalty term is: , in, is the aspect ratio penalty, is the width of the prediction box, is the height of the prediction box, is the width of the real frame, is the height of the actual frame.
[0027] The formula for calculating cIOU is: , in, To improve the intersection-over-union ratio, is the intersection and union ratio, is the center distance penalty term, is the aspect ratio penalty, and is the weight coefficient.
[0028] As a preferred solution of the weakly supervised target detection method for engineering drawings described in the present invention, the prediction includes comprehensively processing the engineering drawing image to be detected and inputting it into the trained network for prediction verification. After each training cycle, the generated significance prediction map is analyzed using the cIOU evaluation index, and the evaluation index threshold is set. When the cIOU is greater than or equal to the evaluation index threshold, it is proved that the model is not the optimal model and the parameters need to be changed. The modified model is evaluated again until it reaches the optimal standard; when the cIOU is less than the evaluation index threshold, it is proved that the model is the optimal model and there is no need to modify the parameters. The changes in model performance under different parameters are verified, the algorithm is optimized, and the optimal parameter value is determined.
[0029] Another object of the present invention is to provide a weakly supervised target detection system for engineering drawings. The present invention effectively improves the performance of the model under different parameter settings. The system of the present invention reduces the dependence on precisely labeled data through weakly supervised learning, and optimizes the model parameters by introducing the cIOU evaluation indicator, thereby improving the accuracy and efficiency of target detection in engineering drawings. The model's ability to parse engineering drawings is improved through block processing and network structure optimization. In addition, by combining partial cross-entropy loss, edge-aware loss, and standard cross-entropy loss, the method of the present invention can effectively improve the performance of the model under different parameter settings, providing a new technical path for automatic parsing and target detection of engineering drawings.
[0030] As a preferred solution of the weakly supervised target detection system for engineering drawings described in the present invention, it is characterized by including a data preparation and preprocessing module, a network construction and initialization module, a loss function definition and optimization module, a model training and verification module, and a performance evaluation module.
[0031] The data preparation and preprocessing module is used to select an engineering drawing data set, annotate the drawing data, perform image channel transformation and splitting processing, adjust image quality, verify the retention of key features and structural information, and group the preprocessed sub-images into a data set.
[0032] The network construction and initialization module is used to construct a saliency prediction module, an edge detection network and a saliency refinement module, initialize the network structure, modify the network structure and perform group convolution, access the ASPP module to capture context information, and perform feature fusion.
[0033] The loss function definition and optimization module is used to define the objective function, optimize the edge detection network performance, calculate and integrate the loss function, calculate the gradient according to the objective function, and update the network parameters using gradient descent.
[0034] The model training and verification module is used to perform supervised training and optimize model parameters. It uses the cross-entropy loss of the part with graffiti annotations as the supervision signal, introduces the edge-aware loss function, suppresses pixels in non-salient areas, and analyzes the generated saliency prediction map using the cIOU evaluation metric after each training cycle.
[0035] The performance evaluation module is used to comprehensively process the engineering drawing images to be inspected, input the trained network for prediction verification, calculate cIOU, verify the changes in model performance under different parameters, optimize the algorithm, and determine the optimal parameter values.
[0036] A computer device includes a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, the steps of a method for weakly supervised target detection for engineering drawings are implemented.
[0037] A computer-readable storage medium stores a computer program, which, when executed by a processor, implements the steps of a method for weakly supervised object detection for engineering drawings.
[0038] The beneficial effects of the present invention are as follows: the present invention effectively suppresses redundant structural information outside the salient area by imposing a specific constraint mechanism in the salient area, thereby enhancing the focus and accuracy of the salient map; and by introducing an edge-aware loss function, successfully strengthens the pixels in the salient area, while suppressing the pixels in the non-salient area, thereby improving the overall detection performance of the model; by defining the objective function and the evaluation indicator cIOU, it also achieves precise optimization of the model performance, ensuring that the model can maintain the optimal state under different parameter settings. BRIEF DESCRIPTION OF THE DRAWINGS
[0039] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0040] Figure 1 An overall flow chart of a weakly supervised object detection method for engineering drawings provided by one embodiment of the present invention.
[0041] Figure 2 A system solution flow chart of a weakly supervised object detection system for engineering drawings provided by one embodiment of the present invention. DETAILED DESCRIPTION
[0042] To make the above-mentioned objects, features, and advantages of the present invention more clearly understood, the following detailed description of the specific embodiments of the present invention is given in conjunction with the accompanying drawings. It is obvious that the described embodiments are only part of the embodiments of the present invention, not all of them. Based on the embodiments of the present invention, all other embodiments obtained by ordinary persons in this field without creative work should fall within the scope of protection of the present invention.
[0043] In the following description, many specific details are set forth to facilitate a full understanding of the present invention. However, the present invention may also be implemented in other ways different from those described herein. Those skilled in the art may make similar generalizations without violating the connotation of the present invention. Therefore, the present invention is not limited to the specific embodiments disclosed below.
[0044] Secondly, the term "one embodiment" or "embodiment" herein refers to a specific feature, structure, or characteristic that may be included in at least one implementation of the present invention. The phrase "in one embodiment" appearing in various places throughout this specification does not necessarily refer to the same embodiment, nor does it individually or selectively refer to an embodiment that is mutually exclusive of other embodiments.
[0045] The present invention is described in detail with reference to schematic diagrams. For ease of illustration, cross-sectional views of device structures may be partially enlarged and not to scale when describing embodiments of the present invention. Furthermore, the schematic diagrams are merely illustrative and should not limit the scope of the present invention. Furthermore, in actual production, the three-dimensional dimensions of length, width, and depth should be included.
[0046] In the description of the present invention, it should be noted that the terms "upper, lower, inner, and outer" and other references to orientations or positional relationships are based on the orientations or positional relationships shown in the accompanying drawings and are intended solely to facilitate and simplify the description of the present invention. They are not intended to indicate or imply that the devices or components referred to must have, be constructed, or operate in a specific orientation, and therefore should not be construed as limitations on the present invention. Furthermore, the terms "first, second, or third" are used for descriptive purposes only and should not be construed as indicating or implying relative importance.
[0047] In this disclosure, unless otherwise specified or limited, the terms "mounted," "connected," and "connected" should be interpreted broadly. For example, they may refer to fixed, removable, or integral connections. They may also refer to mechanical, electrical, or direct connections, indirect connections through an intermediary, or internal communication between two components. Those skilled in the art will understand the specific meanings of these terms in this disclosure.
[0048] Example 1, with reference to Figure 1 , which is the first embodiment of the present invention, provides a weakly supervised object detection method for engineering drawings, comprising: S1: Select an engineering drawing dataset, annotate the drawing data, preprocess the annotated dataset, and decompose the drawings into small blocks.
[0049] The marking of the drawing data includes selecting an ABCCAD dataset from the engineering drawing dataset, and marking each engineering drawing in the ABCCAD dataset using a graffiti tool to mark a target area that needs to be detected.
[0050] It should be noted that the preprocessing includes performing image channel transformation and splitting processing on the labeled data set, adjusting the quality and clarity of the image, verifying whether the preprocessed image retains key features and structural information, and combining the preprocessed sub-images into a data set; The image channel conversion includes reading the red, green and blue channel values of the color image, calculating the grayscale value of the image using a weighted average formula, assigning the calculated grayscale value to each pixel to generate a grayscale image, and adjusting the contrast and brightness of the converted grayscale image; The weighted average formula is: , Among them, R is the red channel value, G is the green channel value, B is the blue channel value, and Gray is the grayscale value after conversion; The splitting process includes analyzing the structure of the drawing, determining a cutting plan, cutting along the determined cutting line using a cutting tool, dividing the drawing into sub-images, and saving each sub-image as an independent file.
[0051] S2: Build a training and testing model, divide the training network into blocks, modify the network structure and group convolution, perform supervised training, and optimize parameter model parameters.
[0052] Furthermore, the construction of the training and testing model includes constructing a saliency prediction module, constructing an edge detection network, and constructing a saliency refinement module; The construction of the saliency prediction module includes initializing the VGG16 network structure and deleting the fifth pooling layer and subsequent layers, constructing a front-end saliency prediction network, and grouping the remaining convolutional layers. The ASPP module is connected after the fifth convolutional layer to capture context information. The output of the ASPP is subjected to feature fusion through two consecutive 1×1 convolutions to generate a preliminary saliency map. The edge detection network is constructed by using a feature map generated by an intermediate layer of a saliency prediction module, calculating horizontal and vertical gradients using a Sobel operator, obtaining gradient amplitudes using the horizontal and vertical gradient calculations, constructing a channel edge map, training the edge detection network using a cross entropy loss function, and minimizing the loss function; The formula for calculating the horizontal gradient is: , The formula for calculating the vertical gradient is: , in, is the center pixel coordinate, is the horizontal gradient, is the vertical gradient, m is the horizontal offset of the domain pixel relative to the center pixel, n is the vertical offset of the domain pixel relative to the center pixel, and F is the vertical gradient of the domain pixel. The pixel value within the local window centered at The formula for obtaining the gradient amplitude is: , in, is the center pixel coordinate, is the horizontal gradient, is the vertical gradient; The cross entropy loss function is expressed as: , in, is the cross entropy loss function, M is the number of edge pixels, is the real edge annotation, is the edge graph predicted by the network, and s is the variable index.
[0053] Furthermore, the constructing of the training and testing model further includes, during the training process, using the cross entropy loss of the portion annotated with graffiti as a supervisory signal, introducing an edge-aware loss function within the salient region, and suppressing pixels outside the salient region; The supervisory signal is expressed as: , Among them, K is the set of pixels in the graffiti annotation area, A real salient map of graffiti annotations, To minimize the partial cross entropy loss function, is a rough saliency map, k is the variable index; The edge-aware loss function is expressed as: , in, is the edge weight matrix, is the final saliency map, is the initial saliency map, is the adjustment coefficient, a and b are variable indices, is the edge-aware loss function.
[0054] S3: The engineering drawing images to be inspected are fully processed and input into the trained network for prediction, verifying the changes in model performance under different parameters and further improving and optimizing the algorithm.
[0055] Furthermore, the optimization parameter model parameters include defining an objective function, optimizing edge detection network performance, and evaluating the indicator cIOU; Defining the objective function includes initializing the weights and biases of the edge detection network, calculating and integrating the loss function, calculating the gradient according to the objective function, and updating the network parameters using gradient descent until the network converges to a preset number of training iterations; The objective function is expressed as: , in, To minimize the partial cross entropy loss function, is the edge-aware loss function, is the cross entropy loss function, is the objective function, To minimize the weight coefficient of the partial cross entropy loss function, is the weight coefficient of the edge-aware loss function, is the weight coefficient of the cross entropy loss function; The evaluation indicator cIOU includes calculating the intersection area and union area of the predicted saliency map and the true saliency map to obtain the intersection-over-union ratio, calculating the distance between the center point coordinates of the predicted box and the true box, converting the distance between the center point coordinates into a penalty term, calculating the aspect ratio penalty term, and combining the IOU, center distance penalty term, and aspect ratio penalty term to calculate cIOU; The formula for obtaining the intersection-over-union ratio is: , Among them, I is the intersection area of the predicted box and the real box, U is the union area of the predicted box and the real box, For intersection and comparison; The distance formula for calculating the center point coordinates is: , in, is the center coordinate of the prediction box, is the center coordinate of the real box, and d is the Euclidean distance between the center point of the predicted box and the real box; The formula for converting to penalty term is: , in, is the center distance penalty term, c is the normalization factor related to the size of the predicted box and the true box, and d is the Euclidean distance between the center point of the predicted box and the true box; The formula for calculating the aspect ratio penalty term is: , in, is the aspect ratio penalty, is the width of the prediction box, is the height of the prediction box, is the width of the real frame, is the height of the real frame; The formula for calculating cIOU is: , in, To improve the intersection-over-union ratio, is the intersection and union ratio, is the center distance penalty term, is the aspect ratio penalty, and is the weight coefficient.
[0056] Furthermore, the prediction includes dividing the target training sample set into a ratio of 8:2 to obtain a training set and a validation set, initializing the modified VGG16 network, randomly sampling the convolutional layer weights from a normal distribution of (0, 0.01), sampling the bias term from (0, 0.1), initializing Xavier to 0, using data augmentation techniques including random cropping, rotation, and scaling, processing the training set image data, inputting the processed image into the improved VGG16 network, obtaining the intermediate layer feature map of the saliency prediction module, and converting the intermediate layer feature map into the saliency prediction module. Input the edge monitoring network, construct the edge map, and calculate the loss based on the cross-entropy loss function to optimize the edge detection accuracy. The rough saliency map output by the edge detection network is input into the saliency refinement module, and the loss is calculated based on the partial cross-entropy loss. The saliency map is supervised and refined, and a constraint mechanism is imposed within the salient area. The edge-aware loss function is introduced to enhance the pixels in the salient area and suppress the pixels in the non-salient area. The comprehensive loss is calculated based on the objective function, the model parameters are updated, and each module is optimized. The engineering drawing image to be detected is fully processed and input into the trained network for prediction and verification.
[0057] Set the evaluation index threshold. When cIOU is greater than or equal to the evaluation index threshold, it proves that the model is not the optimal model and needs to change the parameters. The modified model is evaluated again until it reaches the optimal standard. When cIOU is less than the evaluation index threshold, it proves that the model is the optimal model and no parameter modification is required. The changes in model performance under different parameters are verified, the algorithm is optimized, and the optimal parameter values are determined.
[0058] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit the present invention. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that the technical solutions of the present invention may be modified or replaced by equivalents without departing from the spirit and scope of the technical solutions of the present invention, which should all be included in the scope of the claims of the present invention.
[0059] Example 2, reference Figure 2 , which is the second embodiment of the present invention, provides a weakly supervised target detection system for engineering drawings, including a data preparation and preprocessing module, a network construction and initialization module, a loss function definition and optimization module, a model training and verification module, and a performance evaluation module.
[0060] The data preparation and preprocessing module is used to select an engineering drawing data set, annotate the drawing data, perform image channel transformation and splitting processing, adjust image quality, verify the retention of key features and structural information, and group the preprocessed sub-images into a data set.
[0061] The network construction and initialization module is used to construct a saliency prediction module, an edge detection network and a saliency refinement module, initialize the network structure, modify the network structure and perform group convolution, access the ASPP module to capture context information, and perform feature fusion.
[0062] The loss function definition and optimization module is used to define the objective function, optimize the edge detection network performance, calculate and integrate the loss function, calculate the gradient according to the objective function, and update the network parameters using gradient descent.
[0063] The model training and verification module is used to perform supervised training and optimize model parameters. It uses the cross-entropy loss of the part with graffiti annotations as the supervision signal, introduces the edge-aware loss function, suppresses pixels in non-salient areas, and analyzes the generated saliency prediction map using the cIOU evaluation metric after each training cycle.
[0064] The performance evaluation module is used to comprehensively process the engineering drawing images to be inspected, input the trained network for prediction verification, calculate cIOU, verify the changes in model performance under different parameters, optimize the algorithm, and determine the optimal parameter values.
[0065] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit the present invention. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that the technical solutions of the present invention may be modified or replaced by equivalents without departing from the spirit and scope of the technical solutions of the present invention, which should all be included in the scope of the claims of the present invention.
[0066] Embodiment 3, the third embodiment of the present invention, is different from the first two embodiments in that: If the functions are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the portion that contributes to the prior art, or the portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage media include various media that can store program code, such as USB flash drives, mobile hard drives, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical disks.
[0067] The logic and / or steps represented in the flowcharts or otherwise described herein, for example, can be considered as a sequenced list of executable instructions for implementing the logical functions, and can be embodied in any computer-readable medium for use by, or in conjunction with, an instruction execution system, apparatus, or device (e.g., a computer-based system, a system including a processor, or other system that can fetch and execute instructions from an instruction execution system, apparatus, or device). For purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by, or in conjunction with, an instruction execution system, apparatus, or device.
[0068] More specific examples (a non-exhaustive list) of computer-readable media include the following: an electrical connection with one or more wires (electronic devices), a portable computer disk cartridge (magnetic devices), a random access memory (RAM), a read-only memory (ROM), an erasable and programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable compact disc read-only memory (CDROM). In addition, the computer-readable medium may even be paper or other suitable medium on which the program is printed, since the program may be obtained electronically, for example, by optically scanning the paper or other medium, and then editing, interpreting, or processing in another suitable manner as necessary, and then storing it in a computer memory.
[0069] It should be understood that various components of the present invention may be implemented using hardware, software, firmware, or a combination thereof. In the above-described embodiments, multiple steps or methods may be implemented using software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented using hardware, as in another embodiment, any one of the following technologies known in the art or a combination thereof may be used: a discrete logic circuit having logic gate circuits for implementing logic functions on data signals, an application-specific integrated circuit having suitable combinational logic gate circuits, a programmable gate array (PGA), a field-programmable gate array (FPGA), etc.
Claims
1. A weakly supervised object detection method for engineering drawings, characterized by: include, Select an engineering drawing dataset, annotate the drawing data, pre-process the annotated dataset, and decompose the drawings into small pieces; Build a training and testing model, divide the training network into blocks, modify the network structure and group convolution, perform supervised training, and optimize parameter model parameters; The engineering drawing images to be inspected are fully processed and input into the trained network for prediction, verifying the changes in model performance under different parameters and further improving and optimizing the algorithm.
2. The weakly supervised object detection method for engineering drawings according to claim 1, characterized in that: The marking of the drawing data includes selecting an ABCCAD dataset from the engineering drawing dataset, and marking each engineering drawing in the ABCCAD dataset using a graffiti tool to mark a target area that needs to be detected.
3. The weakly supervised object detection method for engineering drawings according to claim 2, characterized in that: The preprocessing includes performing image channel transformation and splitting processing on the labeled data set, adjusting the quality and clarity of the image, verifying whether the preprocessed image retains key features and structural information, and grouping the preprocessed sub-images into a data set; The image channel conversion includes reading the red, green and blue channel values of the color image, calculating the grayscale value of the image using a weighted average formula, assigning the calculated grayscale value to each pixel to generate a grayscale image, and adjusting the contrast and brightness of the converted grayscale image; The weighted average formula is: , Among them, R is the red channel value, G is the green channel value, B is the blue channel value, and Gray is the grayscale value after conversion; The splitting process includes analyzing the structure of the drawing, determining a cutting plan, cutting along the determined cutting line using a cutting tool, dividing the drawing into sub-images, and saving each sub-image as an independent file.
4. The weakly supervised object detection method for engineering drawings according to claim 3, characterized in that: The construction of the training and testing model includes constructing a saliency prediction module, constructing an edge detection network, and constructing a saliency refinement module; The construction of the saliency prediction module includes initializing the VGG16 network structure and deleting the fifth pooling layer and subsequent layers, constructing a front-end saliency prediction network, and grouping the remaining convolutional layers. The ASPP module is connected after the fifth convolutional layer to capture context information. The output of the ASPP is subjected to feature fusion through two consecutive 1×1 convolutions to generate a preliminary saliency map. The edge detection network is constructed by using a feature map generated by an intermediate layer of a saliency prediction module, calculating horizontal and vertical gradients using a Sobel operator, obtaining gradient amplitudes using the horizontal and vertical gradient calculations, constructing a channel edge map, training the edge detection network using a cross entropy loss function, and minimizing the loss function; The formula for obtaining the gradient amplitude is: , in, is the center pixel coordinate, is the horizontal gradient, is the vertical gradient; The cross entropy loss function is expressed as: , in, is the cross entropy loss function, M is the number of edge pixels, is the real edge annotation, is the edge graph predicted by the network, and s is the variable index.
5. The weakly supervised object detection method for engineering drawings according to claim 4, characterized in that: The constructing of the training and testing model further includes, during the training process, using the cross entropy loss of the portion annotated with scribbles as a supervisory signal, introducing an edge-aware loss function within the salient region, and suppressing pixels outside the salient region; The supervisory signal is expressed as: , Among them, K is the set of pixels in the graffiti annotation area, A real salient map of graffiti annotations, To minimize the partial cross entropy loss function, is a rough saliency map, k is the variable index; The edge-aware loss function is expressed as: , in, is the edge weight matrix, is the final saliency map, is the initial saliency map, is the adjustment coefficient, a and b are variable indices, is the edge-aware loss function.
6. The weakly supervised object detection method for engineering drawings according to claim 5, characterized in that: The optimization parameter model parameters include defining the objective function, optimizing the edge detection network performance, and evaluating the indicator cIOU; Defining the objective function includes initializing the weights and biases of the edge detection network, calculating and integrating the loss function, calculating the gradient according to the objective function, and updating the network parameters using gradient descent until the network converges to a preset number of training iterations; The objective function is expressed as: , in, To minimize the partial cross entropy loss function, is the edge-aware loss function, is the cross entropy loss function, is the objective function, To minimize the weight coefficient of the partial cross entropy loss function, is the weight coefficient of the edge-aware loss function, is the weight coefficient of the cross entropy loss function; The evaluation indicator cIOU includes calculating the intersection area and union area of the predicted saliency map and the true saliency map to obtain the intersection-over-union ratio, calculating the distance between the center point coordinates of the predicted box and the true box, converting the distance between the center point coordinates into a penalty term, calculating the aspect ratio penalty term, and combining the IOU, center distance penalty term, and aspect ratio penalty term to calculate cIOU; The formula for obtaining the intersection-over-union ratio is: , Among them, I is the intersection area of the predicted box and the real box, U is the union area of the predicted box and the real box, For intersection and comparison; The distance formula for calculating the center point coordinates is: , in, is the center coordinate of the prediction box, is the center coordinate of the real box, and d is the Euclidean distance between the center point of the predicted box and the real box; The formula for converting to penalty term is: , in, is the center distance penalty term, c is the normalization factor related to the size of the predicted box and the true box, and d is the Euclidean distance between the center point of the predicted box and the true box; The formula for calculating the aspect ratio penalty term is: , in, is the aspect ratio penalty, is the width of the prediction box, is the height of the prediction box, is the width of the real frame, is the height of the real frame; The formula for calculating cIOU is: , in, To improve the intersection-over-union ratio, is the intersection and union ratio, is the center distance penalty term, is the aspect ratio penalty, and is the weight coefficient.
7. The weakly supervised object detection method for engineering drawings according to claim 6, characterized in that: The prediction includes comprehensively processing the engineering drawing image to be detected and inputting it into the trained network for prediction verification. After each training cycle, the generated significance prediction map is analyzed using the cIOU evaluation indicator, and the evaluation indicator threshold is set. When the cIOU is greater than or equal to the evaluation indicator threshold, it is proved that the model is not the optimal model and the parameters need to be changed. The modified model is evaluated again until it reaches the optimal standard; when the cIOU is less than the evaluation indicator threshold, it is proved that the model is the optimal model and no parameters need to be modified. The changes in model performance under different parameters are verified, the algorithm is optimized, and the optimal parameter value is determined.
8. A system using the weakly supervised object detection method for engineering drawings according to any one of claims 1 to 7, characterized in that: It includes data preparation and preprocessing module, network construction and initialization module, loss function definition and optimization module, model training and verification module, and performance evaluation module; The data preparation and preprocessing module is used to select the engineering drawing data set, annotate the drawing data, perform image channel transformation and split processing, adjust image quality, verify the retention of key features and structural information, and group the preprocessed sub-images into a data set; The network construction and initialization module is used to construct the saliency prediction module, edge detection network and saliency refinement module, initialize the network structure, modify the network structure and perform group convolution, connect to the ASPP module to capture context information, and perform feature fusion; The loss function definition and optimization module is used to define the objective function, optimize the edge detection network performance, calculate and integrate the loss function, calculate the gradient according to the objective function, and update the network parameters using gradient descent; The model training and validation module is used to perform supervised training and optimize model parameters. It uses the cross-entropy loss of the part with graffiti annotations as the supervision signal, introduces the edge-aware loss function to suppress pixels in non-salient areas, and analyzes the generated saliency prediction map using the cIOU evaluation metric after each training cycle. The performance evaluation module is used to comprehensively process the engineering drawing images to be inspected, input the trained network for prediction verification, calculate cIOU, verify the changes in model performance under different parameters, optimize the algorithm, and determine the optimal parameter values.
9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the processor implements the steps of a weakly supervised target detection method for engineering drawings according to any one of claims 1 to 7.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of a weakly supervised target detection method for engineering drawings according to any one of claims 1 to 7 are implemented.