Work order detection method based on target detection and semantic segmentation
Through the work order detection method of object detection and semantic segmentation, the problem of traditional construction quality inspection relying on manual review is solved, automated and accurate construction quality inspection is realized, and quality inspection efficiency and quality compliance are improved.
Patent Information
- Application Number
- CN202410042634.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-01-10
- Publication Date
- 2025-07-11
AI Technical Summary
Traditional construction quality inspection relies on manual photo review, which is inefficient and has high risk, and cannot achieve full quality inspection, and requires intelligent work order inspection technology.
The work ticket detection method based on object detection and semantic segmentation is adopted, including data set acquisition, labeling and standardization processing, and the improved yolov5 port detection model and the improved R2U-Net handheld line recognition model are built, and the work ticket detection is carried out in combination with object detection results and semantic segmentation results.
It realizes automated and accurate work order inspection, improves quality inspection efficiency, reduces labor costs, and ensures that construction quality complies with specifications.
Smart Images

Figure CN120299060A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence computer vision, and particularly to a work order detection model and device based on object detection and semantic segmentation. Background Art
[0002] With the rapid development of artificial intelligence, more and more industrial scenario problems can be solved by artificial intelligence. In many construction scenarios, the core problem lies in how to conduct reasonable and efficient quality inspection of the construction process, so that the staff can meet the construction specifications, ensure the construction accuracy, while reducing the labor cost in installation and maintenance quality inspection and improving the quality inspection efficiency.
[0003] Currently, traditional construction quality inspection relies on manual review of the photos uploaded by installation and maintenance personnel to determine whether the installation is qualified based on the photo information. However, the number of work orders per day is very large, and it is unrealistic to conduct full-scale manual quality inspection. Only sampling inspection can be used to check the installation quality, but this method has certain risks. Therefore, an intelligent work order detection technology that can effectively replace manual operation is particularly important. Summary of the Invention
[0004] To solve the technical problems existing in the above-mentioned prior art, the main purpose of this application is to provide a work order detection method based on object detection and semantic segmentation to solve the deficiencies of the prior art.
[0005] To achieve the above object, this application provides a work order detection method based on object detection and semantic segmentation, including:
[0006] Step 1, obtain a work order detection task scenario data set, which has construction pictures, and the pictures include construction port numbers, two-dimensional code information, and system port statuses;
[0007] Step 2, perform annotation, data cleaning, and normalization processing on the target data set, and divide it into a training set, a validation set, and a test set;
[0008] Step 3, construct and train a work order port status object detection model, which performs port detection based on the improved yolov5;
[0009] Step 4, construct and train a work order construction semantic segmentation model, which performs handheld wire recognition based on the improved R2U-Net;
[0010] Step 5, combine the object detection result and the semantic segmentation result and perform reasoning to finally obtain a work order detection model.
[0011] In addition, to achieve the above object, this application also provides a work order detection model and device based on object detection and semantic segmentation. The work order detection model and device based on object detection and semantic segmentation include:
[0012] Module 1, data acquisition module, which is used to acquire the work order detection task scenario dataset;
[0013] Module 2, data preprocessing module, which is used to label, clean and normalize the target dataset, and divide it into training set, validation set and test set;
[0014] Module 3, target detection module, which is used to implement port and its status detection;
[0015] Module 4, semantic segmentation module, which is used to identify the handheld line and identify the splitter area and cable area in the picture;
[0016] Module 5, model inference module, which uses the connected region algorithm to determine whether the QR code recognized by semantic segmentation and the port position in target detection can be connected by cables, and finally realizes work order detection.
[0017] These aspects or other aspects of the present application will be more clearly understood in the following description of the embodiments. It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present application. Brief Description of the Drawings
[0018] To more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings in the following description are only some embodiments of the present application. In the drawings:
[0019] Figure 1 It is a flow chart of the work order detection method based on target detection and semantic segmentation of the present application;
[0020] Figure 2 It is a schematic diagram of the target detection network structure of the work order detection method based on target detection and semantic segmentation of the present application;
[0021] Figure 3 It is a schematic diagram of the semantic segmentation network structure of the work order detection method based on target detection and semantic segmentation of the present application;
[0022] The implementation, functional features and advantages of the purpose of the present application will be further described with reference to the embodiments and the drawings. Detailed Embodiments
[0023] This section will describe in detail specific embodiments of the present invention. Preferred embodiments of the present invention are shown in the accompanying drawings. The function of the accompanying drawings is to supplement the description in the text part of the specification, enabling people to intuitively and vividly understand each technical feature and the overall technical solution of the present invention. However, it should not be construed as a limitation on the protection scope of the present invention.
[0024] In the description of the present invention, the meaning of several is one or more, and the meaning of multiple is more than two. Understandings such as greater than, less than, exceeding, etc. do not include the present number, and understandings such as above, below, within, etc. include the present number.
[0025] In the description of the present invention, the consecutive numbering of method steps is for the convenience of review and understanding. Considering the overall technical solution of the present invention and the logical relationship between each step, adjusting the execution order between steps will not affect the technical effects achieved by the technical solution of the present invention.
[0026] In the description of the present invention, unless otherwise clearly defined, words such as set should be understood in a broad sense. Those skilled in the art can reasonably determine the specific meaning of the above words in the present invention in combination with the specific content of the technical solution.
[0027] In an embodiment of the present application, a work order detection method based on object detection and semantic segmentation is provided. Among them, the work order detection method based on object detection and semantic segmentation can be applied to a work order detection device based on object detection and semantic segmentation. The work order detection device based on object detection and semantic segmentation can be a device with display and processing functions such as a PC, a portable computer, a mobile terminal, etc., and of course, it is not limited thereto.
[0028] The following will describe in detail some embodiments of the present application with reference to the accompanying drawings. Without conflict, the following embodiments and the features in the embodiments can be combined with each other.
[0029] Please refer to Figure 1 , Figure 1 which is a schematic flowchart of a work order detection device based on object detection and semantic segmentation according to the first embodiment of the present application. In the embodiments of the present application, the work order detection method based on object detection and semantic segmentation includes the following steps S10 to S50:
[0030] S10. Obtain a work order detection task scenario data set, which has construction pictures, and the pictures include construction port numbers, two-dimensional code information, and system port statuses;
[0031] S20. Perform annotation, data cleaning, and normalization processing on the target data set, and divide it into a training set, a validation set, and a test set;
[0032] S30. Construct and train a work order port status target detection model, which is based on the improved yolov5 for port detection;
[0033] S40. Construct and train a work order construction semantic segmentation model, which is based on the improved R2U-Net for handheld line recognition;
[0034] S50. Combine the target detection results with the semantic segmentation results and perform inference to finally obtain a work order detection model.
[0035] In the embodiment of the present application, a work order detection task scenario dataset is obtained. This dataset has construction pictures, and the pictures include construction port numbers, QR code information, and system port status, including the following:
[0036] Obtain a scenario dataset for a work order detection task. This dataset is used to support and improve the performance and efficiency of the work order management system in practical applications. It contains various key information, as described below:
[0037] Construction pictures: Such pictures reflect the actual construction scenarios, which may include ongoing engineering activities, construction equipment, and personnel operations, etc. These pictures help the system to identify and understand the specific situation on site.
[0038] Construction port numbers: Port numbers are important identifiers for identifying and tracking construction equipment or system interfaces. In the work order management system, this information is used to ensure that problems can be accurately located and solved.
[0039] QR code information: QR codes provide a way to quickly identify and track work orders. In the work order management system, QR codes can be used to quickly record and access the detailed information of work orders, including status and historical records.
[0040] System port status: System port status information refers to the occupancy status of each port in the work order management system, which is crucial for monitoring the health and performance of the system.
[0041] Obtaining and integrating these data is crucial for building an efficient and intelligent work order management system. The system can automatically create, update, and track work orders based on these data to ensure that problems are solved in a timely and effective manner. At the same time, through big data analysis and machine learning, the work order processing process can also be optimized, the work efficiency can be improved, resource waste can be reduced, and more accurate decision-making support can be provided.
[0042] In the embodiment of the present application, the target dataset is labeled, data-cleaned, and normalized, and divided into a training set, a validation set, and a test set, including the following:
[0043] The dataset is annotated using professional annotation tools. Annotation is a very important step in tasks such as object detection and image segmentation. In this task, manual annotation of the target objects is required. The annotation includes the location of the object (such as the coordinates of the bounding box) and the category of the object. For the instance segmentation task, the pixel-level segmentation of each object is also part of the annotation. The accuracy of the annotation directly affects the performance of the model.
[0044] Data cleaning is an important part of the preprocessing process, including deleting missing data, abnormal data, duplicate data, and data irrelevant to the analysis goal. This is to improve the data quality and ensure that the trained model can accurately and stably predict.
[0045] The purpose of data normalization is to compare and train different data on the same scale, which is very important for the generalization ability of the model. The normalization process includes operations such as standardization and normalization.
[0046] The purpose of dividing the dataset into three parts is to effectively monitor the performance of the model during the model training process. The training set is used to train the model, the validation set is used to adjust the parameters of the model (such as the learning rate, etc.), and the test set is used to finally evaluate the accuracy of the model. This process distributes the data according to a certain ratio. For example, 70% of the data is used as the training set, 15% as the validation set, and 15% as the test set.
[0047] In the embodiment of this application, a work order port status object detection model is constructed and trained. This model is based on the improved YOLOv5 for port detection, including the following:
[0048] YOLOv5 is a real-time object detection algorithm. Compared with other object detection algorithms, it has the characteristics of fast detection speed and high accuracy. The improved YOLOv5 model is optimized on the original basis. To reduce the computational amount and add a local attention mechanism, a sliding class attention module is proposed; at the same time, to ensure that high-order information of the image is extracted, a third-order transformation operator is proposed to better adapt to the work order detection task.
[0049] Specifically, the sliding class attention module is similar to the attention mechanism. Considering the attention mechanism formula as follows:
[0050]
[0051] Among them, Q is the matrix composed of query vectors, K is the matrix composed of key vectors, V is the matrix composed of value vectors, and d is the dimension of the key vectors. Softmax is the activation function, taking the exponent of each point along the corresponding dimension and then normalizing.
[0052] The above formula performs a similarity operation on the query matrix and the key matrix, and then multiplies it by the value matrix to obtain the final vector representation. However, this operation requires a large amount of computation, which affects the real-time performance of the network. Therefore, it is improved by using a sliding window and cross-attention to collect local features and finally form a feature map. First, the tensor is divided into m×n sub-tensors X ij , and then each sub-tensor X ij and its transpose are respectively multiplied by the prior feature block P k to obtain M kij .
[0053] M kij = X ij P k ,
[0054] where M kij is the sub-tensor of the k-th channel at the i-th row and j-th column, X ij is the sub-vector of the i-th row and j-th column in the m×n sub-tensors obtained by block division, and P k is the prior feature block, which is used as a parameter during training.
[0055] To ensure that the size of the output feature map remains unchanged, the width and height of the above two are set to the same value in this paper. The prior feature block can be regarded as the result of multiplying the transpose of the key matrix and the value matrix and undergoing a certain transformation. Since it will pass through the batch normalization layer later, it is not normalized here. The number of prior feature blocks determines the number of channels of the output feature map, and the number of output channels is twice the number of input channels multiplied by the number of feature blocks.
[0056] Specifically, the third-order transformation operator performs three transformations on each pixel of the feature map to extract the high-order information of the image. For each pixel, the transformation is performed according to the formula y = ax + bx 2 + cx 3 . Then, feature fusion is performed through the convolutional layer, and feature normalization is achieved through the batch normalization layer. In the use of the activation function, max(0, -x) is introduced as the activation function for some layers to expect to retain more information.
[0057] The formula is as follows:
[0058] y = A·X + B·X 2 + C·X 3 , O = Activation(BN(y))
[0059] where X is the input tensor, X 2 is the square of the input tensor, and X 3is the cube of the input tensor, y is the tensor after the input tensor undergoes three operator transformations, · is the Hadamard product, that is, pointwise multiplication, BN is the BatchNorm layer, and Activation is the activation function max(0, -x).
[0060] Add the above operators to the network and fuse them with the original YOLO network to improve the detection accuracy in complex scenarios while maintaining real-time performance.
[0061] In the embodiment of this application, a work order construction semantic segmentation model is constructed and trained. This model is based on the improved R2U-Net for handheld wire recognition and includes the following:
[0062] R2U-Net is a convolutional neural network architecture based on U-Net, specifically for medical image segmentation. It introduces a Recurrent Residual Convolutional Unit on the basis of U-Net, which can enhance the network's expressive ability and segmentation accuracy by introducing loop shortcut connections and skip connections.
[0063] The recursive residual convolutional unit of R2U-Net contains two levels of loop shortcut connections, allowing gradients to flow more directly and efficiently in the network. In addition, R2U-Net also retains the skip connections in U-Net, which combine low-level feature maps with high-level feature maps, helping the network to utilize global information during segmentation.
[0064] To improve the inference speed of the model, the upsampling operator of the semantic segmentation network is improved. Let the size of the depth feature map be M×N, and each position element be a tensor X with a depth of C pq , when performing upsampling on it, use the linear combination of all positions of the deep feature map to represent the element M at the i-th row and j-th column of the upsampled feature map ij . That is,
[0065]
[0066] where, X pq is the sub-tensor with a depth of C at the i-th row and j-th column of the depth feature map with a size of M×N, a ijpq is a scalar, which is the weight of the linear combination. M ij is the element with a depth of C at the i-th row and j-th column of the upsampled feature map.
[0067] In the embodiment of this application, the object detection result is combined with the semantic segmentation result and inference is performed to finally obtain a work order detection model. It includes:
[0068] Input the standardized captured image into the work order detection device based on object detection and semantic segmentation. Use the semantic segmentation algorithm to identify the splitter area and cable area in the image. Then, use the connected component algorithm to determine whether the recognized QR code and the port position detected by object detection can be connected by a cable. If there is a connection path, the recognition is successful; otherwise, it indicates that the port number detection fails.
[0069] Among them, each module in the above work order detection device based on object detection and semantic segmentation corresponds to each step in the above work order detection model and device embodiments based on object detection and semantic segmentation. Its functions and implementation processes will not be elaborated here one by one.
[0070] It should be noted that in this article, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including a..." does not exclude the existence of another identical element in the process, method, article or device including that element.
[0071] The serial numbers of the above embodiments of the present application are only for description and do not represent the superiority or inferiority of the embodiments.
[0072] The present application can be used in many general or special computer device environments or configurations. Further, the method can be implemented in any type of computing platform operably connected, including but not limited to personal computers, minicomputers, mainframes, workstations, network or distributed computing environments, separate or integrated computer platforms, or communicating with charged particle tools or other imaging devices, etc. Aspects of the present invention can be implemented with machine-readable code stored on a non-transitory storage medium or device, whether removable or integrated into a computing platform, such as a hard disk, optical read and / or write storage medium, RAM, ROM, etc., such that it can be read by a programmable computer. When the storage medium or device is read by a computer, it can be used to configure and operate the computer to execute the processes described herein. In addition, the machine-readable code, or a part thereof, can be transmitted via a wired or wireless network. When such media includes instructions or programs that implement the above-described steps in combination with a microprocessor or other data processor, the invention described herein includes these and other different types of non-transitory computer-readable storage media. When programmed according to the methods and techniques of the present invention, the present invention also includes the computer itself.
[0073] A computer program can be applied to input data to perform the functions described herein, thereby transforming the input data to generate output data stored in a non-volatile memory. The output information can also be applied to one or more output devices such as a display. In a preferred embodiment of the present invention, the transformed data represents physical and tangible objects, including specific visual depictions of physical and tangible objects generated on a display.
[0074] Through the description of the above embodiments, those skilled in the art can clearly understand that the above-described embodiment methods can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases, the former is a better implementation method. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product is stored in a storage medium as described above (such as ROM / RAM, magnetic disk, optical disk), and includes several instructions for causing a terminal device (which can be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in various embodiments of the present application.
[0075] The above are only the preferred embodiments of the present application, and do not limit the patent scope of the present application. Any equivalent structure or equivalent process transformation made by using the content of the specification and drawings of the present application, or directly or indirectly applied in other related technical fields, shall be equally included in the patent protection scope of the present application.
Claims
1. A work order detection method based on object detection and semantic segmentation, characterized in that It includes the following steps: Obtain a work order detection task scenario dataset, which has construction pictures, and the pictures include construction port numbers, QR code information, and system port statuses; Annotate, clean, and normalize the target dataset, and divide it into a training set, a validation set, and a test set; Construct and train a work order port status target detection model, which performs port detection based on the improved YOLOv5; Construct and train a work order construction semantic segmentation model, which performs handheld line recognition based on the improved R2U-Net; Combine the target detection results with the semantic segmentation results and perform inference to finally obtain a work order detection model.
2. The work order detection method based on object detection and semantic segmentation according to claim 1, wherein, The obtaining of the work order detection task scenario dataset includes: Collect and organize pictures with construction port numbers, QR code information, and system port statuses uploaded by users.
3. The work order detection method based on object detection and semantic segmentation according to claim 1, wherein The annotating, cleaning, and normalizing of the target dataset and dividing it into a training set, a validation set, and a test set include: For work order pictures in normal and abnormal states, annotate information such as construction ports, system ports, and QR code positions in the pictures; according to the quality and content of the pictures, clean and normalize the target dataset, such as scaling, cropping, flipping, etc., to ensure the quality and consistency of the data; finally divide it into a training set, a validation set, and a test set.
4. The work order detection method based on object detection and semantic segmentation according to claim 1, wherein The constructing and training of the work order port status target detection model, which performs port detection based on the improved YOLOv5, includes: Based on the improved YOLOv5 target detection framework, construct a work order port status detection model. YOLOv5 is a real-time target detection algorithm, which has the characteristics of fast detection speed and high accuracy compared with other target detection algorithms. The improved YOLOv5 model is optimized on the original basis. To reduce the computational amount, a local attention mechanism is added, and a sliding class attention module is proposed; at the same time, to ensure the extraction of high-order information of the image, a third-order transformation operator is proposed to better adapt to the work order detection task.
5. The work order detection method based on object detection and semantic segmentation according to claim 1, wherein The constructing and training of the work order construction semantic segmentation model, which performs handheld line recognition based on the improved R2U-Net, includes: Based on the improved R2U-Net semantic segmentation network, construct a work order construction semantic segmentation model. The improved R2U-Net uses residual units to achieve deeper parsing of the input data while maintaining high accuracy, effectively improving the learning ability of the model; the features are added in the recurrent residual layer to enhance the extraction of segmentation features; by improving the structure of U-Net, the segmentation performance of the network is improved without increasing parameters.
6. The work order detection method based on object detection and semantic segmentation according to claim 1, wherein The combining of the target detection results with the semantic segmentation results and performing inference to finally obtain a work order detection model includes: For the splitter area and cable area in the picture recognized by the semantic segmentation algorithm, use the connected component algorithm to judge whether the QR code recognized by the semantic segmentation and the port position in the target detection can be connected by a cable. If there is a connected path, the recognition is successful; if not, it means that the port number detection fails.
7. A work order detection method based on object detection and semantic segmentation, characterized in that, The work order detection method based on object detection and semantic segmentation includes the work order detection method based on object detection and semantic segmentation described in any one of claims 1-6, and the method includes: (The first step) A data acquisition module, where the acquisition module is used to acquire a work order detection task scenario dataset; (The second step) A data preprocessing module, where the data preprocessing module is used to label, clean, and normalize the target dataset, and divide it into a training set, a validation set, and a test set; (The third step) An object detection module, where the object detection module is used to implement port and its status detection; (The fourth step) A semantic segmentation module, where the semantic segmentation module is used for handheld line recognition to identify the splitter area and cable area in the picture; (The fifth step) A model inference module, where the model inference module uses a connected component algorithm to determine whether the QR code recognized by semantic segmentation and the port position in object detection can be connected by a cable, and finally realizes work order detection.