Muck truck detection method and system
Through the image analysis model and channel attention mechanism based on UNet network, semantic segmentation and feature selection of dump trucks are solved, and the existing dump truck detection methods are inefficient and difficult to guarantee accuracy is achieved, and high-precision and robust dump truck detection are achieved.
Patent Information
- Application Number
- CN202510190757.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-20
- Publication Date
- 2025-06-17
AI Technical Summary
The existing dump truck detection methods rely on a large amount of labeled data, resulting in low detection efficiency and difficult to ensure accuracy.
Semantic segmentation is performed using an image analysis model based on UNet network, and dynamic feature selection is performed on the global feature information through the channel attention mechanism to output detection results of whether the semantic annotation information is correct.
It realizes accurate identification and semantic annotation of dump trucks, improves the robustness and adaptability of detection, and improves the detection accuracy.
Smart Images

Figure CN120164157A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image processing, and particularly to a detection method and system for a muck truck. Background Art
[0002] Currently, muck trucks are widely used in the construction industry. They are easy to operate and have high transportation efficiency. However, due to the special nature of muck trucks, they are prone to causing road pollution and dust during driving, posing a threat to the urban environment and traffic safety. Therefore, the detection and supervision of muck trucks become crucial.
[0003] Existing muck truck detection methods mainly rely on manual inspection and a large amount of labeled data. These methods are inefficient and it is difficult to guarantee the detection effect. Summary of the Invention
[0004] In order to solve the problems of low detection efficiency and difficult-to-guarantee accuracy of existing muck truck detection methods due to reliance on a large amount of labeled data, the present invention proposes a muck truck detection method, including:
[0005] Based on the video stream monitoring data of each road section in the target area collected, using a pre-constructed image analysis model for semantic segmentation to obtain the semantic annotation information corresponding to the muck truck in the video stream monitoring data;
[0006] Extract context information according to the semantic annotation information corresponding to the muck truck to obtain global feature information;
[0007] Perform dynamic feature selection on the global feature information through a channel attention mechanism to obtain the high-level semantic features corresponding to the global feature information;
[0008] According to the high-level semantic features, output the detection result of whether the semantic annotation information is correct;
[0009] Wherein, the image analysis model is constructed based on the UNet network.
[0010] Optionally, the step of based on the video stream monitoring data of each road section in the target area collected, using a pre-constructed image analysis model for semantic segmentation to obtain the semantic annotation information corresponding to the muck truck in the video stream monitoring data includes:
[0011] Preprocess the video stream monitoring data of each road section in the target area collected to obtain preprocessed image data;
[0012] According to the preprocessed image data, use a pre-constructed image analysis model for semantic segmentation to output the semantic segmentation results of each image frame in the preprocessed image data; wherein, the semantic segmentation results include: a segmentation mask image and an embedding representation;
[0013] Extract the pixel region information of the dump truck in the video stream monitoring data from the segmentation mask image;
[0014] Extract the regional feature vector of the dump truck from the embedding representation;
[0015] Integrate the pixel region information and the regional feature vector of the dump truck to generate the semantic annotation information of the dump truck.
[0016] Optionally, the image analysis model includes the following construction process:
[0017] Use the historical image data of the dump truck as the input of the training data;
[0018] Use the annotation information of the historical image data as the output of the training data;
[0019] Train the UNet network based on the input and output of the training data to obtain an image analysis model.
[0020] Optionally, the preprocessing of the video stream monitoring data of each road section in the collected target area to obtain preprocessed image data includes:
[0021] Use a preset video processing tool to cut the video stream monitoring data of each road section in the collected target area into single-frame images at a set time interval;
[0022] Denoise and normalize each single-frame image respectively to obtain primary processed data;
[0023] Convert the color space of the primary processed data to obtain preprocessed image data.
[0024] Optionally, the preprocessed image data includes: RGB image data, HSV image data.
[0025] Optionally, the dynamic feature selection of the global feature information through the channel attention mechanism to obtain the high-level semantic feature corresponding to the global feature information includes:
[0026] Allocate attention weights to the global feature information through the channel attention mechanism to obtain a weighted feature map;
[0027] Fuse the features of the weighted feature map to generate a high-resolution feature map;
[0028] Extract semantic features from the high-resolution feature map to obtain the corresponding high-level semantic features.
[0029] Optionally, outputting a detection result of whether the semantic annotation information is correct according to the height semantic feature includes:
[0030] Upsampling the height semantic feature to obtain sampling characteristic information;
[0031] According to the sampling characteristic information, using the Faster R-CNN network to generate a bounding box and a class label of the muck truck;
[0032] Performing data augmentation on the generated bounding box and class label of the muck truck, and outputting a detection result of whether the enhanced bounding box and class label of the muck truck are correct.
[0033] Based on the same inventive concept, the present invention also provides a muck truck detection system, including:
[0034] A semantic segmentation module, configured to perform semantic segmentation on the video stream monitoring data of each road section in the collected target area by using a pre-constructed image analysis model to obtain semantic annotation information corresponding to the muck truck in the video stream monitoring data;
[0035] An information extraction module, configured to extract context information according to the semantic annotation information corresponding to the muck truck to obtain global feature information;
[0036] A feature selection module, configured to perform dynamic feature selection on the global feature information through a channel attention mechanism to obtain a height semantic feature corresponding to the global feature information;
[0037] A target detection module, configured to output a detection result of whether the semantic annotation information is correct according to the height semantic feature;
[0038] Wherein, the image analysis model is constructed based on the UNet network.
[0039] Optionally, the semantic segmentation module includes:
[0040] A preprocessing sub-module, configured to preprocess the video stream monitoring data of each road section in the collected target area to obtain preprocessed image data;
[0041] A data parsing sub-module, configured to perform semantic segmentation on the preprocessed image data by using a pre-constructed image analysis model and output a semantic segmentation result of each image frame in the preprocessed image data; wherein, the semantic segmentation result includes: a segmentation mask image and an embedding representation;
[0042] A region extraction sub-module, configured to extract pixel region information of the muck truck in the video stream monitoring data from the segmentation mask image;
[0043] A vector extraction sub-module, configured to extract the regional feature vector of the muck truck from the embedding representation;
[0044] A data integration sub-module, configured to integrate the pixel region information and the regional feature vector of the muck truck to generate the semantic annotation information of the muck truck.
[0045] Optionally, the system further includes a model construction module, including:
[0046] An input setting sub-module, configured to use the historical image data of the muck truck as the input of the training data;
[0047] An output setting sub-module, configured to use the annotation information of the historical image data as the output of the training data;
[0048] A network training sub-module, configured to train the UNet network based on the input and output of the training data to obtain an image analysis model.
[0049] Optionally, the preprocessing sub-module includes:
[0050] An image cutting unit, configured to use a preset video processing tool to cut the video stream monitoring data of each section in the collected target area into single-frame images at a set time interval;
[0051] A denoising processing unit, configured to perform denoising and normalization on the single-frame images respectively to obtain primary processed data;
[0052] An image conversion unit, configured to perform color space conversion on the primary processed data to obtain preprocessed image data.
[0053] Optionally, the preprocessed image data includes: RGB image data, HSV image data.
[0054] Optionally, the feature selection module includes:
[0055] A weight assignment sub-module, configured to perform attention weight assignment on the global feature information through a channel attention mechanism to obtain a weighted feature map;
[0056] A feature fusion sub-module, configured to fuse the weighted feature maps to generate a high-resolution feature map;
[0057] A semantic extraction sub-module, configured to perform semantic feature extraction on the high-resolution feature map to obtain corresponding high-level semantic features.
[0058] Optionally, the target detection module includes:
[0059] An upsampling sub-module, configured to perform upsampling on the height semantic feature to obtain sampling characteristic information;
[0060] A label generation sub-module, configured to generate a bounding box and a class label of the muck truck according to the sampling characteristic information by using a Faster R-CNN network;
[0061] A data augmentation sub-module, configured to perform data augmentation on the generated bounding box and class label of the muck truck, and output a detection result indicating whether the augmented bounding box and class label of the muck truck are correct.
[0062] On the other hand, the present invention further provides an electronic device, including: at least one processor and a memory; the memory and the processor are connected through a bus;
[0063] The memory is configured to store one or more programs;
[0064] When the one or more programs are executed by the at least one processor, the method for detecting a muck truck as described above is implemented.
[0065] On the other hand, the present invention further provides a computer-readable storage medium having an execution program stored thereon, and when the execution program is executed, the method for detecting a muck truck as described above is implemented.
[0066] Compared with the prior art, the beneficial effects of the present invention are:
[0067] The present invention provides a method and a system for detecting a muck truck, including: based on the video stream monitoring data of each road section in a target area collected, performing semantic segmentation by using a pre-constructed image analysis model to obtain semantic annotation information corresponding to the muck truck in the video stream monitoring data; extracting context information according to the semantic annotation information corresponding to the muck truck to obtain global feature information; performing dynamic feature selection on the global feature information through a channel attention mechanism to obtain a height semantic feature corresponding to the global feature information; according to the height semantic feature, outputting a detection result indicating whether the semantic annotation information is correct; the present application uses a pre-constructed image analysis model to perform semantic segmentation on the video stream monitoring data of each road section in a target area, and can accurately extract the semantic annotation information of the muck truck; by extracting context information to obtain global feature information, the ability to understand complex scenes can be enhanced, the problem that a single feature is difficult to cope with environmental changes and occlusion can be solved, and it is beneficial to improve the robustness and adaptability of detection; by performing dynamic feature selection on the global feature information through a channel attention mechanism, the feature expression can be optimized, and the detection accuracy is further improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0068] Figure 1 It is a schematic flowchart of a method for detecting a muck truck provided by the present invention;
[0069] Figure 2 Schematic diagram of the feature fusion framework process in a detection method for muck trucks provided by the present invention;
[0070] Figure 3 Schematic diagram of the target detection process in a detection method for muck trucks provided by the present invention;
[0071] Figure 4 Schematic diagram of the structural composition of a detection system for muck trucks provided by the present invention. Detailed implementation manners
[0072] The present invention proposes a detection method, system, device and medium for muck trucks. The following further elaborates on the detailed implementation manners of the present invention with reference to the accompanying drawings.
[0073] Embodiment 1:
[0074] The present invention provides a detection method for muck trucks. The schematic diagram of the process is as Figure 1 shown and includes:
[0075] Step 1: Based on the video stream monitoring data of each road section in the collected target area, use a pre-constructed image analysis model to perform semantic segmentation to obtain the semantic annotation information corresponding to the muck trucks in the video stream monitoring data;
[0076] Step 2: Extract context information according to the semantic annotation information corresponding to the muck trucks to obtain global feature information;
[0077] Step 3: Perform dynamic feature selection on the global feature information through a channel attention mechanism to obtain the high-level semantic features corresponding to the global feature information;
[0078] Step 4: Output the detection result of whether the semantic annotation information is correct according to the high-level semantic features;
[0079] Among them, the image analysis model is constructed based on the UNet network.
[0080] In one implementation manner, the process of performing semantic segmentation on the video stream monitoring data of each road section in the collected target area using a pre-constructed image analysis model to obtain the semantic annotation information corresponding to the muck trucks in the video stream monitoring data in the above step 1 may include:
[0081] Preprocess the video stream monitoring data of each road section in the collected target area to obtain preprocessed image data;
[0082] Based on the preprocessed image data, semantic segmentation is performed using a pre-constructed image analysis model, and the semantic segmentation results of each image frame in the preprocessed image data are output; wherein, the semantic segmentation results include: a segmentation mask image and an embedding representation;
[0083] Extract the pixel region information of the dump truck in the video stream monitoring data from the segmentation mask image;
[0084] Extract the regional feature vector of the dump truck from the embedding representation;
[0085] Integrate the pixel region information and the regional feature vector of the dump truck to generate the semantic annotation information of the dump truck;
[0086] In this implementation manner, through systematic processing and analysis of the video stream monitoring data of each road section in the target area, accurate identification and semantic annotation of the dump truck can be achieved. First, by preprocessing the collected data, such as image denoising and normalization, high-quality input images are provided for the subsequent image analysis model. This step can not only effectively reduce the noise impact caused by external interference but also improve the training and inference efficiency of the model. Based on the preprocessed image data, semantic segmentation is performed using a pre-constructed image analysis model to output the semantic segmentation results of each frame, including the segmentation mask image and the embedding representation. This deep learning-based image processing method can assign clear pixel labels to the dump truck in the image, ensuring its accurate identification in complex scene backgrounds. In addition, by extracting the pixel region information and its corresponding regional feature vector of the dump truck from the segmentation mask image, the ability to capture the features of the dump truck can be further enhanced, making the semantic annotation information more abundant and accurate. By effectively integrating the pixel region information and the feature vector, the finally generated semantic annotation information of the dump truck can be widely applied to subsequent intelligent monitoring, data analysis, and urban management.
[0087] In one implementation manner, the above-mentioned image analysis model may include the following construction process:
[0088] Use the historical image data of the dump truck as the input of the training data;
[0089] Use the annotation information of the historical image data as the output of the training data;
[0090] Based on the input and output of the training data, train the UNet network to obtain an image analysis model;
[0091] In this implementation, during the initial training phase of the object detection network, transfer learning is carried out using a pre-trained deep learning model on a public dataset. This process enables the model to quickly adapt to the task of detecting dump trucks, while significantly reducing the dependence on a large amount of labeled data. Through the method of incremental learning, as new data is continuously collected, the model is gradually fine-tuned to continuously optimize its performance and overcome the challenges brought about by changes in the types and appearances of dump trucks, thereby effectively improving the detection effect of the model. In this process, the construction of an image analysis model based on the UNet network plays a crucial role. This model uses the historical image data of dump trucks as the training input and the annotation information as the output, enabling the model to learn the expression features of dump trucks in different scenarios and automatically identify and segment the pixel regions of dump trucks to generate accurate segmentation masks. Through its powerful feature extraction ability, the UNet network realizes in-depth analysis of each image, hierarchically extracts and retains the important information in the image, enabling the detection result to not only accurately identify dump trucks but also effectively separate elements of different backgrounds and types. By saving the segmentation results of each image, the generated masks of different categories are efficiently stored and managed, laying a solid data foundation for subsequent information processing. This processing flow can significantly improve the recognition accuracy and processing efficiency of dump trucks, greatly enhancing the intelligent level of automatic monitoring and management.
[0092] In one implementation, the process of preprocessing the video stream monitoring data of each road section in the collected target area to obtain preprocessed image data may include:
[0093] Using a preset video processing tool to cut the video stream monitoring data of each road section in the collected target area into single-frame images at a set time interval;
[0094] Denosing and normalizing the single-frame images respectively to obtain primary processed data;
[0095] Performing color space conversion on the primary processed data to obtain preprocessed image data;
[0096] Exemplarily, the above preprocessed image data may include: RGB image data, HSV image data;
[0097] In this implementation, by systematically preprocessing the video stream monitoring data of each road section in the target area, the accuracy and efficiency of subsequent image analysis and object detection can be greatly improved. First, the monitoring data is cut into single-frame images at a set time interval by using a preset video processing tool. This process lays the foundation for subsequent processing, ensuring that each frame of image is independent and operable. After denoising and normalization processing, the obtained primary processed data significantly reduces the influence of noise, making the input image data cleaner and more consistent, and providing a high-quality input source for deep learning algorithms. Subsequently, through color space conversion, the primary processed data is converted into RGB and HSV image data. In this step, the use of the HSV color space is particularly crucial because for construction waste trucks, their color features are obvious, mainly green and red. Using HSV images can better emphasize these color features, thus improving the detection effect. This meticulousness in the image preprocessing process increases the accuracy of subsequent object detection. Based on the processed image data, through object detection networks such as Faster R-CNN, YOLO, SSD, or Centernet, the construction waste trucks in each frame of image are identified. Specifically, the inputs to the model include RGB images, HSV images, and the output of the scene understanding network, which provides two parts of semantic segmentation masks and related semantic embeddings. This diverse input method can provide more comprehensive information for the model, enabling the network to more effectively capture the color and shape features of construction waste trucks during the feature extraction and fusion stages. In the feature extraction link, after processing the RGB and HSV images respectively, the extracted features will be transmitted through the backbone to the neck and decoder for deep fusion. On this basis, the fusion module applies an attention mechanism to effectively fuse the features from different sources, then performs upsampling and splicing through the decoder, and finally outputs to the detection and classification head for object recognition. Generally speaking, this implementation forms an efficient and accurate construction waste truck recognition process through image preprocessing and multi-stage object detection networks. The implementation of this method not only improves the recognition accuracy but also provides reliable data support for urban traffic management and construction waste truck monitoring, laying the foundation for the development of intelligent transportation systems.
[0098] In one implementation, the process of dynamically selecting features of the global feature information through the channel attention mechanism in step 3 above to obtain the high-level semantic features corresponding to the global feature information may include:
[0099] Performing attention weight assignment on the global feature information through the channel attention mechanism to obtain a weighted feature map;
[0100] Fuse the weighted feature maps to generate high-resolution feature maps;
[0101] Extract semantic features from the high-resolution feature maps to obtain corresponding high-level semantic features;
[0102] In this implementation, by applying the channel attention mechanism to perform dynamic feature selection on the global feature information, the extraction process of high-level semantic features is realized. In this process, the channel attention mechanism can not only effectively allocate weights to the global feature information, but also highlight the attention to the features related to dump trucks, ensuring that important features are reasonably amplified, while redundant or irrelevant features are suppressed. Through this dynamic adjustment, the obtained weighted feature maps can better reflect the key information in the image, thus laying a good foundation for subsequent feature fusion and semantic feature extraction. Further, fusing the weighted feature maps can generate high-resolution feature maps, thereby further enhancing the feature representation ability and resolution. This process is particularly important because in the field of image processing, improving the resolution of features can significantly enhance the model's ability to capture details, making the detection of dump trucks in various complex scenarios more accurate. Finally, extracting semantic features from the high-resolution feature maps, high-level semantic features are ultimately obtained. These features not only have rich semantic information but also can effectively describe the specific performance of the target object in the image.
[0103] In one implementation, the process of outputting the detection result of whether the semantic annotation information is correct in step 4 above according to the high-level semantic features may include:
[0104] Upsample the high-level semantic features to obtain sampled feature information;
[0105] According to the sampled feature information, use the Faster R-CNN network to generate the bounding boxes and class labels of the dump trucks;
[0106] Perform data augmentation on the generated bounding boxes and class labels of the dump trucks, and output the detection result of whether the enhanced bounding boxes and class labels of the dump trucks are correct;
[0107] In this implementation, through the processing of high-level semantic features, the process of effective detection of dump trucks and output of semantic annotation information is achieved. In this process, first, the high-level semantic features are upsampled to obtain sampling characteristic information. The design of this step is to restore the spatial resolution of the feature map, thereby enhancing the recognition ability of dump trucks at the detail level. When the sampling characteristic information is further processed, the Faster R-CNN network is used to generate the bounding boxes and class labels of the dump trucks. This link utilizes the powerful detection ability of the deep learning model to ensure the positioning and recognition accuracy of the dump trucks. Then, the generated bounding boxes and class labels are processed by the data augmentation method. The purpose of this process is to further improve the robustness and generalization ability of the model by expanding the diversity of the data. The enhanced bounding boxes and class labels not only enrich the training samples but also ensure that the model can still maintain high recognition performance when facing different environments and conditions. Finally, the output detection results can evaluate whether the generated bounding boxes and class labels of the dump trucks are correct. This feedback mechanism not only provides a basis for further optimization of the model but also improves the adaptive ability of the system.
[0108] Embodiment 2:
[0109] Based on the same inventive concept, the present invention also provides a dump truck detection system, the structural composition schematic diagram of which is as Figure 4 shown, including:
[0110] A semantic segmentation module, configured to perform semantic segmentation on the video stream monitoring data of each section in the collected target area by using a pre-constructed image analysis model to obtain the semantic annotation information corresponding to the dump trucks in the video stream monitoring data;
[0111] An information extraction module, configured to extract context information according to the semantic annotation information corresponding to the dump trucks to obtain global feature information;
[0112] A feature selection module, configured to perform dynamic feature selection on the global feature information through a channel attention mechanism to obtain the high-level semantic features corresponding to the global feature information;
[0113] A target detection module, configured to output a detection result indicating whether the semantic annotation information is correct according to the high-level semantic features;
[0114] Wherein, the image analysis model is constructed based on the UNet network.
[0115] In one implementation, the above-mentioned semantic segmentation module may include:
[0116] A preprocessing sub-module, configured to preprocess the video stream monitoring data of each section in the collected target area to obtain preprocessed image data;
[0117] A data parsing sub-module, configured to perform semantic segmentation on the preprocessed image data by using a pre-constructed image analysis model, and output semantic segmentation results of each image frame in the preprocessed image data; wherein, the semantic segmentation results include: a segmentation mask image and an embedding representation;
[0118] A region extraction sub-module, configured to extract pixel region information of a muck truck in the video stream monitoring data from the segmentation mask image;
[0119] A vector extraction sub-module, configured to extract a region feature vector of the muck truck from the embedding representation;
[0120] A data integration sub-module, configured to integrate the pixel region information and the region feature vector of the muck truck to generate semantic annotation information of the muck truck.
[0121] In one implementation, the above system may further include a model construction module, specifically including:
[0122] An input setting sub-module, configured to use historical image data of a muck truck as the input of training data;
[0123] An output setting sub-module, configured to use annotation information of the historical image data as the output of training data;
[0124] A network training sub-module, configured to train a UNet network based on the input and output of the training data to obtain an image analysis model.
[0125] In one implementation, the above preprocessing sub-module may include:
[0126] An image cutting unit, configured to cut the video stream monitoring data of each road section in the collected target area into single-frame images at a set time interval by using a preset video processing tool;
[0127] A denoising processing unit, configured to perform denoising and standardization on the single-frame images respectively to obtain primary processed data;
[0128] An image conversion unit, configured to perform color space conversion on the primary processed data to obtain preprocessed image data.
[0129] Exemplarily, the above preprocessed image data may include: RGB image data, HSV image data.
[0130] In one implementation, the above feature selection module may include:
[0131] A weight assignment sub-module, configured to perform attention weight assignment on the global feature information through a channel attention mechanism to obtain a weighted feature map;
[0132] A feature fusion sub-module, configured to perform feature fusion on the weighted feature maps to generate high-resolution feature maps;
[0133] A semantic extraction sub-module, configured to perform semantic feature extraction on the high-resolution feature maps to obtain corresponding high-level semantic features.
[0134] In one implementation, the above object detection module may include:
[0135] An upsampling sub-module, configured to perform upsampling on the high-level semantic features to obtain sampling characteristic information;
[0136] A label generation sub-module, configured to generate the bounding box and class label of the dump truck by using the Faster R-CNN network according to the sampling characteristic information;
[0137] A data augmentation sub-module, configured to perform data augmentation on the generated bounding box and class label of the dump truck, and output a detection result indicating whether the augmented bounding box and class label of the dump truck are correct.
[0138] Embodiment 3:
[0139] The present invention further provides an electronic device, which may be a computer device, a single-chip microcomputer device, a smart mobile device, etc. The electronic device in this embodiment may include a processor, a memory, a transceiver component, etc. The memory, the processor, and the transceiver component are connected by a bus; the memory may be used to store an execution program, and an exemplary execution program may include instructions; the processor is used to execute the instructions stored in the memory. The memory may also be used to store data, and the data may be called and / or modified when the instructions are executed.
[0140] The processor may be a Central Processing Unit (CPU), or may also be other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. It is the computing core and control core of the terminal, and is suitable for implementing one or more instructions, specifically suitable for loading and executing one or more instructions in the storage medium to implement the corresponding method flow or corresponding function, so as to implement the steps of a dump truck detection method in the above embodiment.
[0141] Embodiment 4:
[0142] Based on the same inventive concept, the present invention also provides a readable storage medium, specifically an electronic device-readable storage medium (Memory). The electronic device-readable storage medium is a memory device in an electronic device, used to store programs and data. It can be understood that the storage medium here can include both the built-in storage medium in the electronic device and, of course, the extended storage medium supported by the electronic device. The storage medium provides a storage space, and this storage space stores the operating system of the terminal. Moreover, in this storage space, there is also stored one or more instructions suitable for being loaded and executed by a processor. These instructions can be one or more execution programs (including program codes). It should be noted that the storage medium here can be a high-speed RAM memory or a non-volatile memory, such as at least one disk memory. By loading and executing one or more instructions stored in the storage medium by the processor, the steps of a detection method for a muck truck in the above embodiments can be implemented.
[0143] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memories, CD-ROMs, optical memories, etc.) containing computer-usable program codes.
[0144] The present invention is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each flow and / or block in the flowchart and / or block diagram, and the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for implementing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0145] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device, and this instruction device implements the functions in the flow Figure 1 one flow or multiple flows and / or blocks Figure 1The functions specified in one or more boxes.
[0146] These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process. Thus, the instructions executed on the computer or other programmable device provide for implementing the steps of the functions specified in one Figure 1 one process or more processes and / or boxes Figure 1 step of the functions specified in one box or more boxes.
[0147] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention rather than to limit the scope of its protection. Although the present invention has been described in detail with reference to the above embodiments, those of ordinary skill in the art should understand that: after reading the present invention, those skilled in the art can still make various changes, modifications or equivalent replacements to the specific implementation manners of the application. However, these changes, modifications or equivalent replacements are all within the scope of the protection of the claims pending for approval of the application.
Claims
1. A method for detecting a muck truck, characterized in that: include: Based on the collected video stream monitoring data of each road section in the target area, semantic segmentation is performed using a pre-built image analysis model to obtain semantic annotation information corresponding to the muck truck in the video stream monitoring data; Extracting context information according to the semantic annotation information corresponding to the muck truck to obtain global feature information; Performing dynamic feature selection on the global feature information through a channel attention mechanism to obtain a high-level semantic feature corresponding to the global feature information; Outputting a detection result of whether the semantic annotation information is correct according to the high semantic features; Wherein, the image analysis model is constructed based on the UNet network.
2. The method according to claim 1, characterized in that The video stream monitoring data of each road section in the target area is collected, and semantic segmentation is performed using a pre-built image analysis model to obtain semantic annotation information corresponding to the muck truck in the video stream monitoring data, including: Preprocessing the collected video stream monitoring data of each road section in the target area to obtain preprocessed image data; According to the pre-processed image data, semantic segmentation is performed using a pre-built image analysis model, and a semantic segmentation result of each image frame in the pre-processed image data is output; wherein the semantic segmentation result includes: a segmentation mask image and an embedded representation; Extracting pixel area information of the muck truck in the video stream monitoring data from the segmentation mask image; Extracting a regional feature vector of the muck truck from the embedded representation; The pixel area information and the regional feature vector of the muck truck are integrated to generate semantic annotation information of the muck truck.
3. The method according to claim 1 or 2, characterized in that The image analysis model includes the following construction process: The historical image data of muck trucks is used as the input of training data; Using the annotation information of the historical image data as output of training data; Based on the input and output of the training data, the UNet network is trained to obtain an image analysis model.
4. The method according to claim 2, characterized in that The preprocessing of the collected video stream monitoring data of each road section in the target area to obtain preprocessed image data includes: Use the preset video processing tool to cut the collected video stream monitoring data of each road section in the target area into single frame images according to the set time interval; Denoising and standardizing the single-frame images respectively to obtain primary processed data; The primary processed data is subjected to color space conversion to obtain pre-processed image data.
5. The method according to claim 4, characterized in that The pre-processed image data includes: RGB image data and HSV image data.
6. The method according to claim 1, characterized in that The dynamic feature selection of the global feature information by the channel attention mechanism to obtain the high semantic features corresponding to the global feature information includes: Attention weights are assigned to the global feature information through a channel attention mechanism to obtain a weighted feature map; Performing feature fusion on the weighted feature map to generate a high-resolution feature map; Semantic features are extracted from the high-resolution feature map to obtain corresponding high-level semantic features.
7. The method according to claim 1, characterized in that Outputting a detection result of whether the semantic annotation information is correct according to the high semantic feature includes: Upsampling the high semantic features to obtain sampling characteristic information; According to the sampling characteristic information, a bounding box and a category label of the muck truck are generated using a Faster R-CNN network; Data enhancement is performed on the generated bounding box and category label of the muck truck, and a detection result of whether the enhanced bounding box and category label of the muck truck are correct is output.
8. A muck truck detection system, characterized in that: include: A semantic segmentation module is used to perform semantic segmentation based on the collected video stream monitoring data of each road section in the target area using a pre-built image analysis model to obtain semantic annotation information corresponding to the muck truck in the video stream monitoring data; An information extraction module, used to extract context information according to the semantic annotation information corresponding to the muck truck to obtain global feature information; A feature selection module, used to perform dynamic feature selection on the global feature information through a channel attention mechanism to obtain a high-level semantic feature corresponding to the global feature information; An object detection module is used to output a detection result of whether the semantic annotation information is correct based on the high semantic features; Wherein, the image analysis model is constructed based on the UNet network.
9. The system according to claim 8, characterized in that The semantic segmentation module comprises: A preprocessing submodule is used to preprocess the collected video stream monitoring data of each road section in the target area to obtain preprocessed image data; A data analysis submodule, configured to perform semantic segmentation based on the pre-processed image data using a pre-built image analysis model, and output semantic segmentation results of each image frame in the pre-processed image data; wherein the semantic segmentation results include: a segmentation mask image and an embedded representation; A region extraction submodule, used for extracting pixel region information of the muck truck in the video stream monitoring data from the segmentation mask image; A vector extraction submodule, used for extracting a regional feature vector of the muck truck from the embedded representation; The data integration submodule is used to integrate the pixel area information and the regional feature vector of the muck truck to generate the semantic annotation information of the muck truck.
10. The system according to claim 8, characterized in that The feature selection module comprises: A weight allocation submodule, used to allocate attention weights to the global feature information through a channel attention mechanism to obtain a weighted feature map; A feature fusion submodule, used for performing feature fusion on the weighted feature map to generate a high-resolution feature map; The semantic extraction submodule is used to extract semantic features from the high-resolution feature map to obtain corresponding high-level semantic features.