A real-time monitoring method, device, equipment and medium for excavation volume

Through video stream image analysis, the location and status of excavation equipment and self-dumping equipment are identified, and the excavation volume of earth and stone is calculated, which solves the problems of high cost and low efficiency of real-time monitoring in the existing technology, and achieves low-cost and fast real-time monitoring of excavation volume.

CN119068410BActive Publication Date: 2025-07-11CHINA THREE GORGES CORPORATION
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202410998096.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-07-23
Publication Date
2025-07-11
Estimated Expiration
2044-07-23

AI Technical Summary

Technical Problem

现有技术在土石方明挖作业中,开挖量的实时监测方法需要花费大量人力和时间,且实时性差,无法及时跟进施工进度。

Method used

By collecting video stream images, target detection, identifying the location and status of the mining equipment and self-dumping equipment, determining the excavation operation status of the target area, and calculating the excavation volume based on the positional relationship of the equipment and the rated load capacity to achieve real-time monitoring.

Benefits of technology

It realizes low-cost, fast and real-time monitoring of the excavation volume of earth and rock, saves manpower and data processing time, and improves the real-time and efficiency of construction progress follow-up.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119068410B_ABST
    Figure CN119068410B_ABST
Patent Text Reader

Abstract

The embodiments of the present application provide a method, device, equipment and medium for real-time monitoring of excavation volume. The method specifically includes: performing object detection on the video stream image to obtain the target information corresponding to the excavation equipment and dump truck equipment included in the video stream image; determining N target regions in the video stream image; determining the excavation operation status of the current target region; when the excavation operation status of the current target region is updated from the loading state to the non-excavation state or the excavation state, determining the excavation volume corresponding to the current loading of the current target region according to the rated loading capacity of the dump truck equipment in the current target region, and updating the cumulative excavation volume of the current target region; obtaining the total real-time excavation volume of the construction area to be counted according to the cumulative excavation volume of all current target regions in the N target regions. The embodiments of the present application can realize real-time monitoring of the earthwork excavation volume at a low cost and at a fast speed based on the analysis of the video stream image.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present application relate to the technical field of open cut of earthwork and stonework, and particularly to a method, device, equipment and medium for real-time monitoring of excavation volume. Background Art

[0002] Open cut operation of earthwork and stonework is a very common type of construction operation. Real-time monitoring of the excavation volume helps to timely follow up the on-site construction progress, assist in formulating short-term excavation construction plan arrangements, and timely dispatch the number of construction machinery entering the site.

[0003] In the related art, for the calculation method of excavation volume, during the engineering construction stage, usually before and after excavation, an unmanned aerial vehicle (UAV) oblique photography method is used to obtain a digital elevation model, or a three-dimensional laser scanning technology is used to obtain spatial point cloud data to calculate the excavation volume.

[0004] The above on-site measurement methods not only require certain human and data acquisition costs, but also have the deficiencies of long data processing time and poor real-time performance. Summary of the Invention

[0005] The embodiments of the present application provide a method for real-time monitoring of excavation volume, which can realize real-time monitoring of earthwork and stonework excavation volume at a lower cost and faster speed based on the analysis of video stream images.

[0006] Correspondingly, the embodiments of the present application also provide a device for real-time monitoring of excavation volume, an electronic device and a machine-readable medium to ensure the implementation and application of the above method.

[0007] To solve the above problems, the embodiments of the present application disclose a method for real-time monitoring of excavation volume, including:

[0008] Collecting video stream images of the construction area to be counted;

[0009] Performing object detection on the video stream images to obtain first target information corresponding to the excavation equipment included in the video stream images and second target information corresponding to the dump truck equipment included in the video stream images;

[0010] Determining N target areas in the video stream images according to the first target information and the second target information; N is a positive integer;

[0011] For the current target area among the N target areas, determining the excavation operation state of the current target area according to the operation state of the excavation equipment in the current target area, whether the current target area includes a dump truck equipment, and the positional relationship between the excavation equipment and the dump truck equipment in the current target area; the excavation operation state includes one of: unexcavated state, excavating state, and loading state;

[0012] When the excavation operation status of the current target area is updated from the loading state to the unexcavated state or the excavating state, it is considered that the current target area has completed the current loading. According to the rated loading capacity of the dump equipment corresponding to the current target area, the excavation volume corresponding to the current loading of the current target area is determined, and the cumulative excavation volume of the current target area is updated;

[0013] According to the cumulative excavation volume of all current target areas among the N target areas, the total real-time excavation volume corresponding to the construction area to be statistically calculated is determined.

[0014] An embodiment of the present application also discloses a real-time excavation volume monitoring device, and the device includes:

[0015] An acquisition module, configured to acquire video stream images;

[0016] A target detection module, configured to perform target detection on the video stream images to obtain first target information corresponding to the excavation equipment included in the video stream images and second target information corresponding to the dump equipment included in the video stream images;

[0017] A target area determination module, configured to determine N target areas in the video stream images according to the first target information and the second target information; N is a positive integer;

[0018] An excavation operation status determination module, configured to, for the current target area among the N target areas, determine the excavation operation status of the current target area according to the status of the excavation equipment in the current target area, whether the current target area includes a dump equipment, and the positional relationship between the excavation equipment and the dump equipment in the current target area; the excavation operation status includes one of an unexcavated state, an excavating state, and a loading state;

[0019] A target area excavation volume determination module, configured to, when the operation status of the current target area is updated from the loading state to the unexcavated state or the excavating state, consider that the current target area has completed the current loading. According to the rated loading capacity of the dump equipment corresponding to the current target area, determine the excavation volume corresponding to the current loading of the current target area, and update the cumulative excavation volume of the current target area;

[0020] A construction area to be statistically calculated excavation volume determination module, configured to determine the total real-time excavation volume corresponding to the construction area to be statistically calculated according to the cumulative excavation volume of all current target areas among the N target areas.

[0021] Optionally, the first target information includes: a first bounding box, and the second target information includes: a second bounding box;

[0022] The target area determination module includes:

[0023] A target area quantity determination module, configured to determine the quantity N of target areas according to the quantity N of the first bounding boxes;

[0024] A bounding box matching module, configured to perform one-to-one matching on the first bounding boxes and the second bounding boxes to obtain corresponding matching results;

[0025] A first determination module, configured to, if there is a second bounding box that matches a first bounding box, use the smallest closed rectangular area that contains both the first bounding box and the second bounding box as the target area;

[0026] A second determination module, configured to, if there is no second bounding box that matches a first bounding box, use the rectangular area corresponding to the first bounding box as the target area.

[0027] Optionally, the bounding box matching module includes:

[0028] A distance intersection-over-union metric value determination module, configured to determine the distance intersection-over-union metric value between a first bounding box and a second bounding box;

[0029] A matching determination module, configured to select, according to the distance intersection-over-union metric values between a first bounding box and multiple second bounding boxes, the second bounding box with the largest distance intersection-over-union metric value from the multiple second bounding boxes as the second bounding box that matches the first bounding box.

[0030] Optionally, the excavation operation state determination module includes:

[0031] A first judgment module, configured to judge whether the excavation equipment included in the current target area is in a stationary state to obtain a first judgment result;

[0032] A first determination module, configured to, if the first judgment result is yes, determine that the excavation operation state of the current target area is an unexcavated state;

[0033] A second judgment module, configured to, if the first judgment result is no, judge whether a dump truck is included in the current target area to obtain a second judgment result;

[0034] A second determination module, configured to, if the second judgment result is no, determine that the excavation operation state of the current target area is an in-excavation state;

[0035] A third judgment module, configured to, if the second judgment result is yes, judge whether there is an overlap between the dump truck and the excavation equipment in the current target area to obtain a third judgment result;

[0036] A third determination module, configured to, if the third judgment result is yes, determine that the excavation operation state of the current target area is a loading state;

[0037] The fourth determination module is used to determine that the excavation operation status of the current target area is the in-excavation status if the result of the third judgment is negative.

[0038] Optionally, the third judgment module includes:

[0039] The intersection-over-union index value determination module is used to determine the intersection-over-union index value between the second bounding box corresponding to the dump truck equipment and the first bounding box corresponding to the excavating equipment in the current target area.

[0040] The overlap determination module is used to determine that there is an overlap between the dump truck equipment and the excavating equipment in the current target area if the intersection-over-union index value is greater than the first threshold.

[0041] Optionally, the first judgment module is specifically used for: in the video stream images from the i-th frame to the (i + α)-th frame, if the coordinate differences of the excavating equipment in any two video stream images are both less than the second threshold, it is determined that the excavating equipment included in the current target area is in a stationary state; where i and α are positive integers.

[0042] Optionally, the device further includes:

[0043] The model determination module is used to determine the target model corresponding to the dump truck equipment in the current target area by using the dump truck equipment model classification model.

[0044] The search module is used to search in the mapping relationship between the search model and the rated loading capacity according to the target model to determine the rated loading capacity corresponding to the dump truck equipment in the current target area.

[0045] Optionally, the target detection module is specifically used for inputting the video stream image into the target detection model to obtain the first target information and the second target information output by the target detection model.

[0046] Wherein, the target detection model specifically includes: a backbone network, a neck network, and a detection network.

[0047] The backbone network includes: a first convolution module, a first downsampling module, a first feature extraction module, a second downsampling module, a second feature extraction module, a third downsampling module, a third feature extraction module, a fourth downsampling module, a fourth feature extraction module, and a spatial pyramid pooling module connected in sequence.

[0048] The neck network includes: a first upsampling module, a first splicing module, a fifth feature extraction module, a second upsampling module, a second splicing module, a sixth feature extraction module, a fifth downsampling module, a third splicing module, a seventh feature extraction module, a sixth downsampling module, a fourth splicing module, and an eighth feature extraction module connected in sequence.

[0049] The second feature extraction module is connected to the second splicing module, the third feature extraction module is connected to the first splicing module, the spatial pyramid pooling module is respectively connected to the first upsampling module and the fourth splicing module, and the fifth feature extraction module is connected to the third splicing module; the sixth feature extraction module is connected to the detection network, the seventh feature extraction module is connected to the detection network, and the eighth feature extraction module is connected to the detection network.

[0050] Optionally, the first downsampling module includes: an average pooling module, a first segmentation module, a second convolutional module, a max pooling module, a third convolutional module, and a fifth splicing module;

[0051] Among them, the average pooling module receives input features, the first segmentation module is respectively connected to the average pooling module, the second convolutional module, and the max pooling module, the third convolutional module is connected to the max pooling module, and the fifth splicing module is respectively connected to the second convolutional module and the third convolutional module.

[0052] Optionally, the first feature extraction module includes: a fourth convolutional module, a second segmentation module, a first representative multi-branch module, a fifth convolutional module, a second representative multi-branch module, a sixth convolutional module, a sixth splicing module, and a seventh convolutional module;

[0053] Among them, the fourth convolutional module performs convolutional processing on the input features to obtain a first convolutional result feature; the second segmentation module segments the first convolutional result feature to obtain a first partial feature, a second partial feature, and a third partial feature, where the first partial feature and the second partial feature enter the sixth splicing module, and the third partial feature enters the first representative branch module; the first representative multi-branch module, the fifth convolutional module, the second representative multi-branch module, the sixth convolutional module, and the sixth splicing module are connected in sequence; the sixth splicing module is connected to the seventh convolutional module.

[0054] Optionally, the first representative multi-branch module includes: an eighth convolutional module, a representative bottleneck module, a ninth convolutional module, a seventh splicing module, and a tenth convolutional module;

[0055] Among them, the eighth convolutional module and the ninth convolutional module respectively receive input features, the representative bottleneck module is connected to the eighth convolutional module; the seventh splicing module is respectively connected to the representative bottleneck module and the ninth convolutional module; the tenth convolutional module is connected to the seventh splicing module.

[0056] Optionally, the representative bottleneck module includes: a representative convolutional module and an eleventh convolutional module;

[0057] Among them, the input features are sequentially processed by a representational convolution module and an eleventh convolution module, and the output of the eleventh convolution module is linearly fused with the input features.

[0058] An embodiment of the present application also discloses an electronic device, including: a processor; and a memory, on which executable code is stored, and when the executable code is executed, the processor is caused to execute the method as described in the embodiment of the present application.

[0059] An embodiment of the present application also discloses a machine-readable medium, on which executable code is stored, and when the executable code is executed, a processor is caused to execute the method as described in the embodiment of the present application.

[0060] The embodiment of the present application has the following advantages:

[0061] In the technical solution of the embodiment of the present application, first, a video stream image of a construction area to be counted is collected; then, object detection is performed on the above video stream image to obtain object information corresponding to objects such as excavation equipment and dump trucks in the video stream image; then, according to the above object information, N target areas in the above video stream image are determined; then, according to the objective law between the excavation operation state of the target area and the operation state of the excavation equipment, the presence or absence of a dump truck, and the positional relationship between the excavation equipment and the dump truck, the excavation volume corresponding to the current loading of the current target area is determined, and the cumulative excavation volume of the current target area is updated; furthermore, according to the cumulative excavation volumes of all the current target areas in the above N target areas, the total real-time excavation volume corresponding to the construction area to be counted is determined.

[0062] Since the embodiment of the present application determines N target areas in the video stream image according to the object information corresponding to objects such as excavation equipment and dump trucks in the video stream image, and determines the excavation volume corresponding to the current loading of the current target area according to the objective law between the excavation operation state of the target area and the operation state of the excavation equipment, the presence or absence of a dump truck, and the positional relationship between the excavation equipment and the dump truck; in this way, the embodiment of the present application can continuously determine the target areas existing in the video stream image and the excavation operation states corresponding to the target areas according to the continuously updated video stream image; on this basis, the embodiment of the present application can continuously determine the excavation volume corresponding to the current loading of the current target area in the video stream image and the cumulative excavation volumes of all the current target areas in the video stream image. Therefore, the embodiment of the present application can realize real-time monitoring of the earthwork excavation volume based on the analysis of the video stream image at a relatively low cost and at a relatively fast speed. In other words, the embodiment of the present application can save the labor cost, data acquisition cost, and data processing time cost consumed by using the unmanned aerial vehicle oblique photography or three-dimensional laser scanning technology, and can improve the processing efficiency and real-time performance of the earthwork excavation volume. Description of the Drawings

[0063] Figure 1 It is a schematic diagram of the application environment of the real-time monitoring method for excavation volume in an embodiment of the present application;

[0064] Figure 2 It is a schematic diagram of the step flow of the real-time monitoring method for excavation volume in an embodiment of the present application;

[0065] Figure 3 It is a flowchart of the steps of the method for determining the excavation operation state of the current target area in an embodiment of the present application;

[0066] Figure 4 It is a schematic diagram of the structure of the target detection model in an embodiment of the present application;

[0067] Figure 5 It is a schematic diagram of the structure of the first downsampling module in an embodiment of the present application;

[0068] Figure 6 It is a schematic diagram of the structure of the first feature extraction module in an embodiment of the present application;

[0069] Figure 7 It is a schematic diagram of the structure of the first representative multi-branch module in an embodiment of the present application;

[0070] Figure 8 It is a schematic diagram of the structure of the representative bottleneck module in an embodiment of the present application;

[0071] Figure 9 It is a schematic diagram of the structure of the spatial pyramid pooling module in an embodiment of the present application;

[0072] Figure 10 It is a schematic diagram of the structure of the real-time monitoring device for excavation volume in an embodiment of the present application;

[0073] Figure 11 It is a schematic diagram of the structure of the device provided in an embodiment of the present application. Detailed implementation manners

[0074] To make the above objects, features, and advantages of the present application more obvious and understandable, the present application will be further described in detail below with reference to the accompanying drawings and specific implementation manners.

[0075] The embodiments of the present application can be applied to engineering industries such as hydropower and civil engineering, and are used for real-time monitoring of the corresponding excavation volume for the earth-rock open-pit operation in the engineering industry.

[0076] Taking the pumped-storage power station project as an example, the project construction presents a situation where excavation operations are carried out simultaneously at multiple construction sites. At the same time, due to the continuous change of the on-site construction situation, there is a need to continuously carry out real-time monitoring of the excavation volume for multiple construction sites.

[0077] In the traditional calculation method of the excavation volume in the engineering construction stage, usually, before and after excavation, the digital elevation model of the excavation area is obtained by using the oblique photography method of drones, or the spatial point cloud data of the excavation area is obtained by using the three-dimensional laser scanning technology; then, according to the difference between the digital elevation model or the spatial point cloud data of the excavation area, the volume of the excavation area is calculated; then, according to the volume of the excavation area, the excavation volume is calculated. The above on-site measurement methods not only require certain human and data acquisition costs, but also have the deficiencies of long data processing time and poor real-time performance.

[0078] In view of the technical problems of the traditional technology consuming manpower, data acquisition cost, long processing time and poor real-time performance, the embodiment of the present application provides a real-time monitoring method for the excavation volume, and the method specifically includes the following steps:

[0079] Collect the video stream images of the construction area to be counted;

[0080] Perform object detection on the above video stream images to obtain the first target information corresponding to the excavation equipment included in the above video stream images, and the second target information corresponding to the dump truck equipment included in the above video stream images;

[0081] According to the above first target information and the above second target information, determine N target areas in the above video stream images; N is a positive integer;

[0082] For the current target area among the N target areas, according to the operation status of the excavation equipment in the current target area, whether the current target area contains a dump truck equipment, and the positional relationship between the excavation equipment and the dump truck equipment in the current target area, determine the excavation operation status of the current target area; the above excavation operation status includes: one of the unexcavated state, the excavating state and the loading state;

[0083] When the excavation operation status of the current target area is updated from the loading state to the unexcavated state or the excavating state, it is considered that the current target area has completed the current loading, and according to the rated loading capacity corresponding to the dump truck equipment in the current target area, determine the excavation volume corresponding to the current loading of the current target area, and update the cumulative excavation volume of the current target area;

[0084] According to the cumulative excavation volume of all the current target areas among the above N target areas, determine the total real-time excavation volume corresponding to the construction area to be counted.

[0085] In the embodiments of the present application, first, a video stream image of the construction area to be counted is collected; then, object detection is performed on the above video stream image to obtain object information corresponding to objects such as excavation equipment and dump trucks in the video stream image; next, according to the above object information, N target areas in the above video stream image are determined; then, according to the objective laws between the excavation operation status of the target area and the operation status of the excavation equipment, the presence or absence of dump trucks, and the positional relationship between the excavation equipment and the dump trucks, the excavation volume corresponding to the current loading of the current target area is determined, and the cumulative excavation volume of the current target area is updated; furthermore, according to the cumulative excavation volumes of all current target areas in the above N target areas, the total real-time excavation volume corresponding to the construction area to be counted is determined.

[0086] Since the embodiments of the present application determine N target areas in the video stream image according to the object information corresponding to objects such as excavation equipment and dump trucks in the video stream image, and determine the excavation volume corresponding to the current loading of the current target area according to the objective laws between the excavation operation status of the target area and the operation status of the excavation equipment, the presence or absence of dump trucks, and the positional relationship between the excavation equipment and the dump trucks; in this way, the embodiments of the present application can continuously determine the target areas existing in the video stream image and the excavation operation status corresponding to the target areas according to the continuously updated video stream image; on this basis, the embodiments of the present application can continuously determine the excavation volume corresponding to the current loading of the current target area in the video stream image and the cumulative excavation volumes of all current target areas in the video stream image. Therefore, the embodiments of the present application can realize real-time monitoring of the earthwork excavation volume based on the analysis of the video stream image at a relatively low cost and at a relatively fast speed. In other words, the embodiments of the present application can save the labor cost, data acquisition cost, and data processing time cost consumed by using the unmanned aerial vehicle oblique photography or three-dimensional laser scanning technology, and can improve the processing efficiency and real-time performance of the earthwork excavation volume.

[0087] Refer to Figure 1 , which shows a schematic diagram of the application environment of a method for real-time monitoring of excavation volume according to an embodiment of the present invention. Among them, the image acquisition end 101 and the server end 102 can perform data interaction based on a wireless network or a wired network.

[0088] In practical applications, the image acquisition end 101 may be deployed with an image acquisition device having an image acquisition function such as an image sensor. The image acquisition device can collect construction videos and send the construction videos to the server end 102 at a preset time period.

[0089] After receiving the construction video sent by the image acquisition end 101, the server end 102 can extract the video stream image from the construction video and use the method of the embodiments of the present application to process the video stream image to determine the excavation volumes corresponding to all current target areas in the video stream image of the construction area to be counted.

[0090] It should be noted that as the video stream image is updated, the embodiments of the present application can continuously determine the target areas existing in the video stream image and the corresponding excavation operation status of the target areas. On this basis, the embodiments of the present application can continuously determine the excavation volume corresponding to the current loading of the current target area in the video stream image and the cumulative excavation volume of all current target areas in the video stream image. Therefore, the embodiments of the present application can realize real-time monitoring of the earthwork excavation volume based on the analysis of the video stream image at a relatively low cost and at a relatively fast speed.

[0091] Method Embodiment 1

[0092] Reference Figure 2 , which shows the schematic flow chart of the steps of the real-time monitoring method for excavation volume according to an embodiment of the present application. The method specifically includes the following steps:

[0093] Step 201, collect the video stream image of the construction area to be counted;

[0094] Step 202, perform target detection on the above video stream image to obtain the first target information corresponding to the excavation equipment included in the above video stream image and the second target information corresponding to the dump truck equipment included in the above video stream image;

[0095] Step 203, determine N target areas in the above video stream image according to the above first target information and the above second target information; N is a positive integer;

[0096] Step 204, for the current target area among the N target areas, determine the excavation operation status of the current target area according to the operation status of the excavation equipment in the current target area, whether the current target area includes a dump truck equipment, and the positional relationship between the excavation equipment and the dump truck equipment in the current target area; the above excavation operation status includes: one of the non-excavation status, excavation in progress status, and loading in progress status;

[0097] Step 205, in the case where the excavation operation status of the current target area is updated from the loading in progress status to the non-excavation status or the excavation in progress status, it is considered that the current target area has completed the current loading. Determine the excavation volume corresponding to the current loading of the current target area according to the rated loading capacity of the dump truck equipment in the current target area, and update the cumulative excavation volume of the current target area;

[0098] Step 206, determine the total real-time excavation volume corresponding to the construction area to be counted according to the cumulative excavation volume of all current target areas among the above N target areas.

[0099] Figure 2 At least one step included in the method shown can be executed by the server. It can be understood that the embodiments of the present application Figure 2The specific execution entity of the method shown is not limited.

[0100] In step 201, the process of collecting video stream images specifically includes: after the server receives the construction video sent by the image acquisition end, it extracts video stream images from the construction video. Among them, the server can receive video stream images from the image acquisition end according to a preset time period.

[0101] In the scenario of hydropower engineering, multiple video sources can be set, and different video sources can correspond to different video stream images. Among them, the video source can continuously provide continuous video stream images. The embodiments of the present application can process the continuous video stream images provided by any video source to obtain the excavation volume corresponding to the video source. The construction area to be counted can be the construction area represented by a certain video source.

[0102] In step 202, target detection technology can be used to perform target detection on the above video stream images to obtain the first target information corresponding to the excavation equipment included in the above video stream images and the second target information corresponding to the dump equipment included in the above video stream images.

[0103] Target detection technology is an important branch in the field of computer vision, which aims to identify the target information corresponding to the targets in the image. Specifically in the embodiments of the present application, the video stream images can include targets such as excavation equipment and dump equipment. Examples of excavation equipment can include excavators, etc. Examples of dump equipment can include dump trucks, etc. In practical applications, the dump truck drives into the excavation site and stops within the working radius of the excavator so that the excavator can easily load the materials into the cargo box of the dump truck.

[0104] The first target information and the second target information belong to the category of target information. The target information can correspond to category information. For example, the category corresponding to the first target information is the excavation equipment category, and the category corresponding to the second target information is the dump equipment category.

[0105] The target information specifically includes: the bounding box corresponding to the target. The bounding box is used to represent the position and range of the target in the video stream image. The bounding box is usually a rectangular box, and its four sides are respectively aligned with the outermost edges of the target, thus surrounding the entire target. The information of the bounding box can include: the upper left corner coordinates, width, and height of the rectangular box; or, the information of the bounding box can include: the center point coordinates, width, and height of the rectangular box.

[0106] The specific process of using target detection technology to perform target detection on the above video stream images will be described in detail in Method Embodiment 2.

[0107] In step 203, the target area can be the area in the video stream image that contains an excavation device. According to whether the target area contains a dump truck, the target area can be divided into: a first target area and a second target area. Among them, the first target area contains an excavation device and does not contain a dump truck. The second target area contains an excavation device and a dump truck.

[0108] In an implementation manner of the present application, the first target information includes: a first bounding box, and the second target information includes: a second bounding box;

[0109] The process of step 203 for determining N target areas in the video stream image according to the above first target information and the above second target information specifically includes:

[0110] Step A1: Determine the number N of target areas according to the number N of first bounding boxes;

[0111] Step A2: Perform one-to-one matching on the first bounding box and the second bounding box to obtain corresponding matching results;

[0112] Step A3: If there is a second bounding box that matches the first bounding box, then use the smallest closed rectangle area that contains both the first bounding box and the second bounding box as the target area;

[0113] Step A4: If there is no second bounding box that matches the first bounding box, then use the rectangular area corresponding to the first bounding box as the target area.

[0114] In step A1, the number N of target areas can be equal to the number N of first bounding boxes. In other words, the number of target areas contained in a video stream image can be equal to the number of first bounding boxes contained in a video stream image.

[0115] In step A2, performing one-to-one matching on the first bounding box and the second bounding box can be distance matching between the first bounding box and the second bounding box.

[0116] In an implementation manner, the process of step A2 for performing one-to-one matching on the first bounding box and the second bounding box specifically includes: determining the distance intersection over union (DIoU) index value between a first bounding box and a second bounding box; according to the distance intersection over union index values between a first bounding box and multiple second bounding boxes, selecting the one with the largest distance intersection over union index value from the multiple second bounding boxes as the second bounding box that matches the first bounding box.

[0117] Among them, for the second bounding boxes of M dump trucks in the video stream image, one-to-one matching is performed with the first bounding boxes in the video stream image one by one based on the distance intersection over union index value. The distance intersection over union index value DIoU mnThe determination process is shown in Formula (1).

[0118]

[0119] Among them, A m represents the m-th (m ∈ [1, M], and m is a positive integer) second bounding box, and B n represents the n-th (n ∈ [1, N], and n is a positive integer) first bounding box. a m , b n respectively represent the center points of the second bounding box A m and the first bounding box B n . ρ mn is the Euclidean distance between the center points a m and b n . d is the diagonal distance of the smallest closed region that contains both the first bounding box and the second bounding box.

[0120] The distance intersection over union index value reflects the overlapping degree between the first bounding box and the second bounding box and the distance between the center points. Generally speaking, the larger the distance intersection over union index value, the higher the overlapping degree between the two and the closer the distance between the center points.

[0121] For a first bounding box, there are M distance intersection over union index values between it and the M second bounding boxes. Then, the largest one of the M distance intersection over union index values can be selected. The second bounding box corresponding to the largest distance intersection over union index value can be used as the second bounding box that matches the first bounding box.

[0122] Step A3 and Step A4 can be executed in parallel.

[0123] In practical applications, the N first bounding boxes in the video stream image can be traversed in the first preset order to determine the target regions corresponding to the first bounding boxes one by one. The first preset order can be: the order from left to right and from bottom to top, or the order from left to right and from top to bottom.

[0124] In the embodiments of the present application, the N first bounding boxes in the video stream image can also be numbered in the first preset order. The numbers of the N first bounding boxes can be the same as the numbers of the N target regions.

[0125] During the process of traversing the N first bounding boxes in the video stream image, it can be determined whether there is a second bounding box that matches a first bounding box. If so, Step A3 can be executed to use the smallest closed rectangular region that contains both the first bounding box and the second bounding box as the target region; if not, Step A4 can be executed to use the rectangular region corresponding to the first bounding box as the target region.

[0126] After determining the target area based on step A3 and step A4, information about the target area can be obtained. The information about the target area specifically includes: the number of the target area, the position information of the rectangular frame corresponding to the target area, and information such as whether it contains a second bounding box.

[0127] In step 204, for the current target area among the N target areas, the excavation operation status of the current target area can be determined. Wherein, the current target area can represent one of the N target areas. It can be understood that the embodiments of the present application can process multiple current target areas among the N target areas in parallel or serially. The processing of the current target area can include: the processing of step 204 and step 205. The process of serially processing multiple current target areas among the N target areas specifically includes: traversing the N target areas in a second preset order from smallest to largest or from largest to smallest according to the numbers of the target areas to obtain a current target area among the N target areas, and serially processing multiple current target areas among the N target areas in the second preset order.

[0128] The excavation operation status of the present application specifically includes one of: unexcavated status, excavating status, and loading status; wherein, the unexcavated status means that the current target area has not undergone an excavation operation. The excavating status means that the current target area is undergoing an excavation operation. The loading status means that the current target area is undergoing a loading operation.

[0129] The embodiments of the present application can determine the excavation operation status of the current target area according to the operation status of the excavation equipment in the current target area, whether a dump truck is included in the current target area, and the positional relationship between the excavation equipment and the dump truck in the current target area.

[0130] Refer to Figure 3 , which shows the step flowchart of the method for determining the excavation operation status of the current target area in an embodiment of the present application, and specifically may include the following steps:

[0131] Step 301, determine whether the excavation equipment included in the current target area is in a stationary state to obtain a first judgment result;

[0132] If the first judgment result is yes, then execute step 302: determine that the excavation operation status of the current target area is the unexcavated status;

[0133] If the first judgment result is no, then execute step 303: determine whether a dump truck is included in the current target area to obtain a second judgment result;

[0134] If the second judgment result is no, then execute step 304: determine that the excavation operation status of the current target area is the excavating status;

[0135] If the second judgment result is yes, then perform step 305: Determine whether there is an overlap between the dump truck equipment and the excavation equipment in the current target area to obtain a third judgment result;

[0136] If the third judgment result is yes, then perform step 306: Determine that the excavation operation status of the current target area is the loading state;

[0137] If the third judgment result is no, then perform step 307: Determine that the excavation operation status of the current target area is the excavation state.

[0138] The embodiments of the present application provide the following determination rules for the excavation operation status:

[0139] The determination rule for the unexcavated state is: The excavation equipment is in a stationary state.

[0140] The determination rules for the excavation state include: The excavator is in a non-stationary state and the current target area does not contain dump truck equipment; or, the current target area contains dump truck equipment and there is no overlap between the dump truck equipment and the excavation equipment.

[0141] The determination rules for the loading state include: The excavation equipment is in a non-stationary state, the current target area contains dump truck equipment, and there is an overlap between the dump truck equipment and the excavation equipment.

[0142] The process of step 301 for determining whether the excavation equipment included in the current target area is in a stationary state specifically includes:

[0143] In the i-th frame of the video stream image to the (i + α)-th frame of the video stream image, if the coordinate differences of the excavation equipment in any two frames of the video stream images are all less than the second threshold, it is determined that the excavation equipment included in the current target area is in a stationary state; where both i and α can be positive integers, and α is greater than i.

[0144] The embodiments of the present application can select two video stream images to be compared in the α video stream images corresponding to the i-th frame of the video stream image to the (i + α)-th frame of the video stream image, and determine whether the coordinate difference between the two video stream images to be compared is less than the second threshold.

[0145] Two video stream images to be compared are selected from the α video stream images, and the selection scheme can be P types. If the coordinate differences corresponding to the P selection schemes are all less than the second threshold, it can be determined that the excavation equipment included in the current target area is in a stationary state.

[0146] Considering the influence of signal transmission, camera jitter, etc., when the excavation equipment is in a stationary state, the coordinates of its corresponding first bounding box may have a certain offset. Assume that the upper left pixel coordinate of the first bounding box of the excavation equipment in the i-th frame of the video stream image of the current target area is The lower right pixel coordinates are Set the second threshold corresponding to the coordinate offset as δ and the frame number threshold as α. If formula (2) holds, that is, in the video stream images from the i-th frame to the (i + α)-th frame, the coordinate difference between any two frames is less than the second threshold δ, then it can be determined that the excavation equipment is in a stationary state.

[0147] Formula (2) calculates and judges the coordinate difference for the upper left pixel coordinates of the i-th frame video stream image and the (i + j)-th frame video stream image. Therefore, represents the abscissa of the upper left pixel of the (i + j)-th frame video stream image, represents the ordinate of the upper left pixel of the (i + j)-th frame video stream image.

[0148]

[0149] It can be understood that formula (2) calculates and judges the coordinate difference for the upper left pixel coordinates of the i-th frame video stream image and the (i + j)-th frame video stream image. As an optional embodiment, actually, the coordinate difference can be calculated and judged for the upper right pixel coordinates, lower left pixel coordinates, or lower right pixel coordinates of the i-th frame video stream image and the (i + j)-th frame video stream image.

[0150] Since in step 203, after determining the target area, information such as the number of the target area, the position information of the rectangular frame corresponding to the target area, and whether it contains the second bounding box is obtained. Therefore, step 303 can judge whether the current target area contains a dump truck according to the information corresponding to whether the current target area contains the second bounding box. Of course, target detection technology can be used to perform target detection on the area image corresponding to the current target area to judge whether the current target area contains a dump truck.

[0151] The process of step 305 judging whether there is an overlap between the dump truck and the excavation equipment in the current target area specifically includes:

[0152] Step C1: Determine the intersection over union index value between the second bounding box corresponding to the dump truck and the first bounding box corresponding to the excavation equipment in the current target area;

[0153] Step C2: If the intersection over union index value is greater than the first threshold, it is determined that there is an overlap between the dump truck and the excavation equipment in the current target area.

[0154] Formula (3) shows the calculation process of the intersection over union index value between the second bounding box and the first bounding box.

[0155]

[0156] Wherein, A represents the second bounding box, B represents the first bounding box, A ∩ B represents the intersection area between the second bounding box and the first bounding box, and A ∪ B represents the union area between the second bounding box and the first bounding box.

[0157] In practical applications, when the intersection over union index value is greater than 0, it can be considered that there is an overlap between the dump truck equipment and the excavating equipment in the current target area. Therefore, examples of the first threshold can include: 0 or a real number greater than 0.

[0158] In step 205, for the current target area, the change in its excavation operation state can be monitored. If its excavation operation state is updated from the loading state to the non-excavated state or the excavating state, it can be considered that the current target area has completed the current loading. Then, according to the rated loading capacity corresponding to the dump truck equipment in the current target area, the excavation volume corresponding to the current loading of the current target area is determined, and the cumulative excavation volume of the current target area is updated. The cumulative excavation volume of the current target area can be the sum of the excavation volumes corresponding to multiple loadings of the current target area.

[0159] Wherein, the process of determining the rated loading capacity corresponding to the dump truck equipment in the current target area specifically includes:

[0160] Step D1: Use the dump truck equipment model classification model to determine the target model corresponding to the dump truck equipment in the current target area;

[0161] Step D2: According to the target model, search in the mapping relationship between the search model and the rated loading capacity to determine the rated loading capacity corresponding to the dump truck equipment in the current target area.

[0162] The dump truck equipment model classification model of the embodiment of the present application can have the classification ability of the dump truck equipment model. In other words, it can determine the target model corresponding to the dump truck equipment in the regional image according to the regional image corresponding to the current target area.

[0163] In specific implementation, the embodiment of the present application can label the model categories for the dump truck equipment image dataset, and train the dump truck equipment model classification model according to the dump truck equipment image dataset.

[0164] Here provides a process for obtaining a dump truck equipment image dataset. Specifically, based on the video monitoring system, dump truck equipment images covering different construction scene backgrounds, different lighting conditions, different working postures, etc. can be collected. In addition, combined images of the dump truck equipment and the excavating equipment during the collaborative excavation operation need to be collected to reflect the actual situation of mutual occlusion of the two types of construction machinery. Therefore, the dump truck equipment image dataset can include: dump truck equipment images and combined images.

[0165] The embodiments of the present application do not limit the specific structure of the self-unloading equipment model classification model. For example, the self-unloading equipment model classification model may include structures of convolutional neural networks such as VGG (Visual Geometry Group).

[0166] The embodiments of the present application may pre-store the mapping relationship between the model and the rated loading capacity. In this way, according to the target model of the self-unloading equipment in the current target area, the mapping relationship can be searched to obtain the rated loading capacity of the self-unloading equipment in the current target area.

[0167] In the embodiments of the present application, when it is monitored that the excavation operation state of the current target area is updated from the loading state to the non-excavation state or the excavation state, it can be considered that the current target area has completed the current loading, and according to the rated loading capacity of the self-unloading equipment in the current target area, the excavation volume corresponding to the current loading in the current target area is determined. The excavation volume corresponding to the current loading may be equal to the rated loading capacity of the self-unloading equipment in the current target area.

[0168] In step 206, the cumulative excavation volumes of all current target areas in the N target areas may be summed to obtain the total real-time excavation volume corresponding to the construction area to be counted.

[0169] In an example of the present application, at the start time of a natural day, the total real-time excavation volume corresponding to the construction area to be counted may be set to 0. Subsequently, according to the completion of the current loading of the current target area, the total real-time excavation volume is updated to obtain the continuously accumulated total real-time excavation volume. Therefore, the real-time monitoring process of the excavation volume may be a process of continuously updating the total real-time excavation volume according to the completion of the current loading of the N current target areas.

[0170] In practical applications, if there are multiple construction areas to be counted in a hydropower project, the total real-time excavation volume corresponding to each individual construction area to be counted may be monitored according to natural days. After the end of a natural day, the total real-time excavation volumes corresponding to the multiple construction areas to be counted may be fused. It can be understood that the embodiments of the present application do not limit the specific fusion method.

[0171] In summary, in the excavation volume real-time monitoring method of the embodiment of the present application, according to the target information corresponding to targets such as excavation equipment and dump trucks in the video stream image, N target areas in the video stream image are determined, and according to the objective laws between the excavation operation status of the target area and the operation status of the excavation equipment, the presence of dump trucks, and the positions of the excavation equipment and the dump trucks, the excavation volume corresponding to the current loading of the current target area is determined. In this way, the embodiment of the present application can continuously determine the target areas existing in the video stream image and the operation status corresponding to the target areas according to the continuously updated video stream image. On this basis, the embodiment of the present application can continuously determine the excavation volume corresponding to the current loading of the current target area in the video stream image and the cumulative excavation volume of all current target areas in the video stream image. Therefore, the embodiment of the present application can realize the real-time monitoring of the earthwork excavation volume based on the analysis of the video stream image at a low cost and at a high speed. In other words, the embodiment of the present application can save the labor cost, data acquisition cost, and data processing time cost consumed by using the UAV oblique photography or three-dimensional laser scanning technology, and can improve the processing efficiency and real-time performance of the earthwork excavation volume.

[0172] Method Embodiment 2

[0173] In the embodiment of the present application, a target detection model can be used to perform target detection on the above video stream image to obtain the first target information corresponding to the excavation equipment included in the above video stream image and the second target information corresponding to the dump truck included in the above video stream image. In a specific implementation, the video stream image can be input into the target detection model to obtain the first target information and the second target information output by the target detection model.

[0174] The target detection model is a deep learning model in the field of computer vision. It can identify the targets in the video stream image and determine target information such as the position and category of the targets. The target detection model is usually implemented based on a convolutional neural network.

[0175] The embodiment of the present application does not limit the specific target detection model. For example, examples of the target detection model can include: YOLO (You Only Look Once) series, SSD (SingleShotMultiBox Detector), etc.

[0176] In an optional implementation manner of the embodiment of the present application, a lightweight target detection model is provided. Compared with the existing target detection models, this lightweight target detection model greatly reduces the number of parameters, greatly improves the detection speed, and at the same time, can maintain the detection accuracy, improves the detection speed of dump trucks and excavation equipment in the video surveillance image, and is easy to deploy.

[0177] Reference Figure 4 shows a schematic structural diagram of an object detection model according to an embodiment of the present application. The object detection model specifically includes: a backbone network 401, a neck network 402, and a detection network 403.

[0178] Among them, the backbone network 401 can be used to extract features from input images such as video stream images to obtain image features.

[0179] The backbone network 401 further includes: a first convolution module 411, a first downsampling module 412, a first feature extraction module 413, a second downsampling module 414, a second feature extraction module 415, a third downsampling module 416, a third feature extraction module 417, a fourth downsampling module 418, a fourth feature extraction module 419, and a spatial pyramid pooling module 4110 connected in sequence.

[0180] The neck network 402 can be used to perform fusion processing on the image features output by the backbone network 401 to obtain fused image features.

[0181] The neck network 402 further includes: a first upsampling module 421, a first splicing module 422, a fifth feature extraction module 423, a second upsampling module 424, a second splicing module 425, a sixth feature extraction module 426, a fifth downsampling module 427, a third splicing module 428, a seventh feature extraction module 429, a sixth downsampling module 4210, a fourth splicing module 4211, and an eighth feature extraction module 4212 connected in sequence.

[0182] Among them, the second feature extraction module 415 is connected to the second splicing module 425. In this way, the second splicing module 425 can splice the output of the second feature extraction module 415 and the output of the second upsampling module 424.

[0183] The third feature extraction module 417 is connected to the first splicing module 422. In this way, the first splicing module 422 can splice the output of the third feature extraction module 417 and the first upsampling module 421.

[0184] The spatial pyramid pooling module 4110 is respectively connected to the first upsampling module 421 and the fourth splicing module 4211. In this way, the fourth splicing module 4211 can splice the output of the spatial pyramid pooling module 4110 and the output of the sixth downsampling module 4210.

[0185] The fifth feature extraction module 423 is connected to the third splicing module 428; in this way, the third splicing module 428 can splice the output of the fifth feature extraction module 423 and the output of the fifth downsampling module 427.

[0186] The sixth feature extraction module is connected to the detection network, the seventh feature extraction module is connected to the detection network, and the eighth feature extraction module is connected to the detection network, and is used to output three kinds of fused image features to the detection network.

[0187] The detection network 403 is used to determine a detection result according to the fused image features, and the detection result may include: the foregoing target information.

[0188] The first convolution module of the embodiment of the present application can be used to perform convolution processing on the input image. The convolution structures adopted by convolution modules such as the first convolution module and the second convolution module specifically include: at least one convolution layer, at least one batch normalization layer, and at least one activation function. It can be understood that those skilled in the art can adopt the convolution structure required by the convolution module according to actual application requirements, and the embodiment of the present application does not limit the specific convolution structure.

[0189] Downsampling modules such as the first downsampling module and the second downsampling module can adopt a lightweight downsampling structure to reduce the corresponding number of parameters.

[0190] Refer to Figure 5 , which shows a schematic structural diagram of the first downsampling module of an embodiment of the present application. The first downsampling module specifically includes: an average pooling module 501, a first segmentation module 502, a second convolution module 503, a max pooling module 504, a third convolution module 505, and a fifth splicing module 506;

[0191] Among them, the average pooling module 501 receives the input features, and the first segmentation module 502 is respectively connected to the average pooling module 501, the second convolution module 503, and the max pooling module 504. The first segmentation module 502 divides the output of the average pooling module 501 into: feature A and feature B. Feature A enters the second convolution module 503, and feature B enters the max pooling module 504.

[0192] The third convolution module 505 is connected to the max pooling module 504. The fifth splicing module 506 is respectively connected to the second convolution module 503 and the third convolution module 505, and is used to perform splicing processing on the outputs of the second convolution module 503 and the third convolution module 505.

[0193] Feature extraction modules such as the first feature extraction module and the second feature extraction module can adopt the same representative efficient multi-branch processing structure.

[0194] Refer to Figure 6, showing a schematic structural diagram of a first feature extraction module according to an embodiment of the present application, which specifically includes: a fourth convolution module 601, a second segmentation module 602, a first representative multi-branch module 603, a fifth convolution module 604, a second representative multi-branch module 605, a sixth convolution module 606, a sixth splicing module 607, and a seventh convolution module 608.

[0195] Among them, the fourth convolution module 601 performs convolution processing on the input features to obtain a first convolution result feature; the second segmentation module 602 segments the first convolution result feature to obtain a first part feature, a second part feature, and a third part feature, where the first part feature and the second part feature enter the sixth splicing module 607, and the third part feature enters the first representative branch module 603; the first representative multi-branch module 603, the fifth convolution module 604, the second representative multi-branch module 605, the sixth convolution module 606, and the sixth splicing module 607 are connected in sequence; then the sixth splicing module 607 is used to perform splicing processing on the first part feature, the second part feature, and the output of the sixth convolution module 606.

[0196] The sixth splicing module 607 is connected to the seventh convolution module 608, and the seventh convolution module 608 is used to perform convolution processing on the output of the sixth splicing module 607 to obtain an output feature.

[0197] The first representative multi-branch module and the second representative branch module may adopt the same neural network structure.

[0198] Refer to Figure 7 , showing a schematic structural diagram of a first representative multi-branch module according to an embodiment of the present application, which specifically includes: an eighth convolution module 701, a representative bottleneck module 702, a ninth convolution module 703, a seventh splicing module 704, and a tenth convolution module 705;

[0199] Among them, the eighth convolution module and the ninth convolution module respectively receive input features, and the representative bottleneck module is connected to the eighth convolution module; the seventh splicing module is respectively connected to the representative bottleneck module and the ninth convolution module, and is used to perform splicing processing on the output of the representative bottleneck module 702 and the output of the ninth convolution module 703; the tenth convolution module is connected to the seventh splicing module and is used to perform convolution processing on the output of the seventh splicing module to obtain an output feature.

[0200] Refer to Figure 8 , showing a schematic structural diagram of a representative bottleneck module according to an embodiment of the present application, which specifically includes: a representative convolution module 801 and an eleventh convolution module 802;

[0201] Among them, the input features are successively processed by a representative convolution module and an eleventh convolution module, and the output of the eleventh convolution module is linearly fused with the input features.

[0202] In a specific implementation, the representative convolution module may include: two convolution modules connected in parallel. Assuming that the two convolution modules connected in parallel are respectively: a twelfth convolution module and a thirteenth convolution module, after the input features are respectively subjected to convolution processing by the twelfth convolution module and the thirteenth convolution module, the corresponding convolution processing results are linearly added to obtain output features.

[0203] Refer to Figure 9 , which shows a schematic structural diagram of a spatial pyramid pooling module according to an embodiment of the present application. The spatial pyramid pooling module specifically includes: a fourteenth convolution module 901, p max-pooling modules 902, an eighth connection module 903, and a fifteenth convolution module 904. p can be a positive integer greater than 1, and the value of p in the figure is 3. The spatial pyramid pooling module can achieve deep fusion of input features.

[0204] The detection network may include: at least one detection module. The above detection module can be used to perform classification and regression calculations respectively according to the fused image features by using a convolution module and a convolution layer to obtain the target information included in the video stream image.

[0205] Embodiments of the present application can train a target detection model according to a dump truck equipment image dataset and an excavator equipment image dataset.

[0206] Here, a process for obtaining an excavator equipment image dataset is provided. Specifically, based on a video monitoring system, excavator equipment images covering different construction scene backgrounds, different lighting conditions, different working postures, etc. can be collected. In addition, combined images of dump trucks and excavator equipment during cooperative excavation operations need to be collected to reflect the actual situation of mutual occlusion of the two types of construction machinery. Therefore, the excavator equipment image dataset may include: dump truck equipment images and combined images.

[0207] In an embodiment of the present application, the training process of the target detection model may include: forward propagation and backward propagation.

[0208] Among them, forward propagation (Forward Propagation) can calculate the prediction information of the detection result in sequence according to the parameters of the target detection model in the order from the embedding layer to the processing layer. The prediction information is used to determine the loss information

[0209] Backward Propagation can calculate and update the parameters of the object detection model in sequence according to the loss information, in the order from the output layer to the input layer. The object detection model usually adopts the structure of a neural network, and the parameters of the object detection model can include: parameters such as the weights of the neural network. Among them, during the backward propagation process, the gradient information of the parameters of the object detection model can be determined, and this gradient information can be used to update the parameters of the object detection model. For example, backward propagation can calculate and store the gradient information of the parameters of the object detection model in sequence according to the chain rule in calculus, along the order from the detection network to the neck network and then to the backbone network.

[0210] The object detection model of the embodiment of the present application was experimentally analyzed using the construction scene object detection benchmark dataset. For example, the construction scene object detection benchmark dataset specifically includes: 42,000 construction site images and 13 construction object categories.

[0211] The experimental results show that compared with the existing object detection models, the lightweight multi-object detection model constructed in the embodiment of the present application has a 45.9% reduction in the number of parameters, which is close to half, reducing the model deployment cost. The FPS (Frames Per Second) index has improved, the detection speed has increased, and at the same time, the detection accuracy has not decreased.

[0212] In summary, the lightweight object detection model provided by the embodiment of the present application, compared with the existing object detection models, greatly reduces the number of parameters, greatly improves the detection speed, and at the same time, can maintain the detection accuracy, improves the object detection speed of dump equipment and excavation equipment in the video surveillance image, and is easy to deploy.

[0213] It should be noted that for the method embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should know that the embodiments of the present application are not limited by the described action sequence, because according to the embodiments of the present application, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions involved are not necessarily essential for the embodiments of the present application.

[0214] Based on the above embodiments, this embodiment also provides a real-time excavation volume monitoring device. Refer to Figure 10 The device may specifically include: a collection module 1001, an object detection 1002, a target area determination module 1003, an excavation operation status determination module 1004, a target area excavation volume determination module 1005, and a to-be-statistical construction area excavation volume determination module 1006.

[0215] Among them, the acquisition module 1001 is used to acquire the video stream images of the construction area to be counted;

[0216] The target detection module 1002 is used to perform target detection on the video stream images to obtain the first target information corresponding to the excavation equipment included in the video stream images, and the second target information corresponding to the dump truck equipment included in the video stream images;

[0217] The target area determination module 1003 is used to determine N target areas in the video stream images according to the first target information and the second target information; N is a positive integer;

[0218] The excavation operation status determination module 1004 is used to, for the current target area among the N target areas, determine the excavation operation status of the current target area according to the operation status of the excavation equipment in the current target area, whether the current target area contains a dump truck equipment, and the positional relationship between the excavation equipment and the dump truck equipment in the current target area; the excavation operation status includes one of: the non-excavation status, the excavation-in-progress status, and the loading-in-progress status;

[0219] The target area excavation volume determination module 1005 is used to, when the excavation operation status of the current target area is updated from the loading-in-progress status to the non-excavation status or the excavation-in-progress status, consider that the current target area has completed the current loading, determine the excavation volume corresponding to the current loading of the current target area according to the rated loading capacity of the dump truck equipment in the current target area, and update the cumulative excavation volume of the current target area;

[0220] The excavation volume determination module 1006 of the construction area to be counted is used to determine the total real-time excavation volume corresponding to the construction area to be counted according to the cumulative excavation volumes of all the current target areas among the N target areas.

[0221] Optionally, the first target information specifically includes: a first bounding box, and the second target information specifically includes: a second bounding box;

[0222] The target area determination module includes:

[0223] The target area quantity determination module is used to determine the number N of target areas according to the number N of the first bounding boxes;

[0224] The bounding box matching module is used to perform one-to-one matching on the first bounding boxes and the second bounding boxes to obtain the corresponding matching results;

[0225] The first determination module is used to, if there is a second bounding box that matches a first bounding box, use the smallest closed rectangle area that contains both the first bounding box and the second bounding box as the target area;

[0226] A second determination module, configured to use the rectangular area corresponding to the first bounding box as the target area if there is no second bounding box that matches the first bounding box.

[0227] Optionally, the bounding box matching module includes:

[0228] A distance intersection-over-union metric value determination module, configured to determine the distance intersection-over-union metric value between a first bounding box and a second bounding box.

[0229] A matching determination module, configured to select, according to the distance intersection-over-union metric values between a first bounding box and multiple second bounding boxes, the second bounding box with the largest distance intersection-over-union metric value from the multiple second bounding boxes as the second bounding box that matches the first bounding box.

[0230] Optionally, the excavation operation state determination module includes:

[0231] A first judgment module, configured to judge whether the excavation equipment included in the current target area is in a stationary state to obtain a first judgment result.

[0232] A first determination module, configured to determine that the excavation operation state of the current target area is the non-excavated state if the first judgment result is yes.

[0233] A second judgment module, configured to judge whether a dump truck is included in the current target area if the first judgment result is no to obtain a second judgment result.

[0234] A second determination module, configured to determine that the excavation operation state of the current target area is the excavating state if the second judgment result is no.

[0235] A third judgment module, configured to judge whether there is an overlap between the dump truck and the excavation equipment in the current target area if the second judgment result is yes to obtain a third judgment result.

[0236] A third determination module, configured to determine that the excavation operation state of the current target area is the loading state if the third judgment result is yes.

[0237] A fourth determination module, configured to determine that the excavation operation state of the current target area is the excavating state if the third judgment result is no.

[0238] Optionally, the third judgment module includes:

[0239] An intersection-over-union metric value determination module, configured to determine the intersection-over-union metric value between the second bounding box corresponding to the dump truck and the first bounding box corresponding to the excavation equipment in the current target area.

[0240] An overlap determination module, configured to determine that there is an overlap between the dump truck and the excavator in the current target area if the intersection-over-union index value is greater than a first threshold.

[0241] Optionally, the first determination module is specifically configured to: in the i-th frame to the (i + α)-th frame of the video stream images, if the coordinate differences of the excavator in any two frames of the video stream images are both less than a second threshold, determine that the excavator included in the current target area is in a stationary state; where i and α are positive integers.

[0242] Optionally, the apparatus further includes:

[0243] A model determination module, configured to use a dump truck model classification model to determine the target model corresponding to the dump truck in the current target area;

[0244] A search module, configured to search in the mapping relationship between the search model and the rated loading capacity according to the target model, and determine the rated loading capacity corresponding to the dump truck in the current target area.

[0245] Optionally, the target detection module is specifically configured to input the video stream image into a target detection model to obtain first target information and second target information output by the target detection model;

[0246] Wherein, the target detection model specifically includes: a backbone network, a neck network, and a detection network;

[0247] The backbone network includes: a first convolutional module, a first downsampling module, a first feature extraction module, a second downsampling module, a second feature extraction module, a third downsampling module, a third feature extraction module, a fourth downsampling module, a fourth feature extraction module, and a spatial pyramid pooling module connected in sequence;

[0248] The neck network includes: a first upsampling module, a first splicing module, a fifth feature extraction module, a second upsampling module, a second splicing module, a sixth feature extraction module, a fifth downsampling module, a third splicing module, a seventh feature extraction module, a sixth downsampling module, a fourth splicing module, and an eighth feature extraction module connected in sequence;

[0249] The second feature extraction module is connected to the second splicing module, the third feature extraction module is connected to the first splicing module, the spatial pyramid pooling module is respectively connected to the first upsampling module and the fourth splicing module, and the fifth feature extraction module is connected to the third splicing module; the sixth feature extraction module is connected to the detection network, the seventh feature extraction module is connected to the detection network, and the eighth feature extraction module is connected to the detection network.

[0250] Optionally, the first downsampling module includes: an average pooling module, a first splitting module, a second convolutional module, a max pooling module, a third convolutional module, and a fifth splicing module;

[0251] Among them, the average pooling module receives input features, the first splitting module is respectively connected to the average pooling module, the second convolutional module, and the max pooling module, the third convolutional module is connected to the max pooling module, and the fifth splicing module is respectively connected to the second convolutional module and the third convolutional module.

[0252] Optionally, the first feature extraction module includes: a fourth convolutional module, a second splitting module, a first representative multi-branch module, a fifth convolutional module, a second representative multi-branch module, a sixth convolutional module, a sixth splicing module, and a seventh convolutional module;

[0253] Among them, the fourth convolutional module performs convolutional processing on the input features to obtain a first convolutional result feature; the second splitting module splits the first convolutional result feature to obtain a first partial feature, a second partial feature, and a third partial feature, where the first partial feature and the second partial feature enter the sixth splicing module, and the third partial feature enters the first representative branch module; the first representative multi-branch module, the fifth convolutional module, the second representative multi-branch module, the sixth convolutional module, and the sixth splicing module are connected in sequence; the sixth splicing module is connected to the seventh convolutional module.

[0254] Optionally, the first representative multi-branch module includes: an eighth convolutional module, a representative bottleneck module, a ninth convolutional module, a seventh splicing module, and a tenth convolutional module;

[0255] Among them, the eighth convolutional module and the ninth convolutional module respectively receive input features, the representative bottleneck module is connected to the eighth convolutional module; the seventh splicing module is respectively connected to the representative bottleneck module and the ninth convolutional module; the tenth convolutional module is connected to the seventh splicing module.

[0256] Optionally, the representative bottleneck module includes: a representative convolutional module and an eleventh convolutional module;

[0257] Among them, the input features are sequentially processed by the representative convolutional module and the eleventh convolutional module, and the output of the eleventh convolutional module is linearly fused with the input features.

[0258] In summary, for the real-time excavation volume monitoring device according to the embodiments of the present application, based on the target information corresponding to targets such as excavation equipment and dump trucks in the video stream image, N target areas in the video stream image are determined, and according to the objective laws between the excavation operation status of the target area and the operation status of the excavation equipment, the presence or absence of dump trucks, and the positional relationship between the excavation equipment and the dump trucks, the excavation volume corresponding to the current loading of the current target area is determined. In this way, the embodiments of the present application can continuously determine the target areas existing in the video stream image and the operation status corresponding to the target areas according to the continuously updated video stream image. On this basis, the embodiments of the present application can continuously determine the excavation volume corresponding to the current loading of the current target area in the video stream image and the cumulative excavation volume of all current target areas in the video stream image. Therefore, the embodiments of the present application can achieve real-time monitoring of the earthwork excavation volume based on the analysis of the video stream image at a relatively low cost and at a relatively fast speed. In other words, the embodiments of the present application can save the labor cost, data acquisition cost, and data processing time cost consumed by using the drone oblique photography or 3D laser scanning technology, and can improve the processing efficiency and real-time performance of the earthwork excavation volume.

[0259] The embodiments of the present application further provide a non-volatile readable storage medium, in which one or more modules (programs) are stored. When the one or more modules are applied to a device, the device can be caused to execute the instructions (instructions) of the various method steps in the embodiments of the present application.

[0260] The embodiments of the present application provide one or more machine-readable media, on which instructions are stored. When executed by one or more processors, the electronic device is caused to execute the methods as described in one or more of the above embodiments. In the embodiments of the present application, the electronic device includes various types of devices such as terminal devices and servers (clusters).

[0261] The embodiments of the present disclosure can be implemented as a device configured as desired using any suitable hardware, firmware, software, or any combination thereof. The device may include: electronic devices such as terminal devices and servers (clusters). Figure 11 Schematically illustrated is an exemplary device 1300 that can be used to implement the various embodiments described in the present application.

[0262] For one embodiment, Figure 11An exemplary apparatus 1300 is shown, which has one or more processors 1302, a control module (chipset) 1304 coupled to at least one of the one or more processors 1302, a memory 1306 coupled to the control module 1304, a NVM (non-volatile memory) / storage device 1308 coupled to the control module 1304, one or more input / output devices 1310 coupled to the control module 1304, and a network interface 1312 coupled to the control module 1304.

[0263] The processor 1302 may include one or more single-core or multi-core processors, and the processor 1302 may include any combination of general-purpose processors or dedicated processors (such as graphics processors, application processors, baseband processors, etc.). In some embodiments, the apparatus 1300 can act as devices such as the terminal devices, servers (clusters), etc. described in the embodiments of the present application.

[0264] In some embodiments, the apparatus 1300 may include one or more computer-readable media (such as the memory 1306 or the non-volatile memory / storage device 1308) having instructions 1314 and one or more processors 1302 combined with the one or more computer-readable media and configured to execute the instructions 1314 to implement modules so as to perform the actions described in the present disclosure.

[0265] For one embodiment, the control module 1304 may include any suitable interface controller to provide any suitable interface to at least one of the one or more processors 1302 and / or any suitable device or component communicating with the control module 1304.

[0266] The control module 1304 may include a memory controller module to provide an interface to the memory 1306. The memory controller module may be a hardware module, a software module, and / or a firmware module.

[0267] The memory 1306 can be used, for example, to load and store data and / or instructions 1314 for the apparatus 1300. For one embodiment, the memory 1306 may include any suitable volatile memory, such as a suitable DRAM (Dynamic Random Access Memory). In some embodiments, the memory 1306 may include a double data rate type four synchronous dynamic random access memory.

[0268] For one embodiment, the control module 1304 may include one or more input / output controllers to provide an interface to the non-volatile memory / storage device 1308 and the one or more input / output devices 1310.

[0269] For example, the non-volatile memory / storage device 1308 can be used to store data and / or instructions 1314. The non-volatile memory / storage device 1308 can include any suitable non-volatile memory (e.g., flash memory) and / or can include any suitable (one or more) non-volatile storage devices (e.g., one or more hard disk drives, one or more optical disk drives, and / or one or more digital versatile disk drives).

[0270] The non-volatile memory / storage device 1308 can include storage resources that are physically part of the device on which the device 1300 is mounted, or it can be accessed by the device without being part of the device. For example, the non-volatile memory / storage device 1308 can be accessed via the network through the (one or more) input / output devices 1310.

[0271] (One or more) input / output devices 1310 can provide an interface for the device 1300 to communicate with any other suitable device. The input / output devices 1310 can include communication components, audio components, sensor components, etc. The network interface 1312 can provide an interface for the device 1300 to communicate through one or more networks. The device 1300 can wirelessly communicate with one or more components of a wireless network according to any standard and / or protocol among one or more wireless network standards and / or protocols. For example, it can access a wireless network based on a communication standard, such as WiFi (Wireless Fidelity), 2G (2-Generation wireless telephone technology), 3G (3-Generation wireless telephone technology), 4G (4-Generation wireless telephone technology), 5G (5-Generation wireless telephone technology), etc., or a combination thereof for wireless communication.

[0272] For one embodiment, at least one of the (one or more) processors 1302 may be logically encapsulated with one or more controllers (e.g., a memory controller module) of the control module 1304. For one embodiment, at least one of the (one or more) processors 1302 may be logically encapsulated with one or more controllers of the control module 1304 to form a system-in-package. For one embodiment, at least one of the (one or more) processors 1302 may be logically integrated with one or more controllers of the control module 1304 on the same die. For one embodiment, at least one of the (one or more) processors 1302 may be logically integrated with one or more controllers of the control module 1304 on the same die to form a system-on-chip.

[0273] In various embodiments, the device 1300 may be, but is not limited to, a server, a desktop computing device, or a mobile computing device (e.g., a laptop computing device, a handheld computing device, a touchscreen device, a netbook, etc.) and other terminal devices. In various embodiments, the device 1300 may have more or fewer components and / or a different architecture. For example, in some embodiments, the device 1300 includes one or more cameras, a keyboard, a liquid crystal display screen (including a touchscreen display), a non-volatile memory port, multiple antennas, a graphics chip, an application specific integrated circuit, and a speaker.

[0274] Among them, a main control chip can be used as a processor or a control module in the detection device, sensor data, location information, etc. are stored in a memory or a non-volatile memory / storage device, the sensor group can be used as an input / output device, and the communication interface may include a network interface.

[0275] For the device embodiments, since they are basically similar to the method embodiments, the description is relatively simple. For the relevant parts, please refer to the partial description of the method embodiments.

[0276] Each embodiment in this specification is described in a progressive manner. Each embodiment focuses on the differences from other embodiments. For the same and similar parts among the embodiments, please refer to each other.

[0277] Embodiments of the present application are described with reference to the flowcharts and / or block diagrams of methods, terminal devices (systems), and computer program products according to embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram, and the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing terminal device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing terminal device generate a device for implementing the functions specified in one or more flows of the flowchart and / or one or more blocks of the block diagram.

[0278] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing terminal device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device that implements the functions specified in one or more flows of the flowchart and / or one or more blocks of the block diagram.

[0279] These computer program instructions can also be loaded onto a computer or other programmable data processing terminal device, so that a series of operation steps are executed on the computer or other programmable terminal device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable terminal device provide steps for implementing the functions specified in one or more flows of the flowchart and / or one or more blocks of the block diagram.

[0280] Although the preferred embodiments of the embodiments of the present application have been described, those skilled in the art can make additional changes and modifications to these embodiments once they know the basic creative concepts. Therefore, the appended claims are intended to be construed as including the preferred embodiments and all changes and modifications that fall within the scope of the embodiments of the present application.

[0281] Finally, it should also be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or terminal device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or terminal device. Without further limitation, an element defined by the statement "including a..." does not exclude the presence of additional identical elements in the process, method, article or terminal device including the said element.

[0282] The above has introduced in detail a real-time excavation volume monitoring method and device, an electronic device, and a machine-readable medium provided by the present application. Specific examples are used in this article to elaborate on the principle and implementation manner of the present application. The description of the above embodiments is only used to help understand the method and its core idea of the present application; at the same time, for those of ordinary skill in the art, according to the idea of the present application, there will be changes in the specific implementation manner and application scenarios. In summary, the content of this specification should not be construed as a limitation to the present application.

Claims

1. A real-time monitoring method for excavation volume, characterized in that, The method includes: Collecting video stream images of the construction area to be counted; Inputting the video stream images into a target detection model for target detection to obtain first target information corresponding to the excavation equipment included in the video stream images output by the target detection model and second target information corresponding to the dump truck equipment included in the video stream images; Determining N target areas in the video stream images according to the first target information and the second target information; N is a positive integer; For the current target area among the N target areas, determining the excavation operation state of the current target area according to the operation state of the excavation equipment in the current target area, whether the current target area contains a dump truck equipment, and the positional relationship between the excavation equipment and the dump truck equipment in the current target area; the excavation operation state includes one of: an unexcavated state, an excavating state, and a loading state; When the excavation operation state of the current target area is updated from the loading state to the unexcavated state or the excavating state, it is considered that the current target area has completed the current loading. According to the rated loading capacity corresponding to the dump truck equipment in the current target area, determining the excavation volume corresponding to the current loading of the current target area, and updating the cumulative excavation volume of the current target area; Determining the total real-time excavation volume corresponding to the construction area to be counted according to the cumulative excavation volumes of all the current target areas among the N target areas; Wherein, the target detection model specifically includes: a backbone network, a neck network, and a detection network; The backbone network includes: a first convolution module, a first downsampling module, a first feature extraction module, a second downsampling module, a second feature extraction module, a third downsampling module, a third feature extraction module, a fourth downsampling module, a fourth feature extraction module, and a spatial pyramid pooling module connected in sequence; The first downsampling module includes: an average pooling module, a first segmentation module, a second convolution module, a max pooling module, a third convolution module, and a fifth splicing module; Wherein, the average pooling module receives input features, the first segmentation module is respectively connected to the average pooling module, the second convolution module, and the max pooling module, the third convolution module is connected to the max pooling module, and the fifth splicing module is respectively connected to the second convolution module and the third convolution module; The first feature extraction module includes: a fourth convolution module, a second segmentation module, a first representative multi-branch module, a fifth convolution module, a second representative multi-branch module, a sixth convolution module, a sixth splicing module, and a seventh convolution module; Among them, the fourth convolution module performs convolution processing on the input features to obtain the first convolution result features; the second segmentation module segments the first convolution result features to obtain the first partial features, the second partial features, and the third partial features, where the first partial features and the second partial features enter the sixth splicing module, and the third partial features enter the first representational branch module; the first representational multi-branch module, the fifth convolution module, the second representational multi-branch module, the sixth convolution module, and the sixth splicing module are connected in sequence; the sixth splicing module is connected to the seventh convolution module; The neck network includes: a first upsampling module, a first splicing module, a fifth feature extraction module, a second upsampling module, a second splicing module, a sixth feature extraction module, a fifth downsampling module, a third splicing module, a seventh feature extraction module, a sixth downsampling module, a fourth splicing module, and an eighth feature extraction module that are connected in sequence; The second feature extraction module is connected to the second splicing module, the third feature extraction module is connected to the first splicing module, the spatial pyramid pooling module is respectively connected to the first upsampling module and the fourth splicing module, and the fifth feature extraction module is connected to the third splicing module; the sixth feature extraction module is connected to the detection network, the seventh feature extraction module is connected to the detection network, and the eighth feature extraction module is connected to the detection network.

2. The method according to claim 1, wherein The first target information includes: a first bounding box, and the second target information includes: a second bounding box; Determining the N target regions in the video stream image according to the first target information and the second target information includes: Determining the number N of target regions according to the number N of the first bounding boxes; Performing one-to-one matching on the first bounding box and the second bounding box to obtain corresponding matching results; If there is a second bounding box that matches the first bounding box, then taking the smallest closed rectangular region that contains both the first bounding box and the second bounding box as the target region; If there is no second bounding box that matches the first bounding box, then taking the rectangular region corresponding to the first bounding box as the target region.

3. The method according to claim 2, wherein The performing one-to-one matching on the first bounding box and the second bounding box includes: Determining the distance intersection over union index value between a first bounding box and a second bounding box; According to the distance intersection over union index values between a first bounding box and multiple second bounding boxes, selecting the one with the largest distance intersection over union index value from the multiple second bounding boxes as the second bounding box that matches the first bounding box.

4. The method according to claim 1, characterized in that For the current target region among the N target regions, determining the excavation operation state of the current target region according to the operation state of the excavation equipment in the current target region, whether the current target region contains a dump truck, and the positional relationship between the excavation equipment and the dump truck in the current target region includes: Judging whether the excavation equipment included in the current target region is in a stationary state to obtain a first judgment result; If the first judgment result is yes, then determining that the excavation operation state of the current target region is the unexcavated state; If the first judgment result is negative, then determine whether the current target area contains a self-unloading device to obtain a second judgment result; If the second judgment result is negative, then determine that the excavation operation state of the current target area is the in-excavation state; If the second judgment result is positive, then determine whether there is an overlap between the self-unloading device and the excavation device in the current target area to obtain a third judgment result; If the third judgment result is positive, then determine that the excavation operation state of the current target area is the in-loading state; If the third judgment result is negative, then determine that the excavation operation state of the current target area is the in-excavation state.

5. The method according to claim 4, wherein The determination of whether there is an overlap between the self-unloading device and the excavation device in the current target area includes: Determine the intersection-over-union index value between the second bounding box corresponding to the self-unloading device and the first bounding box corresponding to the excavation device in the current target area; If the intersection-over-union index value is greater than the first threshold, then determine that there is an overlap between the self-unloading device and the excavation device in the current target area.

6. The method according to claim 4, characterized in that The determination of whether the excavation device included in the current target area is in a stationary state includes: In the i-th to (i + α)-th video stream images, if the coordinate differences of the excavation device in any two video stream images are both less than the second threshold, then determine that the excavation device included in the current target area is in a stationary state; where i and α are positive integers.

7. The method according to claim 1, wherein The method further includes: Using a self-unloading device model classification model, determine the target model corresponding to the self-unloading device in the current target area; According to the target model, search in the mapping relationship between the search model and the rated loading capacity to determine the rated loading capacity corresponding to the self-unloading device in the current target area.

8. The method according to claim 1, wherein The first representative multi-branch module includes: an eighth convolutional module, a representative bottleneck module, a ninth convolutional module, a seventh splicing module, and a tenth convolutional module; Among them, the eighth convolutional module and the ninth convolutional module respectively receive input features, and the representative bottleneck module is connected to the eighth convolutional module; the seventh splicing module is respectively connected to the representative bottleneck module and the ninth convolutional module; the tenth convolutional module is connected to the seventh splicing module.

9. The method according to claim 8, characterized in that, The representative bottleneck module includes: a representative convolutional module and an eleventh convolutional module; Among them, the input feature is sequentially processed by the representative convolutional module and the eleventh convolutional module, and the output of the eleventh convolutional module is linearly fused with the input feature.

10. A real-time excavation volume monitoring device, characterized in that, The device includes: An acquisition module, configured to acquire video stream images of a construction area to be counted; A target detection module, configured to input the video stream images into a target detection model for target detection to obtain first target information corresponding to the excavation device included in the video stream images output by the target detection model, and second target information corresponding to the self-unloading device included in the video stream images; A target area determination module, configured to determine N target areas in the video stream images according to the first target information and the second target information; N is a positive integer; An excavation operation status determination module, which is used to determine the excavation operation status of the current target area among N target areas according to the operation status of the excavation equipment in the current target area, whether there is a self-unloading equipment in the current target area, and the positional relationship between the excavation equipment and the self-unloading equipment in the current target area; the excavation operation status includes one of: non-excavation status, excavation in progress status, and loading in progress status; A target area excavation volume determination module, which is used to consider that the current target area has completed the current loading when the excavation operation status of the current target area is updated from the loading in progress status to the non-excavation status or the excavation in progress status, determine the excavation volume corresponding to the current loading of the current target area according to the rated loading capacity of the self-unloading equipment in the current target area, and update the cumulative excavation volume of the current target area; A total real-time excavation volume determination module for the construction area to be statistically analyzed, which is used to determine the total real-time excavation volume corresponding to the construction area to be statistically analyzed according to the cumulative excavation volumes of all current target areas among the N target areas; Among them, the target detection model specifically includes: a backbone network, a neck network, and a detection network; The backbone network includes: a first convolution module, a first downsampling module, a first feature extraction module, a second downsampling module, a second feature extraction module, a third downsampling module, a third feature extraction module, a fourth downsampling module, a fourth feature extraction module, and a spatial pyramid pooling module connected in sequence; The first downsampling module includes: an average pooling module, a first splitting module, a second convolution module, a max pooling module, a third convolution module, and a fifth splicing module; Among them, the average pooling module receives input features, the first splitting module is respectively connected to the average pooling module, the second convolution module, and the max pooling module, the third convolution module is connected to the max pooling module, and the fifth splicing module is respectively connected to the second convolution module and the third convolution module; The first feature extraction module includes: a fourth convolution module, a second splitting module, a first representative multi-branch module, a fifth convolution module, a second representative multi-branch module, a sixth convolution module, a sixth splicing module, and a seventh convolution module; Among them, the fourth convolution module performs convolution processing on the input features to obtain a first convolution result feature; the second splitting module splits the first convolution result feature to obtain a first part feature, a second part feature, and a third part feature, where the first part feature and the second part feature enter the sixth splicing module, and the third part feature enters the first representative branch module; the first representative multi-branch module, the fifth convolution module, the second representative multi-branch module, the sixth convolution module, and the sixth splicing module are connected in sequence; the sixth splicing module is connected to the seventh convolution module; The neck network includes: a first upsampling module, a first splicing module, a fifth feature extraction module, a second upsampling module, a second splicing module, a sixth feature extraction module, a fifth downsampling module, a third splicing module, a seventh feature extraction module, a sixth downsampling module, a fourth splicing module, and an eighth feature extraction module, which are connected in sequence; The second feature extraction module is connected to the second splicing module, the third feature extraction module is connected to the first splicing module, the spatial pyramid pooling module is connected to the first upsampling module and the fourth splicing module respectively, and the fifth feature extraction module is connected to the third splicing module; the sixth feature extraction module is connected to the detection network, the seventh feature extraction module is connected to the detection network, and the eighth feature extraction module is connected to the detection network.

11. An electronic device, characterized in that, Comprising: A processor; And A memory having stored thereon executable code that, when executed, causes the processor to perform the method according to any one of claims 1-9.

12. A machine-readable medium having stored thereon executable code that, when executed, causes a processor to perform the method according to any one of claims 1-9.

Citation Information

Patent Citations

  • Excavator and excavator assist device

    CN110446817A