Shield muck volume calculation method and device based on machine vision
Through a dual-mode image segmentation network based on machine vision, the shield soil volume is extracted and fused, and the residue mask image and bounding box are generated, which solves the accuracy problem of shield soil volume calculation under high real-time requirements, and real-time and high-precision shield soil volume calculation is achieved.
Patent Information
- Application Number
- CN202510884133.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-30
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2045-06-30
AI Technical Summary
The existing non-contact measurement technology for shield soil volume is low under high real-time requirements, resulting in insufficient accuracy in the calculation of shield soil volume.
Using a machine vision-based method, the real-time acquisition target depth image and RGB image is extracted and fused through a dual-modal input image segmentation network to generate the coordinates of the slag mask image and the slag bounding box, and the volume calculation method is used to calculate the volume of the shield slag.
Real-time and high-precision calculation of shield soil volume is realized, the calculation accuracy and speed of shield soil volume is improved, and the requirements of high real-timeness are met.
Smart Images

Figure CN120374708A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of image processing, and particularly to a method and device for calculating the volume of shield muck based on machine vision. Background Art
[0002] A shield machine (Tunnel Boring Machine, TBM) is a large-scale underground engineering construction device integrating functions such as rock tunnel excavation, support, ventilation and dust removal, and can form a full-section tunnel at one time by means of hob rock breaking. Compared with the traditional drill and blast method, the tunneling speed of the shield method can reach 3 to 10 times that of the drill and blast method, and the tunnel forming quality is high, and the environmental disturbance is small. At present, shield technology is evolving towards the intelligent direction. By integrating cutting-edge technologies such as machine vision and deep learning, efforts are made to break through key problems such as adaptive tunneling in complex strata and unmanned construction. In dealing with the problem of muck generated during tunneling, real-time monitoring of the muck volume flow can not only reflect the tunneling efficiency, but also establish a rock-soil-equipment interaction feedback mechanism to improve the engineering safety and intelligent level.
[0003] At present, measuring the volume of shield muck in the muck truck carriage through a camera is a common non-contact measurement technology for the volume of shield muck. It takes images of the muck truck carriage through a camera, segments the captured images to obtain a muck mask image, and finally calculates the volume of shield muck according to the muck mask image. However, the inventors of the present application found that: The existing non-contact measurement method for the volume of shield muck has a low segmentation accuracy for the images of the muck truck carriage under high real-time requirements, resulting in a low accuracy of the calculated volume of shield muck under high real-time requirements. Summary of the Invention
[0004] The purpose of the present application is to provide a method and device for calculating the volume of shield muck based on machine vision to solve the problem that the accuracy of the volume of shield muck calculated by the existing non-contact measurement technology for the volume of shield muck is low under high real-time requirements.
[0005] To achieve the above purpose, the present application provides the following solutions: In a first aspect, the present application provides a method for calculating the volume of shield muck based on machine vision, including: Performing feature extraction and fusion on a real-time acquired target depth image and a target RGB image through an image segmentation network based on dual-modal input, and segmenting the target RGB image according to the obtained fusion features to obtain a muck mask image of the target RGB image and the coordinates of the muck bounding box; Calculating the volume of shield muck according to the coordinates of the muck bounding box and the muck mask image through a volume calculation method.
[0006] In a second aspect, the present application provides a computer device, including: a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the method for calculating the volume of shield muck based on machine vision described in any one of the above.
[0007] According to the specific embodiments provided by the present application, the following technical effects are disclosed: The present application provides a method and device for calculating the volume of shield muck based on machine vision. The method includes: obtaining a target depth image and a target RGB image in real time through a depth camera, and performing feature extraction and fusion on the obtained target depth image and target RGB image through an image segmentation network based on dual-modal input. Then, image segmentation is performed on the target RGB image according to the obtained fusion features to obtain the muck mask image of the target RGB image and the coordinates of the muck bounding box. Finally, the volume of the shield muck is calculated through a volume calculation method according to the coordinates of the muck bounding box and the muck mask image. Therefore, the real-time calculation of the volume of the shield muck is realized, and the real-time performance of the calculation result is ensured. Since the boundary colors of the muck and the muck truck carriage are similar, but the depth information is different, the image segmentation network based on dual-modal input supports the input of dual-modal images (RGB images and depth images). By performing feature extraction and fusion on the obtained target depth image and target RGB image through the image segmentation network based on dual-modal input, complementary fusion of dual-modal features (RGB features and depth features) can be performed, and the obtained fusion features can more comprehensively cover the features of the muck and the muck truck carriage, making the segmentation network more accurate. Furthermore, by performing image segmentation on the target RGB image according to the obtained fusion features through the image segmentation network based on dual-modal input, the segmentation accuracy of the target RGB image can be improved, and then the accuracy of the muck mask image of the target RGB image and the coordinates of the muck bounding box can be improved. Thus, the accuracy of the volume of the shield muck calculated through the volume calculation method according to the coordinates of the muck bounding box and the muck mask image can be improved. In summary, the present application realizes the real-time and high-precision calculation of the volume of the shield muck, and solves the problem that the volume of the shield muck calculated by the existing non-contact measurement technology for the volume of the shield muck has low accuracy under high real-time requirements. BRIEF DESCRIPTION OF THE DRAWINGS
[0008] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0009] Figure 1Schematic flowchart of a method for calculating the volume of shield muck based on machine vision provided by an embodiment of the present application; Figure 2 Schematic structural diagram of the backbone network of an image segmentation network based on dual-modal input provided by an embodiment of the present application; Figure 3 Schematic structural diagram of an image segmentation network based on dual-modal input provided by an embodiment of the present application; Figure 4 Provided by an embodiment of the present application Figure 2 Schematic structural diagram of the first dual-modal feature fusion module; Figure 5 Schematic diagram of the shape of the first infinitesimal provided by an embodiment of the present application; Figure 6 Schematic diagram of the shape of the second infinitesimal provided by an embodiment of the present application; Figure 7 Schematic diagram of the functional modules of a device for calculating the volume of shield muck based on machine vision provided by an embodiment of the present application; Figure 8 Schematic structural diagram of a computer device provided by an embodiment of the present application. Detailed implementation manners
[0010] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the protection scope of the present application.
[0011] To make the above objects, features, and advantages of the present application more obvious and understandable, the present application will be further described in detail below with reference to the accompanying drawings and specific implementation manners.
[0012] In an exemplary embodiment, as Figure 1 shown, a method for calculating the volume of shield muck based on machine vision is provided. This method is executed by a computer device and includes the following steps 101 to 102. Among them: Step 101: Through an image segmentation network based on dual-modal input, perform feature extraction and fusion on the real-time acquired target depth image and target RGB image, and perform image segmentation on the target RGB image according to the obtained fusion features to obtain the muck mask image of the target RGB image and the coordinates of the muck bounding box.
[0013] In the embodiments of the present application, the target depth image and the target RGB image are the depth image and the RGB image obtained by using a depth camera installed at the slag outlet of the shield machine to photograph the carriage of the muck truck at the slag outlet of the shield machine.
[0014] The RGB image can provide appearance features such as texture and color, but it is prone to incorrect segmentation when faced with objects with similar colors and textures; the depth image can reflect the depth information of the object from the depth camera and provide geometric structure features. Since the boundary colors of the muck and the muck truck carriage are similar, but the depth information is different, the image segmentation network based on bimodal input supports the input of bimodal images (RGB image and depth image). Through the image segmentation network based on bimodal input, feature extraction and fusion are performed on the target depth image and the target RGB image, and complementary fusion of bimodal features (RGB features and depth features) can be achieved. The obtained fusion features can more comprehensively cover the features of the muck and the muck truck carriage, making the segmentation network more accurate. Furthermore, through the image segmentation network based on bimodal input, image segmentation is performed on the target RGB image according to the obtained fusion features, which can improve the segmentation accuracy of the target RGB image, reduce the error between the muck mask image and the muck area in the target RGB image, and further improve the accuracy of calculating the volume of the shield muck. The muck bounding box refers to the smallest rectangle that encloses the muck area in the target RGB image. The coordinates of the muck bounding box include the coordinates of the four vertices of the muck bounding box.
[0015] Step 102, calculate the volume of the muck according to the coordinates of the muck bounding box and the muck mask image through a volume calculation method.
[0016] In the embodiments of the present application, the volume calculation method is not limited and can be set according to actual needs. For example, an existing volume calculation method is used to calculate the volume of the shield muck.
[0017] By implementing the above-mentioned Step 101 to Step 102, since the target depth image and the target RGB image are obtained in real time, and through the image segmentation network based on bimodal input, feature extraction and fusion are performed on the target depth image and the target RGB image obtained in real time, and the target RGB image is segmented according to the obtained fusion features, the muck mask image of the target RGB image and the coordinates of the muck bounding box are obtained, and according to the coordinates of the muck bounding box and the muck mask image, the volume of the shield muck is calculated by the volume calculation method. Therefore, the real-time calculation of the shield muck volume is realized, and the real-time performance of the calculation result is ensured. Since the boundary color between the muck and the muck truck carriage is similar, but the depth information is different, the image segmentation network based on bimodal input supports the input of bimodal images (RGB image and depth image). Through the image segmentation network based on bimodal input, feature extraction and fusion are performed on the target depth image and the target RGB image obtained in real time, and the complementary fusion of bimodal features (RGB feature and depth feature) can be performed. The obtained fusion features can more comprehensively cover the features of the muck and the muck truck carriage, making the segmentation network more accurate. Furthermore, through the image segmentation network based on bimodal input, the target RGB image is segmented according to the obtained fusion features, which can improve the segmentation accuracy of the target RGB image, and further improve the accuracy of the muck mask image of the target RGB image and the coordinates of the muck bounding box. Thus, the accuracy of the volume of the shield muck calculated by the volume calculation method according to the coordinates of the muck bounding box and the muck mask image can be improved. In summary, this application realizes the real-time and high-precision calculation of the shield muck volume, and solves the problem that the volume of the shield muck calculated by the existing non-contact measurement technology of the shield muck volume has low accuracy under the requirement of high real-time performance. In addition, the segmentation speed of the target RGB image is improved through the image segmentation network based on bimodal input (with reduced number of parameters), and further the calculation speed of the shield muck volume is improved.
[0018] In another exemplary embodiment of this application, in order to improve the segmentation accuracy of the image segmentation network based on bimodal input, the image segmentation network based on bimodal input in the above-mentioned Step 101 includes an input layer, a backbone network, a neck network, and a segmentation layer connected in sequence. Among them, the input layer is used to merge the channels of each input target RGB image and the corresponding target depth image, and convert the feature map obtained by merging the channels into a matrix , B is the batch dimension, which is the number of target RGB images (the same as the number of target depth images) input into the image segmentation network based on bimodal input at one time, H is the width of the feature map obtained by merging the channels, W is the length of the feature map obtained by merging the channels. The backbone network is used to process the matrix Perform dual-modal feature extraction and fusion to obtain shallow features and deep features; among them, the dual-modal features include depth image features and RGB image features. The neck network is used to perform cross-scale fusion operations on the shallow features and deep features output by the backbone network, and obtain fusion features with the same scales as the shallow features and deep features respectively. The segmentation layer is used to generate and output at least the coordinates of the muck bounding box and the muck mask image of the target RGB image according to the fusion features output by the neck network.
[0019] In the embodiment of the present application, the target depth image corresponding to each target RGB image refers to the depth image collected at the same time as each target RGB image. Specifically, the target depth image corresponding to each target RGB image refers to the depth image with the same name as each target RGB image.
[0020] In another exemplary embodiment of the present application, in order to further improve the segmentation accuracy of the improved image segmentation algorithm, as Figure 2 shown, the above-mentioned backbone network includes a first feature extraction module 201, a second feature extraction module 202, a first dual-modal feature fusion module 203, a third feature extraction module 204, a fourth feature extraction module 205, a second dual-modal feature fusion module 206, a fifth feature extraction module 207, a sixth feature extraction module 208, a third dual-modal feature fusion module 209, and an SPPF spatial pyramid pooling module 2010. Among them, the first feature extraction module 201 is used to separate the target depth image part of the matrix and perform feature extraction on the target depth image part to obtain depth image features. The second feature extraction module 202 is used to separate the matrix The target RGB image part, and perform feature extraction on the target RGB image part to obtain RGB image features. The first bimodal feature fusion module 203 is used to perform feature fusion on the bimodal features output by the first feature extraction module 201 and the second feature extraction module 202 based on a triple attention mechanism to obtain the first shallow feature (P3). The third feature extraction module 204 is used to perform feature extraction on the output of the first feature extraction module 201. The fourth feature extraction module 205 is used to perform feature extraction on the output of the first bimodal feature fusion module 203. The second bimodal feature fusion module 206 is used to perform feature fusion on the bimodal features output by the third feature extraction module 204 and the fourth feature extraction module 205 based on a triple attention mechanism to obtain the second shallow feature (P4). The fifth feature extraction module 207 is used to perform feature extraction on the output of the third feature extraction module 204. The sixth feature extraction module 208 is used to perform feature extraction on the output of the second bimodal feature fusion module 206. The third bimodal feature fusion module 209 is used to perform feature fusion on the bimodal features output by the fifth feature extraction module 207 and the sixth feature extraction module 208 based on a triple attention mechanism. The SPPF spatial pyramid pooling module 2010 is used to perform pooling operations of different sizes on the output of the third bimodal feature fusion module 209, and splice the feature maps obtained by the pooling operations of different sizes in the channel dimension to obtain the above-mentioned deep feature (P5).
[0021] In another exemplary embodiment of the present application, in order to further improve the segmentation accuracy of the improved image segmentation algorithm, as Figure 3 shown, the first feature extraction module 201 includes a seventh feature extraction module and an eighth feature extraction module. The seventh feature extraction module includes a first separation module D-part, a first convolution module, a second convolution module, and a first cross-stage feature fusion module (C2f) connected in sequence. Among them, the first separation module D-part is used to separate the matrix of the target depth image part. The first convolution module is used to perform a convolution operation on the feature map output by the first separation module D-part. The second convolution module is used to perform a convolution operation on the feature map output by the first convolution module. The first cross-stage feature fusion module (C2f) is used to perform feature extraction of different sizes on the feature map output by the second convolution module and perform cross-stage feature fusion on the feature maps of different scales obtained by the extraction.
[0022] The eighth feature extraction module includes a third convolution module and a second cross-stage feature fusion module. The third convolution module is used to perform a convolution operation on the output of the first cross-stage feature fusion module. The second cross-stage feature fusion module is used to perform feature extraction of different sizes on the feature map output by the third convolution module and perform cross-stage feature fusion on the feature maps of different scales obtained by the extraction.
[0023] In another exemplary embodiment of the present application, in order to further improve the segmentation accuracy of the improved image segmentation algorithm, as Figure 3 shown, the second feature extraction module 202 includes the first feature extraction module 201, the ninth feature extraction module, the fourth bimodal feature fusion module ( Figure 3 the first BTA of ) and the tenth feature extraction module. The ninth feature extraction module includes a second separation module RGB-part, a fourth convolutional module, a fifth convolutional module, and a third cross-stage feature fusion module (C2f) connected in sequence. Among them, the second separation module RGB-part is used to separate the target RGB image part of the matrix
[0024] In the embodiment of the present application, the output of the second cross-stage feature fusion module and the output of the fourth bimodal feature fusion module are used as the input of the first bimodal feature fusion module, and the first bimodal feature fusion module outputs P3.
[0025] The feature map sizes output by the second separation module RGB-part and the first separation module D-part are the same, the feature map sizes output by the first cross-stage feature fusion module and the third cross-stage feature fusion module are the same. The feature map sizes output by the second cross-stage feature fusion module and the fourth cross-stage feature fusion module are the same. The fourth bimodal feature fusion module has the same structure as the first bimodal feature fusion module, the second bimodal feature fusion module, and the third bimodal feature fusion module.
[0026] In another exemplary embodiment of the present application, in order to further improve the segmentation accuracy of the improved image segmentation algorithm, as Figure 3As shown, the third feature extraction module 204 includes a seventh convolutional module and a fifth cross-stage feature fusion module. The seventh convolutional module is used to perform a convolutional operation on the output of the second cross-stage feature fusion module. The fifth cross-stage feature fusion module is used to perform feature extraction of different sizes on the output of the seventh convolutional module and perform cross-stage feature fusion on the feature maps of different scales obtained by the extraction.
[0027] The fourth feature extraction module 205 includes an eighth convolutional module and a sixth cross-stage feature fusion module. The eighth convolutional module is used to perform a convolutional operation on the output of the first bimodal feature fusion module. The sixth cross-stage feature fusion module is used to perform feature extraction of different sizes on the output of the eighth convolutional module and perform cross-stage feature fusion on the feature maps of different scales obtained by the extraction.
[0028] In the embodiment of the present application, the sizes of the feature maps output by the fifth cross-stage feature fusion module and the sixth cross-stage feature fusion module are the same.
[0029] The output of the fifth cross-stage feature fusion module and the output of the sixth cross-stage feature fusion module are used as the input of the second bimodal feature fusion module, and the second bimodal feature fusion module outputs P4.
[0030] In another exemplary embodiment of the present application, in order to further improve the segmentation accuracy of the improved image segmentation algorithm, as Figure 3 shown, the fifth feature extraction module 207 includes a ninth convolutional module and a seventh cross-stage feature fusion module. The ninth convolutional module is used to perform a convolutional operation on the output of the fifth cross-stage feature fusion module. The seventh cross-stage feature fusion module is used to perform feature extraction of different sizes on the output of the ninth convolutional module and perform cross-stage feature fusion on the feature maps of different scales obtained by the extraction.
[0031] The sixth feature extraction module 208 includes a tenth convolutional module and an eighth cross-stage feature fusion module. The tenth convolutional module is used to perform a convolutional operation on the output of the second bimodal feature fusion module. The seventh cross-stage feature fusion module is used to perform feature extraction of different sizes on the output of the tenth convolutional module and perform cross-stage feature fusion on the feature maps of different scales obtained by the extraction.
[0032] In the embodiment of the present application, the sizes of the feature maps output by the seventh cross-stage feature fusion module and the eighth cross-stage feature fusion module are the same.
[0033] The output of the seventh cross-stage feature fusion module and the output of the eighth cross-stage feature fusion module are used as the input of the second bimodal feature fusion module.
[0034] In another exemplary embodiment of the present application, in order to further improve the segmentation accuracy of the improved image segmentation algorithm, as Figure 3As shown, the SPPF spatial pyramid pooling module 2010 includes an eleventh convolution module, multiple serial max pooling layers, a first connection layer, and a twelfth convolution module in sequence. Among the multiple serial max pooling layers, except for the first max pooling layer, the other max pooling layers are used to splice the output of the eleventh convolution module and the output of the previous max pooling layer and then further perform max pooling operations.
[0035] In the embodiment of the present application, the twelfth convolution module outputs the deep feature P5.
[0036] In another exemplary embodiment of the present application, in order to further improve the segmentation accuracy of the improved image segmentation algorithm, as Figure 3 shown, the Neck network includes a second connection layer, a ninth cross-stage feature fusion module, a thirteenth convolution module, a third connection layer, a tenth cross-stage feature fusion module, a first upsampling module, a fourth connection layer, an eleventh cross-stage feature fusion module, a fourteenth convolution module, a second upsampling module, a fourth connection layer, an eleventh cross-stage feature fusion module, a fourteenth convolution module, a second upsampling module, a fifth connection layer, and a twelfth cross-stage feature fusion module. Among them, the second upsampling module is used to upsample the output of the SPPF spatial pyramid pooling module 2010. The third connection layer is used to splice the output of the second upsampling module and the output of the second bimodal feature fusion module. The tenth cross-stage feature fusion module is used to extract features of different sizes from the output of the third connection layer and perform cross-stage feature fusion on the extracted feature maps of different scales. The first upsampling module is used to upsample the output of the tenth cross-stage feature fusion module. The second connection layer is used to splice the output of the first bimodal feature fusion module and the output of the first upsampling module. The ninth cross-stage feature fusion module is used to extract features of different sizes from the output of the second connection layer and perform cross-stage feature fusion on the extracted feature maps of different scales. The thirteenth convolution module is used to perform convolution operations on the output of the ninth cross-stage feature fusion module. The fourth connection layer is used to splice the output of the thirteenth convolution module and the output of the tenth cross-stage feature fusion module. The eleventh cross-stage feature fusion module is used to extract features of different sizes from the output of the fourth connection layer and perform cross-stage feature fusion on the extracted feature maps of different scales. The fourteenth convolution module is used to perform convolution operations on the output of the eleventh cross-stage feature fusion module. The fifth connection layer is used to splice the output of the second upsampling module and the output of the fourteenth convolution module. The twelfth cross-stage feature fusion module is used to extract features of different sizes from the output of the fifth connection layer and perform cross-stage feature fusion on the extracted feature maps of different scales.
[0037] In the embodiment of the present application, the outputs of the ninth cross-stage feature fusion module, the eleventh cross-stage feature fusion module, and the twelfth cross-stage feature fusion module are used as the inputs of the segmentation layer.
[0038] In another exemplary embodiment of the present application, in order to further improve the segmentation accuracy of the improved image segmentation algorithm, as Figure 3 shown, the segmentation layer includes three Segment modules. The first Segment module is used to perform a convolution operation on the input first shallow feature, predict the muck bounding box, and output the pixel coordinates of the four vertices of the predicted muck bounding box. The second Segment module is used to perform a convolution operation on the input second shallow feature, predict and output the category of the target object in each target RGB image (the category includes muck, and if the category label of muck is set to 0). The third Segment module is used to perform a convolution operation on the input deep feature, predict and output the mask coefficient of each target RGB image.
[0039] In another exemplary embodiment of the present application, as Figure 3 shown, the structures of the first cross-stage feature fusion module to the twelfth cross-stage feature fusion module are the same, and each includes a fifteenth convolution module, a Split module, a plurality of serial BottleNeck modules, a sixth connection layer, and a sixteenth convolution module connected in sequence. The spilt module is used to divide the output of the fifteenth convolution module (with a shape of I B×C×H×W ) along the channel dimension C into two parts, I1 B×C / 2×H×W and I2 B×C / 2×H×W . I1 B×C / 2×H×W is input into the first BottleNeck module, and the output of each BottleNeck module will be used as the input of the next BottleNeck module. The output of each BottleNeck module and I2 B×C / 2×H×W output by the Split module are concatenated in the sixth connection layer and then input into the sixteenth convolution module.
[0040] In the embodiment of the present application, each BottleNeck module includes a seventeenth convolution module, an eighteenth convolution module, and a first adder connected in sequence. The other input of the adder is the input of the seventeenth convolution module of this BottleNeck module.
[0041] In another exemplary embodiment of the present application, the structures of the three Segment modules are the same, as Figure 3 shown, and each includes three convolution branches. Each convolution branch includes a nineteenth convolution module, a twentieth convolution module, and a first convolution layer connected in sequence.
[0042] In the embodiments of the present application, for Segments of three sizes, a bounding box (Bbox) tensor with a shape of [B, N, 4], a class (Class) tensor with a shape of [B, N, num] (where B is the batch size, N represents the number of targets in each image, 4 represents the coordinates of the bounding box of the target, and num represents the number of target classes to be generated), and a mask coefficient (Masks Coefficients) representing target-specific weights with a shape of [B, N, Q] are generated. Q is the number of channels of the mask prototype. The deep features pass through a convolutional layer (Conv2d) to generate a prototype mask. The prototype mask is a general mask template, and all instance masks are reconstructed based on the linear combination of the prototype mask and the mask coefficients. The bounding box data generated by the segment modules of the three sizes are concatenated and then filtered by a threshold (Conf) to remove the bounding boxes with low confidence. Then, non-maximum suppression (NMS) is performed to screen the overlapping prediction boxes for each target, and finally, the best prediction box (bbox) of the target is retained for output. For the class data (Class) generated by the segment modules of the three sizes, after concatenation, the class with the highest score (argmax) is taken as the class output (class). For the mask coefficients generated by the segment modules of the three sizes, after concatenation respectively, matrix multiplication is performed between the prototype mask and the concatenated mask coefficients (Masks Coefficients). After passing through the activation function Sigmoid, it is reshaped into an instance mask with a shape of [B, N, H, W]. The instance mask is scaled to the size of the target RGB image through a cropping operation (Crop), and then through a binarization operation (Threshold), the pixel value of the target muck pixels is set to 255 (white), and the background pixel value is set to 0 (black), and the final mask image is output.
[0043] In another exemplary embodiment of the present application, the first dual-modal feature fusion module 203, the second dual-modal feature fusion module 206, the third dual-modal feature fusion module 209, and the fourth dual-modal feature fusion module have the same structure, and each includes a first attention weight extraction module, a second attention weight extraction module, an adder module, a first attention feature map extraction module, a second attention feature map extraction module, and a seventh connection layer. Among them, the first attention weight extraction module is used to extract a first attention weight, a second attention weight, and a third attention weight. The first attention weight refers to the correlation weight between the channel dimension and the width of the input RGB feature map. The second attention weight refers to the correlation weight between the channel dimension and the length of the input RGB feature map. The third attention weight refers to the spatial attention weight of the input RGB feature map. The second attention weight extraction module is used to extract a fourth attention weight, a fifth attention weight, and a sixth attention weight. The fourth attention weight refers to the correlation weight between the channel dimension and the width of the input depth feature map. The fifth attention weight refers to the correlation weight between the channel dimension and the length of the input depth feature map. The sixth attention weight refers to the spatial attention weight of the input depth feature map. The adder module is used to add the first attention weight and the fourth attention weight to obtain a first combined attention weight, add the second attention weight and the fifth attention weight to obtain a second combined attention weight, and add the third attention weight and the sixth attention weight to obtain a third combined attention weight. The first attention feature map extraction module is used to perform average weighted fusion on the RGB feature map by using the first combined attention weight, the second combined attention weight, and the third combined attention weight to obtain an RGB attention feature map. The second attention feature map extraction module is used to perform average weighted fusion on the depth feature map by using the first combined attention weight, the second combined attention weight, and the third combined attention weight to obtain a depth attention feature map. The seventh connection layer is used to splice the RGB attention feature map and the depth attention feature map to achieve cross-dimensional and multi-dimensional weight distribution of the extracted feature maps, which can enhance the model's utilization of multi-modal features and improve the performance of the segmentation task.
[0044] In the embodiment of the present application, the input RGB feature map and depth feature map refer to the RGB feature map and depth feature map input to the first dual-modal feature fusion module to the fourth dual-modal feature fusion module.
[0045] The core idea of the first dual-modal feature fusion module 203, the second dual-modal feature fusion module 206, the third dual-modal feature fusion module 209, and the fourth dual-modal feature fusion module is to dynamically capture the dependencies in each modality by separately analyzing the feature interactions in the three dimensions of channel-width, channel-height, and space of each modality (RGB feature map modality and depth feature map modality), and then applying the idea of shared weights to calculate the joint attention, and then applying the attention to each modality feature map to achieve feature fusion between modalities.
[0046] In another exemplary embodiment of the present application, the first attention weight extraction module includes a first weight extraction branch, a second weight extraction branch, and a third weight extraction branch, and the second attention weight extraction module includes a fourth weight extraction branch, a fifth weight extraction branch, and a sixth weight extraction branch. Among them, the structures of the first weight extraction branch and the sixth weight extraction branch are the same, as Figure 4 shown, both include a first dimension adjustment module (Permute), a first pooling module (Zpool), and a second convolutional layer (Conv). Among them, the first dimension adjustment module is used to rearrange the order of each dimension of the input RGB feature map or depth feature map according to the following formula: ; Among them, I is the input RGB feature map or the depth feature map, and its shape is , 0 represents I the batch size of B , 1 represents I the length of H , 2 represents I the number of channels of C , 3 represents I the width of W , I cw is the output of the first dimension adjustment module, and its shape is .
[0047] The first pooling module is used to perform a pooling operation on I cw . The second convolutional layer is used to perform a convolutional operation on the output of the first pooling module.
[0048] The structures of the second branch and the fifth branch are the same, and both include a second dimension adjustment module, a second pooling module, and a third convolutional layer (Conv). Among them, the second dimension adjustment module is used to rearrange the order of each dimension of the input RGB feature map or depth feature map according to the following formula: ; Among them, I hcis the output of the second - dimension adjustment module, and its shape is .
[0049] The second pooling module is used to perform a pooling operation on I hc . The third convolutional layer is used to perform a convolutional operation on the output of the second pooling module.
[0050] The structures of the third branch and the fourth branch are the same, and both include a third pooling module and a fourth convolutional layer (Conv). The third pooling module is used to perform a pooling operation on the input RGB feature map or depth feature map. The fourth convolutional layer is used to perform a convolutional operation on the output of the third pooling module.
[0051] In another exemplary embodiment of the present application, as Figure 4 shown, the structures of the first pooling module to the third pooling module are the same, and all include a max - pooling layer, an average - pooling layer, and an eighth connection layer. Among them, the max - pooling layer and the average - pooling layer are in parallel. The max - pooling layer is used to perform a max - pooling operation on the input RGB feature map or depth feature map. The average - pooling layer is used to perform an average - pooling operation on the input RGB feature map or depth feature map. The eighth connection layer is used to splice the outputs of the max - pooling layer and the average - pooling layer.
[0052] In another exemplary embodiment of the present application, as Figure 4 shown, the adder module includes a second adder, a third adder, and a fourth adder. Among them, the second adder is used to add the first attention weight AR cw and the fourth attention weight AD cw to obtain the first joint attention weight AS cw . The third adder is used to add the second attention weight AR hc and the fifth attention weight AD hc to obtain the second joint attention weight AS hc . The fourth adder is used to add the third attention weight AR hw and the sixth attention weight AD hw to obtain the third joint attention weight AS hw .
[0053] In the embodiment of the present application, , , .
[0054] In another exemplary embodiment of the present application, the structures of the first attention feature map extraction module and the second attention feature map extraction module are the same, as Figure 4As shown, all include a first activation function module, a second activation function module, a third activation function module, a first multiplier, a second multiplier, a third multiplier, a fifth adder, and a sixth adder. Among them, the inputs of the first activation function module, the second activation function module, and the third activation function module correspond to the first joint attention weight AS cw , the second joint attention weight AS hc , and the third joint attention weight AS hw . The first inputs of the first multiplier, the second multiplier, and the third multiplier correspond to the outputs of the first activation function module, the second activation function module, and the third activation function module respectively. The second inputs of the first multiplier, the second multiplier, and the third multiplier are the RGB feature map input to the first attention weight extraction module or the depth feature map input to the second attention weight extraction module. The inputs of the fifth adder and the sixth adder are the outputs of the first multiplier, the second multiplier, and the third multiplier, and the outputs of the fifth adder and the sixth adder are the inputs of the seventh connection layer.
[0055] In another exemplary embodiment of the present application, the structures of the first convolutional module to the twentieth convolutional module are the same. As Figure 3 shown, all include a fifth convolutional layer, a batch normalization layer (BatchNorm2d), and an activation function (SiLu). Among them, the fifth convolutional layer is used to extract local features of the input feature map, the batch normalization layer is used to prevent gradient explosion and gradient disappearance, and the activation function (SiLu) improves the expression ability of the network through non-linear transformation.
[0056] In another exemplary embodiment of the present application, step 102 described above includes the following steps 301 to 306. Among them: Step 301, in each target depth image, form a rectangular sliding window starting from the upper left vertex of the muck bounding box.
[0057] Step 302, perform a micro-element classification operation. Among them, the micro-element classification operation includes: If the pixel depth value of any one of the four vertices of the rectangular sliding window is missing, the pixel area included in the rectangular sliding window is recorded as a missing micro-element, and the volume of the missing micro-element is recorded as 0.
[0058] If none of the four vertices of the rectangular sliding window have missing pixel depth values, and at the same time: If all the pixel points within the rectangular sliding window belong to muck pixel points, the pixel area included in the rectangular sliding window is recorded as a first micro-element; among them, if the pixel value of any pixel point within the rectangular sliding window in the corresponding muck mask image is 255, then the pixel point within the rectangular sliding window belongs to muck pixel points; If the number of pixels belonging to muck pixels in the rectangular sliding window is less than a preset first threshold, the pixel region included in the rectangular sliding window is recorded as an invalid microelement, and the volume of the invalid microelement is recorded as nan. If the number of pixels belonging to muck pixels in the rectangular sliding window is not less than the preset first threshold but less than n, the pixel region included in the rectangular sliding window is recorded as a second microelement, where n is the total number of pixels in the rectangular sliding window.
[0059] In the embodiments of the present application, the size, moving step, and first threshold of the rectangular sliding window are not specifically limited and can be set according to actual needs. For example, the size of the rectangular sliding window is set to 3 ×3. At this time, the vertical and horizontal moving steps of the rectangular sliding window are both 3, n is 9, and the first threshold can be set to 5 at this time.
[0060] Step 303, after performing the above microelement classification operation, perform a microelement volume calculation operation: Calculate the volume of the first microelement according to the three-dimensional space coordinate system of the four vertices of the first microelement. Calculate the volume of the second microelement according to the four vertices of the second microelement and the three-dimensional space coordinate system of the muck pixels.
[0061] Step 304, after performing the microelement volume calculation operation, perform a loop operation: In each target depth image, move the rectangular sliding window in sequence in the moving direction from left to right and from top to bottom until the rectangular sliding window moves to the final position, and perform the microelement classification operation and the microelement volume calculation operation after each movement of the rectangular sliding window; the final position is the position where the coordinate of the lower right vertex of the rectangular sliding window exceeds the lower right vertex of the muck bounding box.
[0062] Step 305, after performing the loop operation, perform a volume assignment operation: Assign a volume to each missing microelement of each target depth image according to the volumes of the first microelement and / or the second microelement adjacent to each missing microelement of each target depth image.
[0063] In the embodiments of the present application, since the shape of the muck falling into the carriage at the slag outlet will not have a particularly prominent microelement, and the heights between adjacent microelements are relatively similar, it is selected to fill the microelement volume with the volumes of the first microelement and / or the second microelement adjacent to each missing microelement after completely calculating the microelements within the muck bounding box by the rectangular sliding window, which can improve the accuracy of volume calculation. By accumulating the volumes of all non-invalid microelements, an approximate value of the captured volume can be obtained.
[0064] Step 306, after performing the volume assignment operation, accumulate the volumes of all missing micro-elements, first micro-elements, and second micro-elements in each target depth image to obtain the volume of shield muck in each target depth image.
[0065] In the embodiment of the present application, the volumes of all missing micro-elements, first micro-elements, and second micro-elements in each target depth image are accumulated through the following formula to obtain the volume of shield muck in each target depth image: ; where, is the total number of missing micro-elements, is the volume value of the th missing micro-element after assignment, is the total number of first micro-elements, is the volume value of the th first micro-element, is the total number of second micro-elements, is the volume value of the th second micro-element, V E is the volume of shield muck in each target depth image.
[0066] In another exemplary embodiment of the present application, calculating the volume of the first micro-element according to the three-dimensional space coordinate system of the four vertices of the first micro-element includes the following steps 401 to 402. Wherein: Step 401, convert the pixel coordinates of the four vertices of the first micro-element to the three-dimensional space coordinate system to obtain the spatial three-dimensional coordinates of the four vertices of the first micro-element.
[0067] Step 402, calculate the volume of the first micro-element according to the spatial three-dimensional coordinates of the four vertices of the first micro-element and the height of the first micro-element according to the volume calculation method of a quadrangular prism; wherein, the average height of the four vertices of the first micro-element is the absolute value of the difference between the first height value and the second height value, the first height value is the average value of the spatial Z-axis coordinate values of the four vertices of the first micro-element, and the second height value is the height of the depth camera from the bottom of the muck truck carriage.
[0068] In the embodiment of the present application, according to the number of pixels belonging to muck in the sliding window, the muck micro-elements in the sliding window can be regarded as different shapes. The shape of the first micro-element is as Figure 5 shown, approximately a quadrangular prism, with different heights at the four points, and two triangles with A2B2C2 and B2C2D2 as the bottom surfaces can be formed in the projection on the vehicle bottom. The vertical height Z 5 of the vehicle bottom from the center of the depth camera can be obtained during installation, then the volume of each first micro-element is the sum of the volumes of two triangular prisms , and the volume calculation of the first micro-element is shown in the following formula: ; ; ; Among them, S 1 represents the area of triangle A2B2C2 on the bottom surface, S 2 represents the area of triangle B2C2D2 on the bottom surface, and the coordinates of point A1 are , the coordinates of point B1 are , the coordinates of point C1 are , the coordinates of point C1 are , the height of the depth camera from the bottom of the vehicle frame is . Multiply the bottom surface area of triangle A2B2C2 and the bottom surface area of triangle B2C2D2, and then multiply by the average height of the four points A1, B1, C1, and C1, which is the volume of the first microelement.
[0069] In another exemplary embodiment of the present application, calculating the volume of the second microelement according to the three-dimensional space coordinate system of the four vertices of the second microelement and the muck pixel points includes the following steps 501 to step 502. Among them: Step 501: Convert the pixel coordinates of the four vertices of the second microelement and the muck pixel points to the three-dimensional space coordinate system to obtain the three-dimensional space coordinates of the four vertices of the second microelement and the muck pixel points.
[0070] Step 502: Calculate the volume of the second microelement according to the three-dimensional space coordinates of the four vertices of the second microelement and the height of the second microelement; among them, the height of the second microelement is the absolute value of the difference between the third height value and the second height value, and the third height value is the average value of the spatial Z-axis coordinate values of all muck pixel points of the second microelement.
[0071] In the embodiment of the present application, the shape of the second microelement is as Figure 6 shown, and the volume of the second microelement is as shown in the following formula: ; Among them, avg(Z) represents the average height of the pixel points belonging to the muck pixel points in the second microelement (the average value of the Z w of the soil pixel points).
[0072] In another exemplary embodiment of the present application, convert the pixel coordinates of the target pixel points to the three-dimensional space coordinate system according to the following steps 601 to step 603. The target pixel points include any vertex of the first microelement, any vertex of the second microelement, or any muck pixel point of the second microelement. Among them: Step 601: Convert the pixel coordinates of the target pixel point to the camera coordinate system to obtain the coordinates of the target pixel point in the camera coordinate system. The conversion formula is as follows: ; where, ( u , v ) represents the pixel coordinates of the target pixel point, represents the coordinates of the optical axis center of the depth camera in the pixel coordinate system, is the focal length of the depth camera on the x-axis (unit: pixel), is the focal length of the depth camera on the y-axis (unit: pixel), ( , , ) represents the coordinates of the target pixel point in the camera coordinate system, and the Z-axis coordinate of the target pixel point in the camera coordinate system is the depth value obtained by the depth camera.
[0073] In the embodiments of the present application, first, convert the pixel coordinates of the target pixel point to the image coordinate system according to the following formula to obtain the coordinates of the target pixel point in the image coordinate system. The conversion formula is as follows: ; where, and respectively represent the scale factors for converting the actual distances in the x-axis direction and y-axis direction in the image coordinate system into pixel numbers, and ( x , y ) represents the coordinates of the target pixel point in the image coordinate system.
[0074] Then, convert the coordinates of the target pixel point in the image coordinate system to the camera coordinate system according to the following formula to obtain the coordinates of the target pixel point in the camera coordinate system. The conversion formula is as follows: ; where, and represent the focal lengths of the x-axis and y-axis of the depth camera (unit: mm).
[0075] After the above three coordinate system conversions, the calculation formula for the position ( , , ) of the target pixel point relative to the center of the depth camera in the camera coordinate system can be obtained.
[0076] The pixel coordinates ( u , v ) are the coordinates describing the position of the target pixel point in the target depth image in a pixel coordinate system with the upper left corner of the target depth image as the origin (0, 0), where the abscissa and ordinate represent the number of rows and columns respectively. The coordinates of the image coordinate system (x , y ), in a coordinate system with the optical axis center of the camera as the coordinate origin, the horizontal direction perpendicular to the optical axis as the X-axis, and the vertical direction perpendicular to the optical axis as the Y-axis, describes the coordinates of the target pixel point relative to the optical axis center position. The coordinates of the camera coordinate system ( , , ), in a coordinate system with the optical axis center of the camera as the origin, the optical axis direction as the Z-axis, the horizontal direction of the camera perpendicular to the optical axis as the X-axis, and the vertical direction of the camera perpendicular to the optical axis as the Y-axis, describes the coordinates of the target pixel point relative to the optical axis center position. When the depth camera is placed on a horizontal table with the lens facing up, the pitch angle is 0 at this time. When the depth camera lens faces the gravity direction, the pitch angle is 180°. When the depth camera lens is perpendicular to the gravity direction and parallel to the ground, the pitch angle is 90°. When the depth camera is placed horizontally and the optical axis is perpendicular to the ground, the pitch angle is 180° and the roll angle is 0° at this time, and the three-dimensional actual distance of the position represented by the pixel point relative to the camera center can be obtained.
[0077] Step 602, calculate the angle of the spatial point corresponding to the target pixel point relative to the center of the depth camera according to the following formula: ; where, represents the angle of the target pixel point relative to the center of the depth camera on the axis, represents the angle of the target pixel point relative to the center of the depth camera on the axis.
[0078] Step 603, convert the coordinates of the target pixel point in the camera coordinate system to the spatial three-dimensional coordinate system according to the following formula to obtain the spatial three-dimensional coordinates of the target pixel point : ; where, is the pitch angle of the depth camera.
[0079] In the embodiments of the present application, according to the actual engineering requirements, the pitch angle of the depth camera may not be perpendicular to the ground or parallel to the ground, and it is impossible to ensure that the camera position remains unchanged. It is necessary to use the IMU module in the depth camera to determine the current camera attitude. Determine that when the imaging center of the depth camera is the origin, the direction perpendicular to the ground is the Z-axis, the pitch angle is 90° and the roll angle is 0°, the horizontal direction of the camera parallel to the ground is the X-axis, and the optical axis position of the camera is the Y-axis to establish a world coordinate system. At this time, according to the pitch angle of the camera attitude obtained by the IMU module of the depth camera, the conversion from the coordinates of the target pixel point in the camera coordinate system to the world coordinates can be achieved through the calculation formula of the spatial three-dimensional coordinates .
[0080] In another exemplary embodiment of the present application, in step 305, the volume of each missing microelement is assigned according to the volume values of the first microelement and / or the second microelement adjacent to each missing microelement, which specifically includes: Assign the average volume value of the t microelements closest to each missing microelement to the volume value of the missing microelement, and the microelements include the first microelement and / or the second microelement.
[0081] In the embodiment of the present application, the average volume value of the t microelements closest to each missing microelement is calculated according to the following formula: ; where is the volume value of the th closest microelement.
[0082] In another exemplary embodiment of the present application, to verify the feasibility of applying the improved image segmentation algorithm of the present application to the calculation of muck volume, experiments are conducted on the superiority of the improved image segmentation algorithm and the feasibility of the solution in sequence.
[0083] To simulate the situation of a muck truck loading muck, the present application uses a box scaled in proportion to the size of a single carriage as the muck carrier and uses sand as the content to build an experimental platform. By taking the volume, shape, camera pitch angle, and camera height as variables, a dataset including RGB images, depth images (Depth images), depth information, and actual volume data as the content is captured. After screening and cleaning, a total of 1639 groups of data are retained, and 1500 of them are used as the training set and validation set datasets for training the image segmentation network based on bimodal input, and the labelme software is used to make segmentation labels for the muck in the pictures.
[0084] To verify the superiority of the improved YOLOv8, a comparative experiment is now conducted between the improved YOLO algorithm and the original model. This experiment is based on the windows10 system, and the programming platform is pycharm2023. The hardware used for training the model: the GPU is Nvidia RTX1080TI, the CPU is Inter's 10400F, the parameter batch_size is set to 32, and the number of training rounds is 100.
[0085] The adopted rapidity performance evaluation indicators are: the number of parameters Parameters / M, and the number of frames processed per second for a single image FPS (Frames Per Second). These two indicators can intuitively judge the size and rapidity of the model.
[0086] Evaluation metrics for the accuracy of muck image segmentation: Recall (recall, ) represents the proportion of correctly detected objects among all real objects, Precision (precision, ) represents the proportion of correctly detected objects among the detected objects, and Average Precision (AP) represents the performance of the network at different recall rates. Its calculation formula is: ; ; ; Among them, represents the number of correctly detected objects, represents the number of real objects not detected, represents the amount of non-target data wrongly detected.
[0087] The performance of the improved double-branch OursYOLOv8-seg and the original YOLOv8-seg in the collected muck image dataset is shown in Table 1 below.
[0088] Table 1 Index table before and after improvement
[0089] As can be seen from Table 1 above, the improved double-branch OursYOLOv8-seg of this application has improved in both speed and accuracy, and the number of parameters and computational volume have decreased compared to the original version. The accuracy of muck segmentation can reach 99.5%, and it can accurately separate muck from the background.
[0090] The dataset made during the experiment contains all the information required for calculation, including two types of modal images, depth data, camera angles, the height of the bottom of the muck box from the camera center, and the real volume of the muck. By comparing the volume calculated by the method of this application with the real volume through the above data, the accuracy of the method of this application in muck volume calculation can be obtained. The evaluation metric for evaluating volume accuracy is the average relative accuracy, which can be expressed by the following formula: ; Among them, is the number of experimental data. A total of 723 groups of data are used for accuracy verification. Divide the volume calculated from a group of data by the corresponding real volume , calculate its percentage, calculate the average relative error of 723 groups of data after obtaining the relative error, and the average relative accuracy can be obtained 。The experimental results in the self-built dataset show that the average relative volume accuracy calculated by the method of this application in the dataset can reach 95.5%, and the calculation time for a single set of data is about 0.4s, which can meet the actual application requirements.
[0091] Based on the same inventive concept, the embodiment of this application also provides a machine vision-based shield muck volume calculation device for implementing the above-mentioned machine vision-based shield muck volume calculation method. The solution provided by this device to solve the problem is similar to the solution described in the above method. Therefore, the specific limitations in one or more embodiments of the machine vision-based shield muck volume calculation device provided below can refer to the limitations on the machine vision-based shield muck volume calculation method in the above text, and will not be repeated here.
[0092] In an exemplary embodiment, as Figure 7 shown, a machine vision-based shield muck volume calculation device 70 is provided, including: An image segmentation module 701, configured to perform feature extraction and fusion on the real-time acquired target depth image and target RGB image through an image segmentation network based on bimodal input, and perform image segmentation on the target RGB image according to the obtained fusion features to obtain the muck mask image of the target RGB image and the coordinates of the muck bounding box; A volume calculation module 702, configured to calculate the volume of the shield muck through a volume calculation method according to the coordinates of the muck bounding box and the muck mask image.
[0093] In the embodiment of this application, the relevant introduction of the image segmentation network based on bimodal input can be found in the description of the above method embodiment, and will not be repeated here.
[0094] In another exemplary embodiment of this application, the above volume calculation module 702 is further configured to: In each target depth image, form a rectangular sliding window starting from the upper left vertex of the muck bounding box; Perform a micro-element classification operation: If the pixel depth value of any one of the four vertices of the rectangular sliding window is missing, the pixel area included in the rectangular sliding window is recorded as a missing micro-element, and the volume of the missing micro-element is recorded as 0; If none of the four vertices of the rectangular sliding window have missing pixel depth values, and at the same time: If all pixel points within the rectangular sliding window belong to muck pixel points, the pixel area included in the rectangular sliding window is recorded as a first micro-element; wherein, if the pixel value of any pixel point within the rectangular sliding window in the corresponding muck mask image is 255, then the pixel point within the rectangular sliding window belongs to muck pixel points; If the number of pixels belonging to muck pixels within the rectangular sliding window is less than a preset first threshold, then mark the pixel region included in the rectangular sliding window as an invalid micro-element, and record the volume of the invalid micro-element as nan; If the number of pixels belonging to muck pixels within the rectangular sliding window is not less than the preset first threshold and less than n, then mark the pixel region included in the rectangular sliding window as a second micro-element, where n is the total number of pixels in the rectangular sliding window; After performing the above micro-element classification operation, perform a micro-element volume calculation operation: Calculate the volume of the first micro-element according to the three-dimensional space coordinate system of the four vertices of the first micro-element; Calculate the volume of the second micro-element according to the four vertices of the second micro-element and the three-dimensional space coordinate system of the muck pixels; After performing the micro-element volume calculation operation, perform a loop operation: In each target depth image, move the rectangular sliding window in sequence in the moving direction from left to right and from top to bottom until the rectangular sliding window moves to the final position, and perform the micro-element classification operation and the micro-element volume calculation operation after each movement of the rectangular sliding window; the final position is the position when the coordinate of the lower right vertex of the rectangular sliding window exceeds the lower right vertex of the muck bounding box; Assign a value to the volume of each missing micro-element in each target depth image according to the volume of the first micro-element and / or the second micro-element of each missing micro-element neighbor in each target depth image.
[0095] In the embodiments of the present application, the size, moving step length, and first threshold of the rectangular sliding window are not specifically limited and can be set according to actual needs. For example, set the size of the rectangular sliding window to 3 3. At this time, the vertical and horizontal moving step lengths of the rectangular sliding window are both 3, n is 9, and the first threshold can be set to 5 at this time.
[0096] In another exemplary embodiment of the present application, the above volume calculation module 702 is further configured to: Convert the pixel coordinates of the four vertices of the first micro-element to a three-dimensional space coordinate system to obtain the three-dimensional space coordinates of the four vertices of the first micro-element; Calculate the volume of the first micro-element according to the three-dimensional space coordinates of the four vertices of the first micro-element and the height of the first micro-element according to the volume calculation method of a quadrangular prism; wherein, the average height of the four vertices of the first micro-element is the absolute value of the difference between the first height value and the second height value, the first height value is the average value of the spatial Z-axis coordinate values of the four vertices of the first micro-element, and the second height value is the height of the depth camera from the bottom of the muck truck carriage.
[0097] In another exemplary embodiment of the present application, the above volume calculation module 702 is further configured to: Convert the pixel coordinates of the four vertices of the second micro-element and the muck pixel points to a three-dimensional space coordinate system to obtain the three-dimensional space coordinates of the four vertices of the second micro-element and the muck pixel points; Calculate the volume of the second micro-element according to the volume calculation method of a quadrangular prism based on the three-dimensional space coordinates of the four vertices of the second micro-element and the height of the second micro-element; wherein, the height of the second micro-element is the absolute value of the difference between the third height value and the second height value, and the third height value is the average value of the spatial Z-axis coordinate values of all muck pixel points of the second micro-element.
[0098] In another exemplary embodiment of the present application, the above volume calculation module 702 is further configured to: Convert the pixel coordinates of the target pixel point to the camera coordinate system to obtain the coordinates of the target pixel point in the camera coordinate system. The conversion formula is: ; wherein, ( u , v ) represents the pixel coordinates of the target pixel point, represents the coordinates of the optical axis center of the depth camera in the pixel coordinate system, is the focal length of the depth camera on the x-axis (unit: pixel), is the focal length of the depth camera on the y-axis (unit: pixel), ( , , ) represents the coordinates of the target pixel point in the camera coordinate system, and the Z-axis coordinate of the target pixel point in the camera coordinate system is the depth value obtained by the depth camera; Calculate the angle of the spatial point corresponding to the target pixel point relative to the center of the depth camera according to the following formula: ; wherein, represents the angle of the target pixel point relative to the center of the depth camera on the axis, represents the angle of the target pixel point relative to the center of the depth camera on the axis; Convert the coordinates of the target pixel point in the camera coordinate system to the three-dimensional space coordinate system according to the following formula to obtain the three-dimensional space coordinates of the target pixel point : ; wherein, is the pitch angle of the depth camera.
[0099] In another exemplary embodiment of the present application, the above volume calculation module 702 is further configured to: The one closest to each missing micro-elementt The average volume value of a microelement is assigned to the volume value of the missing microelement, which includes the first microelement and / or the second microelement.
[0100] In an exemplary embodiment, a computer device is provided. The computer device can be a server or a terminal, and its internal structure diagram can be as Figure 8 shown. The computer device includes a processor, a memory, an input / output interface (Input / Output, abbreviated as I / O), and a communication interface. Among them, the processor, the memory, and the input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store shield muck volume calculation data based on machine vision. The input / output interface of the computer device is used to exchange information between the processor and external devices. The communication interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, it implements a method for calculating the volume of shield muck based on machine vision.
[0101] Those skilled in the art can understand that Figure 8 the structure shown in
[0102] is only a block diagram of some structures related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.
[0103] In an exemplary embodiment, a computer device is further provided, including a memory and a processor. A computer program is stored in the memory, and when the processor executes the computer program, the steps in the above method embodiments are implemented.
[0104] In an exemplary embodiment, a computer-readable storage medium is provided, storing a computer program, and when the computer program is executed by the processor, the steps in the above method embodiments are implemented.
[0105] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use, and processing of relevant data need to comply with relevant regulations.
[0106] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, database, or other medium used in the embodiments provided in this application can include at least one of non-volatile and volatile memories. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc.
[0107] The databases involved in the embodiments provided in this application can include at least one of relational databases and non-relational databases. Non-relational databases can include distributed databases based on blockchain, etc., without limitation. The processors involved in the embodiments provided in this application can be general-purpose processors, central processors, graphics processors, digital signal processors, programmable logic devices, data processing logics based on quantum computing, etc., without limitation.
[0108] The technical features of the above embodiments can be combined arbitrarily. For the sake of concise description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.
[0109] In this text, specific examples are used to elaborate on the principles and implementation manners of the present application. The description of the above embodiments is only used to help understand the method of the present application and its core idea. At the same time, for those of ordinary skill in the art, according to the idea of the present application, there will be changes in the specific implementation manners and application scopes. In summary, the content of this specification should not be construed as a limitation to the present application.
Claims
1. A method for calculating the volume of shield muck based on machine vision, characterized in that, The method for calculating the volume of shield muck based on machine vision includes: Through an image segmentation network based on dual-modal input, feature extraction and fusion are performed on the obtained real-time target depth image and target RGB image, and based on the obtained fusion features, image segmentation is performed on the target RGB image to obtain the muck mask image of the target RGB image and the coordinates of the muck bounding box; According to the coordinates of the muck bounding box and the muck mask image, the volume of the shield muck is calculated through a volume calculation method.
2. The method for calculating the volume of shield muck based on machine vision according to claim 1, wherein The image segmentation network based on dual-modal input includes: The input layer is used to merge channels for each of the input target RGB images and the corresponding target depth images, and convert the feature map obtained by channel merging into a matrix , where B is the batch dimension, H is the width of the feature map, W is the length of the feature map; Backbone network, used to perform dual-modal feature extraction and fusion on the matrix to obtain shallow features and deep features; wherein, the dual-modal features include depth image features and RGB image features; A neck network, which is used to perform cross-scale fusion operations on the shallow features and deep features output by the backbone network to obtain fusion features with the same scales as the shallow features and the deep features respectively; A segmentation layer, which is used to generate and output at least the coordinates of the muck bounding box of the target RGB image and the muck mask image according to the fusion features output by the neck network.
3. The method for calculating the volume of shield muck based on machine vision according to claim 2, wherein The backbone network includes: The first feature extraction module is used to separate the target depth image part of the matrix and perform feature extraction on the target depth image part to obtain depth image features; The second feature extraction module is used to separate the target RGB image part of the matrix and perform feature extraction on the target RGB image part to obtain RGB image features; A first dual-modal feature fusion module, which is used to perform feature fusion based on a triple attention mechanism on the dual-modal features output by the first feature extraction module and the second feature extraction module to obtain a first shallow feature; A third feature extraction module, which is used to perform feature extraction on the output of the first feature extraction module; A fourth feature extraction module, which is used to perform feature extraction on the output of the first dual-modal feature fusion module; A second dual-modal feature fusion module, which is used to perform feature fusion based on a triple attention mechanism on the dual-modal features output by the third feature extraction module and the fourth feature extraction module to obtain a second shallow feature; A fifth feature extraction module, which is used to perform feature extraction on the output of the third feature extraction module; A sixth feature extraction module, which is used to perform feature extraction on the output of the second dual-modal feature fusion module; A third dual-modal feature fusion module, which is used to perform feature fusion based on a triple attention mechanism on the dual-modal features output by the fifth feature extraction module and the sixth feature extraction module; An SPPF spatial pyramid pooling module, which is used to perform pooling operations of different sizes on the output of the third dual-modal feature fusion module and splice the feature maps obtained by the pooling operations of different sizes in the channel dimension to obtain the deep features.
4. The method for calculating the volume of shield muck based on machine vision according to claim 3, characterized in that, The structures of the first dual-modal feature fusion module, the second dual-modal feature fusion module, and the third dual-modal feature fusion module are the same, and each includes: A first attention weight extraction module, which is used to extract a first attention weight, a second attention weight, and a third attention weight, where the first attention weight refers to the correlation weight between the channel dimension and the width of the input RGB feature map, the second attention weight refers to the correlation weight between the channel dimension and the length of the input RGB feature map, and the third attention weight refers to the spatial attention weight of the input RGB feature map; The second attention weight extraction module is used to extract the fourth attention weight, the fifth attention weight, and the sixth attention weight. Among them, the fourth attention weight refers to the correlation weight between the channel dimension and the width of the input depth feature map, the fifth attention weight refers to the correlation weight between the channel dimension and the length of the input depth feature map, and the sixth attention weight refers to the spatial attention weight of the input depth feature map; The adder module is used to add the first attention weight and the fourth attention weight to obtain the first combined attention weight, add the second attention weight and the fifth attention weight to obtain the second combined attention weight, and add the third attention weight and the sixth attention weight to obtain the third combined attention weight; The first attention feature map extraction module is used to perform average weighted fusion on the RGB feature map by using the first combined attention weight, the second combined attention weight, and the third combined attention weight to obtain the RGB attention feature map; The second attention feature map extraction module is used to perform average weighted fusion on the depth feature map by using the first combined attention weight, the second combined attention weight, and the third combined attention weight to obtain the depth attention feature map; The seventh connection layer is used to splice the RGB attention feature map and the depth attention feature map.
5. The method for calculating the volume of shield muck based on machine vision according to claim 4, wherein The first attention weight extraction module includes a first branch, a second branch, and a third branch, and the second attention weight extraction module includes a fourth branch, a fifth branch, and a sixth branch, where: The structures of the first branch and the sixth branch are the same, and both include: The first dimension adjustment module is used to rearrange the order of each dimension of the input RGB feature map or depth feature map according to the following formula: ; Among them, I is the input RGB feature map or the depth feature map, and its shape is , where 0 represents I the batch size of B , 1 represents I the length of H , 2 represents I the number of channels of C , 3 represents I the width of W , I cw is the output of the first dimension adjustment module, and its shape is ; The first pooling module is used to perform a pooling operation on I cw ; The second convolutional layer is used to perform a convolutional operation on the output of the first pooling module; The structures of the second branch and the fifth branch are the same, and both include: The second dimension adjustment module is used to rearrange the order of each dimension of the input RGB feature map or depth feature map according to the following formula: ; Among them, I hc is the output of the second-dimensional adjustment module, and its shape is ; The second pooling module is used to I hc perform a pooling operation; The third convolutional layer is used to perform a convolutional operation on the output of the second pooling module; The structures of the third branch and the fourth branch are the same, and both include: The third pooling module is used to perform a pooling operation on the input RGB feature map or depth feature map; The fourth convolutional layer is used to perform a convolutional operation on the output of the third pooling module.
6. The method for calculating the volume of shield muck based on machine vision according to claim 1, characterized in that, Calculating the volume of the shield muck through a volume calculation method according to the coordinates of the muck bounding box and the muck mask image specifically includes: In each of the target depth images, a rectangular sliding window is formed starting from the upper left vertex of the muck bounding box; Performing a micro-element classification operation: If the pixel depth value of any one of the four vertices of the rectangular sliding window is missing, the pixel area included in the rectangular sliding window is recorded as a missing micro-element, and the volume of the missing micro-element is recorded as 0; If none of the four vertices of the rectangular sliding window have missing pixel depth values, and at the same time: If all the pixel points within the rectangular sliding window belong to muck pixel points, then the pixel region enclosed by the rectangular sliding window is denoted as the first microelement; wherein, if the pixel value of any pixel point within the rectangular sliding window in the muck mask image is 255, then this pixel point within the rectangular sliding window belongs to a muck pixel point; If the number of muck pixel points within the rectangular sliding window is less than a preset first threshold, then the pixel region enclosed by the rectangular sliding window is denoted as an invalid microelement, and the volume of the invalid microelement is denoted as nan; If the number of muck pixel points within the rectangular sliding window is not less than the preset first threshold but less than n, where n is the total number of pixel points in the rectangular sliding window, then the pixel region enclosed by the rectangular sliding window is denoted as the second microelement; After performing the microelement classification operation, perform the microelement volume calculation operation: Calculate the volume of the first microelement according to the three-dimensional space coordinate system of the four vertices of the first microelement; Calculate the volume of the second microelement according to the four vertices of the second microelement and the three-dimensional space coordinate system of the muck pixel points; After performing the microelement volume calculation operation, perform a loop operation: In each of the target depth images, move the rectangular sliding window in the moving direction from left to right and from top to bottom until the rectangular sliding window moves to the final position, and perform the microelement classification operation and the microelement volume calculation operation after each movement of the rectangular sliding window; the final position is the position where the coordinate of the lower right vertex of the rectangular sliding window exceeds the lower right vertex of the muck bounding box; After performing the loop operation, perform the volume assignment operation: Assign a volume to each missing microelement of each target depth image according to the volumes of the first microelement and / or the second microelement of each missing microelement neighbor of each target depth image; After performing the volume assignment operation, accumulate the volumes of all the missing microelements, the first microelements, and the second microelements of each target depth image to obtain the volume of the shield muck.
7. The method for calculating the volume of shield muck based on machine vision according to claim 6, characterized in that, The calculating the volume of the first microelement according to the three-dimensional space coordinate system of the four vertices of the first microelement specifically includes: Convert the pixel coordinates of the four vertices of the first microelement to the three-dimensional space coordinate system to obtain the spatial three-dimensional coordinates of the four vertices of the first microelement; Calculate the volume of the first microelement according to the spatial three-dimensional coordinates of the four vertices of the first microelement and the height of the first microelement according to the volume calculation method of a quadrangular prism; wherein, the average height of the four vertices of the first microelement is the absolute value of the difference between a first height value and a second height value, the first height value is the average value of the spatial Z-axis coordinate values of the four vertices of the first microelement, and the second height value is the height of the depth camera from the bottom of the muck truck carriage; The calculating the volume of the second microelement according to the four vertices of the second microelement and the three-dimensional space coordinate system of the muck pixel points specifically includes: Convert the pixel coordinates of the four vertices of the second microelement and the muck pixel points to the three-dimensional space coordinate system to obtain the spatial three-dimensional coordinates of the four vertices of the second microelement and the muck pixel points; Calculate the volume of the second micro-element according to the three-dimensional spatial coordinates of the four vertices of the second micro-element and the height of the second micro-element; wherein, the height of the second micro-element is the absolute value of the difference between the third height value and the second height value, and the third height value is the average value of the spatial Z-axis coordinate values of all the muck pixel points of the second micro-element.
8. The method for calculating the volume of shield muck based on machine vision according to claim 7, characterized in that Convert the pixel coordinates of the target pixel point to the three-dimensional space coordinate system according to the following steps, where the target pixel point includes any vertex of the first micro-element, any vertex of the second micro-element, or any muck pixel point of the second micro-element: Convert the pixel coordinates of the target pixel point to the camera coordinate system to obtain the coordinates of the target pixel point in the camera coordinate system, and the conversion formula is: ; Among them, ( u , v ) represents the pixel coordinates of the target pixel point, represents the coordinates of the optical axis center of the depth camera in the pixel coordinate system, is the focal length of the depth camera on the x-axis, is the focal length of the depth camera on the y-axis, ( , , ) represents the coordinates of the target pixel point in the camera coordinate system, and the Z-axis coordinate of the target pixel point in the camera coordinate system is the depth value obtained by the depth camera; Calculate the angle of the spatial point corresponding to the target pixel point relative to the center of the depth camera according to the following formula: ; Among them, represents the angle of the target pixel point relative to the center of the depth camera in axis, represents the angle of the target pixel point relative to the center of the depth camera in axis; Convert the coordinates of the target pixel point in the camera coordinate system to the three-dimensional space coordinate system according to the following formula to obtain the three-dimensional space coordinates of the target pixel point : ; Wherein, is the pitch angle of the depth camera.
9. The method for calculating the volume of shield muck based on machine vision according to claim 6, wherein, Assign values to the volumes of each missing micro-element according to the volume values of the first micro-element and / or the second micro-element adjacent to each missing micro-element, specifically including: Assign the average volume value of the t micro-elements that are the nearest neighbors to each of the missing micro-elements to the volume value of the missing micro-elements, where the micro-elements include first micro-elements and / or second micro-elements.
10. A computer device, comprising: A memory, a processor, and a computer program stored on the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the machine vision-based shield muck volume calculation method according to any one of claims 1-9.
Citation Information
Patent Citations
Computer-implemented method of analyzing an image to segment article of interest therein
CA3140924A1
Real-time earth volume calculation method based on binocular vision
CN112819882A
Image segmentation method, electronic equipment and computer readable storage medium
CN112949641A
Intelligent reinforcement detection method and system based on convolutional neural network and binocular vision
CN116703835A
Intelligent dynamic identification method for deslagging volume of earth pressure balance shield
CN118351526A
Cited By
Shield slag hole real-time analysis method based on image stabilization enhancement and interactive detection driving
CN121982620A
Real-time analysis method for shield slagging port based on image stabilization enhancement and interaction detection driving
CN121982620B
Complex stratum shield tunnel face image intelligent identification and analysis method and system
CN122135045A