Coarse aggregate grading detection method and device based on multi-modal fusion target detection

By fusing RGB images and depth images, using a multimodal fusion target detection model to detect the aggregate particle size interval, and combining depth information for volume calculation, the particle size calculation error and complexity problems in the existing aggregate grading detection methods are solved, and a higher precision aggregate grading detection is achieved.

CN120047447AActive Publication Date: 2025-05-27FUJIAN SOUTHERN HIGHWAY MECHANICAL CO LTD +1
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202510528570.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-25
Publication Date
2025-05-27
Estimated Expiration
2045-04-25

AI Technical Summary

Technical Problem

The existing aggregate grading detection methods have problems such as particle size calculation error and complex and cumbersome calculation process, which affect the detection accuracy.

Method used

The method based on multimodal fusion target detection is adopted to fuse the RGB image and the depth image, and the multimodal fusion target detection model is used to automatically detect the particle size interval of the aggregate, and the aggregate volume calculation is performed based on image depth information.

Benefits of technology

The accuracy and robustness of aggregate grading detection are improved, and the calculated aggregate volume is more accurate, supporting subsequent grading calculations, and being able to determine whether the aggregate grading is qualified.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120047447A_ABST
    Figure CN120047447A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of concrete production, in particular to a coarse aggregate grading detection method and device based on multi-modal fusion target detection. The invention discloses a coarse aggregate grading detection method based on multi-modal fusion target detection. The method comprises the following steps: S1, aggregate pretreatment; s2, establishing a multi-modal fusion target detection network; s3, carrying out volume calculation; and S4, production line application. According to the method, the depth image and the RGB image are fused, the particle size interval of the aggregate is automatically detected by using the multi-modal fusion target detection model, and the fusion of the RGB image and the depth image can enrich the features of the image and improve the detection precision and robustness of the model because the single-modal aggregate image contains few features and lacks important height information; according to the aggregate volume calculation method in combination with the image depth information, due to the height information of the added aggregate, the calculated aggregate volume is higher in accuracy, subsequent grading calculation is facilitated, and whether the aggregate grading in the batch is qualified or not can be correspondingly judged.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of concrete production, and in particular, to a detection method and device for coarse aggregate gradation based on multi-modal fusion object detection. Background Art

[0002] For the current intelligent detection methods of aggregate gradation, most of them adopt the image method. For example, a camera is used to collect RGB images, the collected images are segmented, the contours of the aggregates are extracted, the particle sizes of the aggregates are fitted, the particle grades of the aggregates are selected, and the gradation is calculated by calculating the proportion of the volume of the coarse aggregates falling into each particle grade in the total volume of the coarse aggregates. Since the coarse aggregate particles lack height information, an equivalent ellipsoid is used to replace the volume, and the equivalent particle size obtained by characterization is used to replace the height of the equivalent ellipse. However, there are defects in this aggregate gradation detection method. Its particle size is approximately obtained by the circumscribed rectangle, and the volume is calculated by the equivalent ellipse. The data obtained equivalently all have certain errors, which affect the accuracy of the calculated gradation; the process of calculating the particle grade requires a series of algorithms such as segmentation and equivalence, and the calculation process is complex and cumbersome.

[0003] In this context, the present invention proposes a method for detecting aggregate gradation based on a deep learning object detection algorithm that combines RGB images and depth images. Through the innovative object detection algorithm, the particle grade distribution of the aggregates can be efficiently detected. By adding height information to the original volume calculation method, a more accurate volume of the aggregates can be obtained, and the detection accuracy of the gradation can be improved. Summary of the Invention

[0004] Other features and advantages of the present invention will be described in the following specification, and will be partially obvious from the specification, or will be understood by implementing the present invention. The objectives and other advantages of the present invention can be achieved and obtained through the structures specifically pointed out in the specification and other specification drawings.

[0005] The objective of the present invention is to overcome the above deficiencies, and provide a detection method and device for coarse aggregate gradation based on multi-modal fusion object detection. The depth image and the RGB image are fused, and the multi-modal fusion object detection model is used to automatically detect the particle size range of the aggregates. Since the features contained in a single-modal aggregate image are less and lack important height information, the fusion of the RGB image and the depth image can enrich the features of the image, improve the detection accuracy and robustness of the model, and combine the aggregate volume calculation method with the image depth information. Since the height information of the aggregates is added, the accuracy of the calculated aggregate volume is higher, which is beneficial to the subsequent gradation calculation. Correspondingly, it can be determined whether the aggregate gradation in this batch is qualified.

[0006] The present invention provides a detection method for coarse aggregate gradation based on multi-modal fusion object detection, including: S1. Aggregate pretreatment: Use a permeable sieve to screen out aggregates of different particle sizes. Collect RGB images and depth images for each particle size of aggregates for sampling. After establishing a dataset, perform data augmentation, and input the augmented dataset into the designed multi-modal fusion object detection network; S2. Establish a multi-modal fusion object detection network: The multi-modal fusion object detection network includes an image feature extraction module, three Transformer fusion modules with pyramid structures, and an interpretation module. Input the datasets of RGB images and depth images into the image feature extraction module to extract features from the data. Input the two modal feature maps obtained into the three Transformer fusion modules with pyramid structures. Each Transformer fusion module will obtain a spliced hierarchical feature image. Decode the three spliced hierarchical feature maps and input them into the yolo detection head, and the yolo detection head can output the particle size information of the corresponding level finally; S3. Volume calculation: Collect RGB images and depth images of the aggregate raw materials passing through the production line. Process the depth image for contour, detect the particle size information of each aggregate within the current view, and calculate their respective volumes; S4. Application in the production line: According to the volume, judge the particle size information of each aggregate in the image, classify them according to the particle size range, calculate the sieve residue value of each particle size range, draw a particle shape gradation curve, and judge whether this batch of aggregates is qualified.

[0007] In some embodiments, in step S1, the specific operation of data augmentation is to use the copypast algorithm to simultaneously cut out the aggregates from the RGB image and the depth image and perform the same transformation and deformation, and splice the aggregates in different particle size ranges onto a new background image to obtain depth images and RGB images containing aggregates of different particle sizes as the dataset for the multi-modal fusion object detection network in step S2.

[0008] In some embodiments, in step S2, the image feature extraction module specifically includes a channel attention module, a convolution module, and a neck module. After the datasets of RGB images and depth images are input into the image feature extraction module, they first enter the channel attention module to split and splice the pictures along the channel direction to achieve the purpose of downsampling and improving the representation ability of the model. Then, the data is input into the convolution module for feature extraction. Finally, it is input into the neck module to reduce the calculation amount of the high-dimensional feature map and improve the representation ability of the model. The RGB image obtains an RGB image feature map after passing through the image feature extraction module, and the depth image obtains a depth image feature map after passing through the image feature extraction module.

[0009] In some embodiments, the three pyramid-structured Transformer fusion modules specifically include the Transformer1 fusion module, the Transformer2 fusion module, and the Transformer3 fusion module. Before inputting into the Transformer fusion module, an average pooling operation needs to be performed on the two modal feature maps obtained in step S2. After the average pooling operation, the feature maps of the two modalities are serialized and concatenated along the channel direction to obtain a long sequence.

[0010] In some embodiments, the long sequence is input into the Transformer1 fusion module. After being processed by the multi-head attention and self-attention in the module, the feature map jointly concerned by the two modalities is obtained after decoding. The jointly concerned feature map is added to the RGB image feature map and the depth image feature map respectively to obtain two feature images feature1_rgb and feature1_dep that have both common features and individual modal features.

[0011] In some embodiments, for the feature images feature1_rgb and feature1_dep, convolution and neck module processing are respectively performed to extract high-level features. The processed results are input into the Transformer2 fusion module for fusion again to obtain a relatively high-level fusion feature map. The obtained relatively high-level fusion feature map is added to the feature images feature1_rgb and feature1_dep respectively to obtain relatively high-level feature images feature2_rgb and feature2_dep. Similarly, after the fusion addition steps for the relatively high-level feature images feature2_rgb and feature2_dep, relatively high-level feature images feature3_rgb and feature3_dep are obtained.

[0012] In some embodiments, in step S3, the collected depth image is binarized. The depth image is essentially a two-dimensional matrix, and the numerical points in the matrix are the distances from the depth camera to the shooting point. First, measure the distance from the depth camera to the rolling belt plane, set this value as a threshold, and determine whether the numerical points in the depth image are greater than this threshold. If greater, the value of this point is set to 255; if less, the value of this point is set to 0 to form a binary image.

[0013] In some embodiments, contour processing is performed on the binary image, that is, the number of pixel points with a value of 255 in the binary image is counted, and contour recognition processing is performed on the obtained binary image. Based on the contour map, the contour areas of each aggregate are calculated, and the contour area of the corresponding aggregate is divided by the number of pixel points to obtain the area contribution value of each pixel point.

[0014] In some embodiments, each pixel point is fitted into a small rectangle. The volume of the rectangle is the base area multiplied by the height. The area contribution value of the pixel point is equivalent to the base area of each pixel point, and the volume contribution value of each pixel point is obtained by multiplying it by the corresponding height information. Finally, volume summation is performed. Since the depth image can only collect the height information of the aggregate exposed to the camera, the summation operation only obtains half of the volume of the aggregate. The aggregate is fitted into a symmetric figure along the projection plane, and the summation volume is multiplied by 2 to obtain the complete volume of the aggregate.

[0015] The coarse aggregate gradation detection device based on multi-modal fusion object detection includes: A production line on which a rolling belt is provided for transporting aggregates; An RGB linear array camera for collecting RGB images of aggregates; A depth camera for collecting depth images of aggregates; A light source provided above the production line to provide illumination light for the aggregates; An industrial control computer electrically connected to the RGB linear array camera and the depth camera to detect the coarse aggregate gradation.

[0016] By adopting the above technical solutions, the beneficial effects of the present invention are as follows: The present invention fuses the depth image and the RGB image, and uses the multi-modal fusion object detection model to automatically detect the particle size range of the aggregates. Since the single-modal aggregate image contains fewer features and lacks important height information, the fusion of the RGB image and the depth image can enrich the image features, improve the detection accuracy and robustness of the model. The aggregate volume calculation method combined with the image depth information has higher accuracy in calculating the aggregate volume due to the addition of the height information of the aggregates, which is beneficial to the subsequent gradation calculation. Correspondingly, it can be judged whether the aggregate gradation in this batch is qualified.

[0017] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present disclosure.

[0018] Undoubtedly, such objects of the present invention and other objects will become more apparent after the details of the preferred embodiments described in multiple drawings and figures below.

[0019] To make the above and other objects, features and advantages of the present invention more obvious and understandable, one or several preferred embodiments are specifically exemplified below and described in detail in conjunction with the accompanying drawings as follows. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] The drawings are used to provide a further understanding of the present invention, and constitute a part of the specification. They are used together with the embodiments of the present invention to explain the present invention, and do not constitute a limitation to the present invention.

[0021] In the accompanying drawings, like parts are designated by like reference numerals, and the drawings are schematic and not necessarily drawn to scale.

[0022] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the accompanying drawings required for the description of the embodiments or the prior art will be briefly introduced below. Obviously, the accompanying drawings in the following description are only one or several embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on such drawings.

[0023] Figure 1 Schematic diagram of an RGB image and depth image acquisition device in some embodiments of the present invention; Figure 2 Schematic diagram of a multi-modal fusion network structure in some embodiments of the present invention; Figure 3 Schematic diagram of a volume calculation process in some embodiments of the present invention; Figure 4 Schematic diagram of the working process of a grading calculation system in some embodiments of the present invention.

[0024] Main reference numeral description: 1. Production line; 2. RGB line array camera; 3. Depth camera; 4. Light source; 5. Industrial control computer. Detailed implementation manners

[0025] In order to make the objectives, technical solutions and advantages of the present invention more clearly understood, the present invention will be further described in detail below in conjunction with specific implementation manners. It should be understood that the specific implementation manners described herein are only used to explain the present invention, but not to limit the present invention.

[0026] In addition, in the description of the present invention, it should be understood that the orientation or positional relationships indicated by the terms "center", "upper", "lower", "front", "rear", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer", "axial", "radial", "circumferential", etc. are based on the orientation or positional relationships shown in the accompanying drawings, and are only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and thus should not be construed as a limitation of the present invention.

[0027] In the present invention, unless otherwise clearly specified or defined, terms such as "installed", "connected", "linked", "fixed", etc. shall be construed in a broad sense. For example, it may be a fixed connection, a detachable connection, or integrated; it may be directly connected, or indirectly connected through an intermediate medium, and may be the internal communication of two components or the interaction relationship between two components. However, indicating a direct connection means that there is no connection relationship constructed through a transition structure between the two connected main bodies, and they are only connected through the connection structure to form a whole. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.

[0028] In the present invention, unless otherwise clearly specified or defined, a first feature being "on" or "under" a second feature may be that the first and second features are in direct contact, or the first and second features are in indirect contact through an intermediate medium. In the description of this specification, the description referring to terms such as "one embodiment", "some embodiments", "example", "specific example", or "some examples" means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic expressions of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples.

[0029] Refer to Figures 1-4 , Figure 1 is a schematic diagram of an RGB image and depth image acquisition device in some embodiments of the present invention; Figure 2 is a schematic diagram of a multimodal fusion network structure in some embodiments of the present invention; Figure 3 is a schematic diagram of a volume calculation process in some embodiments of the present invention; Figure 4 is a schematic diagram of the working process of a grading calculation system in some embodiments of the present invention.

[0030] According to some embodiments of the present invention, the present invention provides a coarse aggregate grading detection method based on multimodal fusion object detection, including: S1. Aggregate pretreatment: Use a permeable sieve to screen out aggregates of different particle sizes, sample RGB images and depth images for each particle size of aggregates, perform data augmentation after establishing a data set, and input the augmented data set into a designed multimodal fusion object detection network; Since there is only a single particle size grade of aggregate in the RGB image and depth image at this time, in order to enrich the dataset and avoid overfitting of the model, it is also necessary to perform data augmentation on the standard dataset. The specific operation for data augmentation is to use the copypast algorithm to simultaneously extract the aggregate from the RGB image and depth image, perform the same transformation and deformation, and splice the aggregates in different particle size ranges onto a new background image to obtain depth images and RGB images containing aggregates of different particle size grades as the dataset for the multi-modal fusion object detection network in step S2. The dataset includes a training set and a validation set; Input the training set obtained by data augmentation into the multi-modal fusion object detection network for training. After 1000 rounds of training, the corresponding model is obtained. Then, use the validation set to verify the accuracy and stability of the model. If the performance of the model is good, it can be input into the distribution images of aggregates of different particle size grades in normal working conditions for detection, and the particle size distribution range of each aggregate can be detected.

[0031] S2. Establish a multi-modal fusion object detection network: The multi-modal fusion object detection network includes an image feature extraction module, three Transformer fusion modules with pyramid structures, and an interpretation module. Input the datasets of the RGB image and depth image into the image feature extraction module to extract features from the data. Input the two modal feature maps obtained into the three Transformer fusion modules with pyramid structures. Each Transformer fusion module will obtain a spliced hierarchical feature image. Decode the three spliced hierarchical feature maps and input them into the yolo detection head, and the yolo detection head can output the particle size grade information corresponding to the final level; In the image preprocessing step, use the semi-automatic annotation algorithm to pre-label the RGB image with the particle size range label in advance. Then, the established dataset also has the particle size label. During the process of inputting into the multi-modal fusion object detection network, the RGB image also carries the particle size information. Therefore, the particle size grade information corresponding to the final level can be visually observed when the yolo detection head outputs the particle size grade information corresponding to the final level; The image feature extraction module specifically includes a channel attention module, a convolution module, and a neck module. After the datasets of the RGB image and depth image are input into the image feature extraction module, they first enter the channel attention module to split and splice the pictures along the channel direction to achieve the purpose of downsampling and improving the representation ability of the model. Then, input the data into the convolution module for feature extraction. Finally, input it into the neck module to reduce the computational amount of the high-dimensional feature map and improve the representation ability of the model. The RGB image obtains the RGB image feature map after passing through the image feature extraction module, and the depth image obtains the depth image feature map after passing through the image feature extraction module; The three Transformer fusion modules with pyramid structures specifically include the Transformer1 fusion module, the Transformer2 fusion module, and the Transformer3 fusion module. Since the two-modal images need to be stitched and fused subsequently, the computational complexity increases exponentially. Therefore, before inputting into the Transformer fusion module, average pooling operations need to be performed on the two modal feature maps obtained in step S2 to reduce the computational complexity. After the average pooling operation, the two modal feature maps are serialized and stitched along the channel direction to obtain a long sequence; As Figure 2 shown, the long sequence is input into the Transformer1 fusion module. After being processed by the multi-head attention and self-attention in the module, the feature map jointly concerned by the two modalities is obtained after decoding. The jointly concerned feature map is added to the RGB image feature map and the depth image feature map respectively to obtain two feature images feature1_rgb and feature1_dep that have both common features and individual modal features; For the feature images feature1_rgb and feature1_dep, convolution and neck module processing are performed respectively to increase the receptive field and extract high-level features. The processed results are input into the Transformer2 fusion module for fusion again to obtain a higher-level fusion feature map. The obtained higher-level fusion feature map is added to the feature images feature1_rgb and feature1_dep respectively to obtain higher-level feature images feature2_rgb and feature2_dep. Similarly, after performing the fusion addition steps on the higher-level feature images feature2_rgb and feature2_dep, even higher-level feature images feature3_rgb and feature3_dep are obtained; Among them, in the fusion addition step for the higher-level feature images feature2_rgb and feature2_dep, an SPP module is additionally set between the convolution and neck processing modules. After performing downsampling on the data after convolution, through pooling layers with pooling kernel sizes of 5, 9, and 13, different feature layers of the input feature map are captured and fused together, and then neck module processing is performed to reduce the computational complexity; After passing through three pyramid-structured Transformer fusion modules, image features with different receptive fields are obtained. Then, a concatenation operation is performed to concatenate different modality feature maps with the same receptive field, that is, feature1_rgb and feature1_dep are concatenated, and so on to obtain three concatenated hierarchical feature maps. These three concatenated hierarchical feature maps are input into the YOLO detection head, and finally, the object detection results of the RGB image, that is, the particle size information of each aggregate, are output.

[0032] S3. Volume calculation: Collect RGB images and depth images of the aggregate raw materials passing through production line 1. Process the depth image for contour, detect the particle size information of each aggregate within the current view, and calculate their respective volumes. As Figure 3 shown, perform binary processing on the collected depth image. The depth image is essentially a two-dimensional matrix, and the numerical points in the matrix are the distances from the depth camera 3 to the shooting point. First, measure the distance from the depth camera 3 to the rolling belt plane, set this value as a threshold, and judge whether the numerical points in the depth image are greater than this threshold. If greater, set the value of this point to 255; if less, set the value of this point to 0 to form a binary image. Perform contour processing on the binary image, that is, count the number of pixel points with a value of 255 in the binary image, and perform contour recognition processing on the obtained binary image. Based on the contour map, calculate the contour area of each aggregate, divide the contour area of the corresponding aggregate by the number of pixel points to obtain the area contribution value of each pixel point. Fit each pixel point into a small rectangle. The volume of the rectangle is the base area multiplied by the height. The area contribution value of the pixel point is equivalent to the base area of each pixel point. Multiply it by the corresponding height information to obtain the volume contribution value of each pixel point, and finally perform volume summation. Since the depth image can only collect the height information of the part of the aggregate exposed to the camera, the summation operation only obtains half of the volume of the aggregate. Fit the aggregate into a symmetric figure along the projection plane, and multiply the summation volume by 2 to obtain the complete volume of the aggregate.

[0033] S4. Production line application: As Figure 4 shown, based on the volume, judge the particle size information of each aggregate in the image, perform classification processing according to the particle size range, calculate the sieve residue value in each particle size range, draw the particle shape grading curve, and judge whether this batch of aggregates is qualified. The specific calculation method of the sieve residue value is as follows: First, calculate the sum of the volumes of the aggregates in each particle size range, compare it with the sum of the volumes of all the aggregates, calculate the cumulative sieve residue value and the cumulative sieve residue value of the aggregates in each particle size range, and draw a curve based on the cumulative sieve residue value of the aggregates in each particle size range obtained, so as to judge whether this batch of aggregates is qualified. Among them, the matplotlib library is used to draw the curve, with the particle size range of the aggregates as the horizontal axis and the cumulative sieve residue value as the vertical axis; the actual grading curve is drawn, and the actual grading curve is compared with the required grading curve to judge whether this batch of aggregates is qualified.

[0034] As Figure 1 shown, the present invention also provides a coarse aggregate grading detection device based on multi-modal fusion object detection, including: A production line 1, on which a rolling belt is provided for transporting aggregates; An RGB linear array camera 2, which is used to collect RGB images of aggregates; A depth camera 3, which is used to collect depth images of aggregates; A light source 4, which is arranged above the production line 1 to provide illumination light for the aggregates; An industrial control computer 5, which is electrically connected to the RGB linear array camera 2 and the depth camera 3 to detect the grading of coarse aggregates.

[0035] It should be understood that the embodiments disclosed in the present invention are not limited to the specific processing steps or materials disclosed herein, but should extend to equivalent alternatives of such features understood by those of ordinary skill in the relevant art. It should also be understood that the terms used herein are only for the purpose of describing specific embodiments and do not mean to limit.

[0036] The "embodiments" mentioned in the specification mean that the specific features or characteristics described in connection with the embodiments are included in at least one embodiment of the present invention. Therefore, the phrase "embodiments" that appears throughout the specification does not necessarily refer to the same embodiment.

[0037] In addition, the described features or characteristics can be combined into one or more embodiments in any other suitable way. In the above description, some specific details, such as thickness, quantity, etc., are provided to provide a comprehensive understanding of the embodiments of the present invention. However, those skilled in the relevant art will understand that the present invention can be implemented without one or more of the above specific details or can also be implemented using other methods, components, materials, etc.

Claims

1. A coarse aggregate gradation detection method based on multimodal fusion target detection, characterized in that: include: S1. Aggregate preprocessing: Use a permeable screen to screen out aggregates of different particle sizes, collect RGB images and depth images for each particle size of aggregate, perform data enhancement after establishing the data set, and input the enhanced data set into the designed multimodal fusion target detection network; S2. Establish a multimodal fusion target detection network: The multimodal fusion target detection network includes an image feature extraction module, a three-pyramid structure Transformer fusion module, and an interpretation module. The data sets of RGB images and depth images are input into the image feature extraction module, and features are extracted from the data. The obtained two-modal feature maps are input into the three-pyramid structure Transformer fusion modules. Each Transformer fusion module will obtain a spliced ​​level feature image. The three spliced ​​level feature maps are interpreted and input into the yolo detection head, which can output the granularity information of the corresponding level. S3, volume calculation: collect RGB images and depth images of the aggregate raw materials passing through the production line, perform contour processing on the depth image, detect the particle size information of each aggregate within the current viewing angle, and calculate their respective volumes; S4. Production line application: According to the volume, determine the particle size information of each aggregate in the image, classify it according to the particle size range, calculate the sieve residue value of each particle size range, draw the particle shape grading curve, and determine whether this batch of aggregates is qualified.

2. The coarse aggregate gradation detection method based on multimodal fusion target detection according to claim 1 is characterized in that: In step S1, the specific operation of data enhancement is to use the copypast algorithm to simultaneously deduct aggregates from the RGB image and the depth image and perform the same transformation deformation, and to splice aggregates of different particle size ranges into the new background image, so as to obtain depth images and RGB images containing aggregates of different particle sizes as the data set of the multimodal fusion target detection network in step S2.

3. The coarse aggregate gradation detection method based on multimodal fusion target detection according to claim 2 is characterized in that: In step S2, the image feature extraction module specifically includes a channel attention module, a convolution module and a neck module. After the data sets of RGB images and depth images are input into the image feature extraction module, they first enter the channel attention module to split and splice the images along the channel direction to achieve the purpose of reducing the sampling and improving the representation ability of the model. The data is then input into the convolution module for feature extraction and finally input into the neck module to reduce the amount of high-dimensional feature map calculation and improve the representation ability of the model. The RGB image obtains the RGB image feature map after passing through the image feature extraction module, and the depth image obtains the depth image feature map after passing through the image feature extraction module.

4. The coarse aggregate gradation detection method based on multimodal fusion target detection according to claim 3 is characterized in that: The three pyramid-structured Transformer fusion modules specifically include Transformer1 fusion module, Transformer2 fusion module and Transformer3 fusion module. Before entering the Transformer fusion module, the two modal feature maps obtained in step S2 need to be average pooled. After the average pooling operation, the feature maps of the two modalities are serialized and spliced ​​along the channel direction to obtain a long sequence.

5. The coarse aggregate gradation detection method based on multimodal fusion target detection according to claim 4 is characterized in that: The long sequence is input into the Transformer1 fusion module. After multi-head attention and self-attention processing in the module, the feature map of common attention of the two modalities is obtained after decoding. The feature map of common attention is added to the RGB image feature map and the depth image feature map respectively to obtain two feature images feature1_rgb and feature1_dep with both common features and separate modality features.

6. The coarse aggregate gradation detection method based on multimodal fusion target detection according to claim 5 is characterized in that: For the feature images feature1_rgb and feature1_dep, convolution and neck module processing are performed respectively to extract high-level features. The processed results are input into the Transformer2 fusion module and fused again to obtain a higher-level fused feature map. The obtained higher-level fused feature map is added to the feature images feature1_rgb and feature1_dep respectively to obtain higher-level feature images feature2_rgb and feature2_dep. Similarly, after the fusion and addition steps of the higher-level feature images feature2_rgb and feature2_dep, higher-level feature images feature3_rgb and feature3_dep are obtained.

7. The coarse aggregate gradation detection method based on multimodal fusion target detection according to claim 1 is characterized in that: In step S3, the collected depth image is binarized. The depth image is essentially a one-dimensional matrix. The numerical point in the matrix is ​​the distance between the depth camera and the shooting point. First, the distance from the depth camera to the rolling belt plane is measured, and the value is set as a threshold to determine whether the numerical point in the depth image is greater than the threshold. If greater than, the value of the point is set to 255, if less than, the value of the point is set to 0, forming a binary image.

8. The coarse aggregate gradation detection method based on multimodal fusion target detection according to claim 7 is characterized in that: The binary image is processed with contours, that is, the number of pixels with a value of 255 in the binary image is counted, and the obtained binary image is processed with contour recognition. Based on the contour image, the contour area of ​​each aggregate is calculated, and the contour area of ​​the corresponding aggregate is divided by the number of pixels to obtain the area contribution value of each pixel.

9. The coarse aggregate gradation detection method based on multimodal fusion target detection according to claim 8, characterized in that: Each pixel is fitted into a small rectangle. The volume of the rectangle is the product of the bottom area and the height. The area contribution of the pixel is equivalent to the bottom area of ​​each pixel. Multiplying it by the corresponding height information gives the volume contribution of each pixel, and finally the volume is summed up. Since the depth image can only collect the height information of the part of the aggregate exposed to the camera, the summation operation only obtains half of the volume of the aggregate. The aggregate is fit into a symmetrical figure along the projection plane. The summed volume is multiplied by 2 to obtain the complete volume of the aggregate.

10. A coarse aggregate gradation detection device based on multimodal fusion target detection, characterized in that: The coarse aggregate gradation detection method based on multimodal fusion target detection according to any one of claims 1 to 9 is applied, and the device comprises: A production line, on which a rolling belt is arranged for conveying aggregate; An RGB line array camera, which is used to collect RGB images of aggregates; A depth camera, which is used to collect depth images of aggregates; A light source is disposed above the production line to provide illumination light for the aggregate; The industrial computer is electrically connected to the RGB linear array camera and the depth camera to detect the gradation of coarse aggregate.

Citation Information

Patent Citations

  • Blast furnace sintered ore particle size detection method and system based on RGB and laser feature fusion

    CN113870341A

  • Grain composition rapid detection method based on YOLO-V4

    CN114022474A

  • Rapid coal quantity detection method based on depth information fusion

    CN115953449A

  • Feature-guided multi-modal fusion RGB-D saliency target detection based on coordinate attention filtering

    CN116246058A

  • Aggregate volume calculation method and device based on visual image detection

    CN119180855A