Coarse Aggregate Gradation Detection Method and Device Based on Multimodal Fusion Object Detection

Through the multimodal object detection model that integrates depth images and RGB images, the problems of large particle size calculation errors and complex calculations in aggregate grading detection are solved, and higher precision grading detection is achieved.

CN120047447BActive Publication Date: 2025-07-18FUJIAN SOUTHERN HIGHWAY MECHANICAL CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510528570.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-25
Publication Date
2025-07-18
Estimated Expiration
2045-04-25

AI Technical Summary

Technical Problem

In the prior art, the aggregate grading detection method has the problem of large particle size calculation errors, complex and cumbersome calculation process, and the lack of height information, resulting in low grading detection accuracy.

Method used

The method based on multimodal fusion target detection is adopted to fuse the depth image and RGB image, and the multimodal fusion target detection model is used to automatically detect the particle size interval of the aggregate, and the volume calculation accuracy is improved by adding height information, and grading calculation is performed in combination with image depth information.

Benefits of technology

It improves the accuracy and robustness of aggregate grading detection, can accurately determine whether the batch of aggregate grading is qualified, and simplifies the calculation process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120047447B_ABST
    Figure CN120047447B_ABST
Patent Text Reader

Abstract

The present invention relates to the technical field of concrete production, and particularly to a detection method and device for coarse aggregate gradation based on multi-modal fusion object detection. The detection method for coarse aggregate gradation based on multi-modal fusion object detection includes S1, aggregate pretreatment; S2, establishing a multi-modal fusion object detection network; S3, volume calculation; S4, production line application. The present invention fuses depth images and RGB images, and uses a multi-modal fusion object detection model to automatically detect the particle size range of aggregates. Since the features contained in a single-modal aggregate image are less and lack important height information, the fusion of RGB images and depth images can enrich the features of the images, improve the detection accuracy and robustness of the model. The aggregate volume calculation method combined with the image depth information, due to the addition of the height information of the aggregates, makes the calculated aggregate volume more accurate, which is beneficial to subsequent gradation calculation, and accordingly can determine whether the aggregate gradation in this batch is qualified.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of concrete production, and particularly to a coarse aggregate gradation detection method and device based on multi-modal fusion object detection. Background Art

[0002] For the current intelligent detection methods of aggregate gradation, most of them adopt the image method. For example, a camera is used to collect RGB images, the collected images are segmented, the contours of the aggregates are extracted, the particle sizes of the aggregates are fitted, the particle grades of the aggregates are selected, and the gradation is calculated by calculating the proportion of the volume of the coarse aggregates in each particle grade to the total volume of the coarse aggregates. Since the coarse aggregate particles lack height information, an equivalent ellipsoid is used to replace the volume, and the equivalent particle size obtained by characterization is used to replace the height of the equivalent ellipse. However, there are defects in this aggregate gradation detection method. Its particle size is approximately obtained by the circumscribed rectangle, and the volume is calculated by the equivalent ellipse. The data obtained equivalently all have certain errors, which affect the accuracy of the calculated gradation; the process of calculating the particle grade requires a series of algorithms such as segmentation and equivalence, and the calculation process is complex and cumbersome.

[0003] In this context, the present invention proposes a method for detecting aggregate gradation based on a deep learning object detection algorithm combining RGB images and depth images. Through the innovative object detection algorithm, the particle grade distribution of the aggregates can be efficiently detected. By adding height information to the original volume calculation method, a more accurate volume of the aggregates can be obtained, and the detection accuracy of the gradation can be improved. Summary of the Invention

[0004] Other features and advantages of the present invention will be described in the following specification, and will become apparent in part from the specification, or will be understood by implementing the present invention. The objectives and other advantages of the present invention can be realized and obtained through the structures specifically pointed out in the specification and other specification drawings.

[0005] The objective of the present invention is to overcome the above deficiencies, and provide a coarse aggregate gradation detection method and device based on multi-modal fusion object detection. The depth image and the RGB image are fused, and a multi-modal fusion object detection model is used to automatically detect the particle size range of the aggregates. Since the features contained in a single-modal aggregate image are less and lack important height information, the fusion of the RGB image and the depth image can enrich the features of the image, improve the detection accuracy and robustness of the model. The aggregate volume calculation method combined with the image depth information can make the calculated aggregate volume more accurate due to the addition of the height information of the aggregates, which is beneficial to the subsequent gradation calculation. Correspondingly, it can be judged whether the aggregate gradation in this batch is qualified.

[0006] The present invention provides a coarse aggregate gradation detection method based on multi-modal fusion object detection, including:

[0007] S1. Aggregate pretreatment: Use a permeable sieve to screen out aggregates of different particle sizes. Collect RGB images and depth images for each particle size of aggregates for sampling. After establishing a dataset, perform data augmentation, and input the augmented dataset into the designed multi-modal fusion object detection network;

[0008] S2. Establish a multi-modal fusion object detection network: The multi-modal fusion object detection network includes an image feature extraction module, three Transformer fusion modules with pyramid structures, and an interpretation module. Input the datasets of RGB images and depth images into the image feature extraction module to extract features from the data. Input the two modal feature maps obtained into the three Transformer fusion modules with pyramid structures. Each Transformer fusion module will obtain a spliced hierarchical feature image. Interpret the three spliced hierarchical feature maps and input them into the YOLO detection head, and the YOLO detection head can output the particle size information of the corresponding level finally;

[0009] S3. Volume calculation: Collect RGB images and depth images of the aggregate raw materials passing through the production line. Process the depth image for contour, detect the particle size information of each aggregate within the current view, and calculate their respective volumes;

[0010] S4. Application in the production line: According to the volume, judge the particle size information of each aggregate in the image. Classify and process according to the particle size range, calculate the sieve residue value of each particle size range, draw the particle shape gradation curve, and judge whether this batch of aggregates is qualified.

[0011] In some embodiments, in step S1, the specific operation of data augmentation is to use the copypast algorithm to simultaneously cut out the aggregates from the RGB image and the depth image and perform the same transformation and deformation, and splice the aggregates in different particle size ranges onto a new background image to obtain depth images and RGB images containing aggregates of different particle sizes as the datasets of the multi-modal fusion object detection network in step S2.

[0012] In some embodiments, in step S2, the image feature extraction module specifically includes a channel attention module, a convolutional module, and a neck module. After the datasets of RGB images and depth images are input into the image feature extraction module, they first enter the channel attention module to split and splice the pictures along the channel direction to achieve the purpose of downsampling and improving the representation ability of the model. Then input the data into the convolutional module for feature extraction, and finally input it into the neck module to reduce the computational amount of the high-dimensional feature map and improve the representation ability of the model. The RGB image obtains an RGB image feature map after passing through the image feature extraction module, and the depth image obtains a depth image feature map after passing through the image feature extraction module.

[0013] In some embodiments, the three pyramid-structured Transformer fusion modules specifically include the Transformer1 fusion module, the Transformer2 fusion module, and the Transformer3 fusion module. Before inputting into the Transformer fusion module, an average pooling operation needs to be performed on the two modal feature maps obtained in step S2. After the average pooling operation, the feature maps of the two modalities are serialized and concatenated along the channel direction to obtain a long sequence.

[0014] In some embodiments, the long sequence is input into the Transformer1 fusion module. After being processed by the multi-head attention and self-attention in the module and decoded, a feature map that is commonly concerned by the two modalities is obtained. The feature map that is commonly concerned is added to the RGB image feature map and the depth image feature map respectively to obtain two feature images, feature1_rgb and feature1_dep, which have both common features and individual modal features.

[0015] In some embodiments, for the feature images feature1_rgb and feature1_dep, convolution and neck module processing are respectively performed to extract high-level features. The processed results are input into the Transformer2 fusion module for fusion again to obtain a relatively high-level fusion feature map. The obtained relatively high-level fusion feature map is added to the feature images feature1_rgb and feature1_dep respectively to obtain relatively high-level feature images feature2_rgb and feature2_dep. Similarly, after the fusion addition steps for the relatively high-level feature images feature2_rgb and feature2_dep, even higher-level feature images feature3_rgb and feature3_dep are obtained.

[0016] In some embodiments, in step S3, the collected depth image is binarized. The depth image is essentially a two-dimensional matrix, and the numerical points in the matrix are the distances from the depth camera to the shooting point. First, measure the distance from the depth camera to the rolling belt plane, set this value as a threshold, and determine whether the numerical points in the depth image are greater than this threshold. If greater, the value of this point is set to 255; if less, the value of this point is set to 0 to form a binary image.

[0017] In some embodiments, contour processing is performed on the binary image, that is, count the number of pixel points with a value of 255 in the binary image, and perform contour recognition processing on the obtained binary image. Based on the contour map, calculate the contour areas of each aggregate, and divide the contour area of the corresponding aggregate by the number of pixel points to obtain the area contribution value of each pixel point.

[0018] In some embodiments, each pixel is fitted into a small rectangle. The volume of the rectangle is the base area multiplied by the height. The area contribution value of the pixel is equivalent to the base area of each pixel, and the volume contribution value of each pixel is obtained by multiplying it by the corresponding height information. Finally, the volumes are summed up. Since the depth image can only collect the height information of the exposed part of the aggregate to the camera, the summation operation only obtains half of the volume of the aggregate. The aggregate is fitted into a symmetric figure along the projection plane, and the summed volume is multiplied by 2 to obtain the complete volume of the aggregate.

[0019] A coarse aggregate gradation detection device based on multi-modal fusion object detection includes:

[0020] A production line on which a rolling belt is provided for transporting aggregates;

[0021] An RGB line array camera for collecting RGB images of aggregates;

[0022] A depth camera for collecting depth images of aggregates;

[0023] A light source arranged above the production line to provide illumination light for the aggregates;

[0024] An industrial control computer electrically connected to the RGB line array camera and the depth camera to detect the coarse aggregate gradation.

[0025] By adopting the above technical solutions, the beneficial effects of the present invention are:

[0026] The present invention fuses depth images and RGB images, and uses a multi-modal fusion object detection model to automatically detect the particle size range of aggregates. Since the features contained in a single-modal aggregate image are less and lack important height information, the fusion of RGB images and depth images can enrich the features of the images, improve the detection accuracy and robustness of the model. The aggregate volume calculation method combined with the image depth information has higher accuracy in calculating the aggregate volume due to the addition of the height information of the aggregates, which is beneficial to subsequent gradation calculation. Correspondingly, it can be judged whether the aggregate gradation in this batch is qualified.

[0027] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present disclosure.

[0028] Undoubtedly, such objects of the present invention and other objects will become more obvious after the detailed description of the preferred embodiments described in multiple drawings and diagrams below.

[0029] To make the above and other objects, features and advantages of the present invention more obvious and understandable, one or several preferred embodiments are specifically given below, and detailed descriptions are made in conjunction with the accompanying drawings as follows. BRIEF DESCRIPTION OF THE DRAWINGS

[0030] The accompanying drawings are used to provide a further understanding of the present invention and form a part of the specification. They are used in conjunction with the embodiments of the present invention to explain the present invention, but do not constitute a limitation to the present invention.

[0031] In the accompanying drawings, the same components are denoted by the same reference numerals, and the drawings are schematic and not necessarily drawn to actual scale.

[0032] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the accompanying drawings required for the description of the embodiments or the prior art. Obviously, the accompanying drawings in the following description are only one or several embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can be obtained based on such drawings.

[0033] Figure 1 Schematic diagram of an RGB image and depth image acquisition device in some embodiments of the present invention;

[0034] Figure 2 Schematic diagram of a multi-modal fusion network structure in some embodiments of the present invention;

[0035] Figure 3 Schematic diagram of a volume calculation process in some embodiments of the present invention;

[0036] Figure 4 Schematic diagram of the working process of a grading calculation system in some embodiments of the present invention.

[0037] Main reference numeral description:

[0038] 1. Production line; 2. RGB line array camera; 3. Depth camera; 4. Light source; 5. Industrial control computer. Detailed implementation manners

[0039] In order to make the objectives, technical solutions and advantages of the present invention clearer, the following further details the present invention in conjunction with specific implementation manners. It should be understood that the specific implementation manners described herein are only used to explain the present invention, but not to limit the present invention.

[0040] In addition, in the description of the present invention, it should be understood that the orientation or positional relationships indicated by the terms "center", "upper", "lower", "front", "rear", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer", "axial", "radial", "circumferential", etc. are based on the orientation or positional relationships shown in the accompanying drawings. They are only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and thus should not be construed as a limitation to the present invention.

[0041] In the present invention, unless otherwise clearly defined and limited, terms such as "installation", "connection", "connection", "fixation" and the like shall be understood in a broad sense. For example, it may be a fixed connection, a detachable connection, or integrated; it may be directly connected, or indirectly connected through an intermediate medium, and may be the communication inside two components or the interaction relationship between two components. However, indicating a direct connection means that there is no connection relationship constructed through a transition structure between the two connected main bodies, and they are only connected through the connection structure to form a whole. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.

[0042] In the present invention, unless otherwise clearly defined and limited, the first feature being "above" or "below" the second feature may be that the first and second features are in direct contact, or the first and second features are in indirect contact through an intermediate medium. In the description of this specification, the description with reference to terms such as "one embodiment", "some embodiments", "examples", "specific examples", or "some examples" means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic expressions of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples.

[0043] Referring to Figures 1-4 , Figure 1 is a schematic diagram of an RGB image and depth image acquisition device in some embodiments of the present invention; Figure 2 is a schematic diagram of a multi-modal fusion network structure in some embodiments of the present invention; Figure 3 is a schematic diagram of a volume calculation process in some embodiments of the present invention; Figure 4 is a schematic diagram of the working process of a grading calculation system in some embodiments of the present invention.

[0044] According to some embodiments of the present invention, the present invention provides a method for detecting the grading of coarse aggregates based on multi-modal fusion object detection, including:

[0045] S1. Aggregate pretreatment: Use a permeable sieve to screen out aggregates of different particle sizes, sample RGB images and depth images for each particle size of aggregates, perform data augmentation after establishing a data set, and input the augmented data set into a designed multi-modal fusion object detection network;

[0046] Since there is only a single grain size of aggregate in the RGB image and depth image at this time, in order to enrich the dataset and avoid overfitting of the model, it is also necessary to perform data augmentation on the standard dataset. The specific operation for data augmentation is to use the copypast algorithm to simultaneously cut out the aggregate from the RGB image and depth image and perform the same transformation and deformation, and splice the aggregates in different particle size ranges onto a new background image to obtain depth images and RGB images containing aggregates of different grain sizes as the dataset for the multi-modal fusion object detection network in step S2. The dataset includes a training set and a validation set;

[0047] Input the training set obtained by data augmentation into the multi-modal fusion object detection network for training. After 1000 rounds of training, the corresponding model is obtained. Then, use the validation set to verify the accuracy and stability of the model. If the performance of the model is good, it can be input into the images of the distribution of aggregates of different grain sizes in normal working conditions for detection, and the particle size distribution range of each aggregate can be detected.

[0048] S2. Establish a multi-modal fusion object detection network: The multi-modal fusion object detection network includes an image feature extraction module, three Transformer fusion modules with pyramid structures, and an interpretation module. Input the datasets of the RGB image and depth image into the image feature extraction module to extract features from the data. Input the two modal feature maps obtained into the three Transformer fusion modules with pyramid structures. Each Transformer fusion module will obtain a spliced hierarchical feature image. Decode the three spliced hierarchical feature maps and input them into the yolo detection head, and the yolo detection head can output the grain size information corresponding to the final level;

[0049] In the image preprocessing step, use the semi-automatic annotation algorithm to pre-label the RGB image with the particle size range label. Then, the established dataset also has the particle size label. During the process of inputting into the multi-modal fusion object detection network, the RGB image also carries the particle size information. Therefore, the grain size information corresponding to the final level can be visually observed when the yolo detection head outputs the grain size information corresponding to the final level;

[0050] The image feature extraction module specifically includes a channel attention module, a convolution module, and a neck module. After the datasets of the RGB image and depth image are input into the image feature extraction module, they first enter the channel attention module, where the pictures are split and spliced along the channel direction to achieve the purpose of downsampling and improving the representation ability of the model. Then, the data is input into the convolution module for feature extraction. Finally, it is input into the neck module to reduce the computational amount of the high-dimensional feature map and improve the representation ability of the model. The RGB image obtains the RGB image feature map after passing through the image feature extraction module, and the depth image obtains the depth image feature map after passing through the image feature extraction module;

[0051] The three pyramid-structured Transformer fusion modules specifically include the Transformer1 fusion module, the Transformer2 fusion module, and the Transformer3 fusion module. Since the two-modal images need to be stitched and fused later, the computational complexity increases exponentially. Therefore, before inputting into the Transformer fusion module, average pooling operations need to be performed on the two modal feature maps obtained in step S2 to reduce the computational complexity. After the average pooling operation, the two modal feature maps are serialized and stitched along the channel direction to obtain a long sequence;

[0052] As Figure 2 shown, the long sequence is input into the Transformer1 fusion module. After being processed by the multi-head attention and self-attention in the module and decoded, the feature map jointly attended to by the two modalities is obtained. The jointly attended feature map is added to the RGB image feature map and the depth image feature map respectively to obtain two feature images feature1_rgb and feature1_dep that have both common features and individual modal features;

[0053] For the feature images feature1_rgb and feature1_dep, convolution and neck module processing are performed respectively to increase the receptive field and extract high-level features. The processed results are input into the Transformer2 fusion module for fusion again to obtain a higher-level fusion feature map. The obtained higher-level fusion feature map is added to the feature images feature1_rgb and feature1_dep respectively to obtain higher-level feature images feature2_rgb and feature2_dep. Similarly, after performing the fusion addition steps on the higher-level feature images feature2_rgb and feature2_dep, even higher-level feature images feature3_rgb and feature3_dep are obtained;

[0054] Among them, in the fusion addition step for the higher-level feature images feature2_rgb and feature2_dep, an SPP module is additionally set between the convolution and the neck processing module. After downsampling the data passing through the convolution, it passes through pooling layers with pooling kernel sizes of 5, 9, and 13 to capture different feature layers of the input feature map and fuse them together, and then the neck module processing is performed to reduce the computational complexity;

[0055] After passing through three pyramid-structured Transformer fusion modules, image features under different receptive fields are obtained. Then, a concatenation operation is performed to concatenate different modality feature maps with the same receptive field, that is, feature1_rgb and feature1_dep are concatenated, and so on to obtain three concatenated hierarchical feature maps. These three concatenated hierarchical feature maps are input into the YOLO detection head, and finally, the object detection results of the RGB image, namely the particle size information of each aggregate, are output.

[0056] S3. Volume calculation: Collect RGB images and depth images of the aggregate raw materials passing through on production line 1. Process the depth image for contours, detect the particle size information of each aggregate within the current view, and calculate their respective volumes.

[0057] As Figure 3 shown, perform binarization processing on the collected depth image. The depth image is essentially a two-dimensional matrix, and the numerical points in the matrix are the distances of the depth camera 3 from the shooting point. First, measure the distance from the depth camera 3 to the rolling belt plane, set this value as a threshold, and determine whether the numerical points in the depth image are greater than this threshold. If greater, set the value of this point to 255; if less, set the value of this point to 0 to form a binary image.

[0058] Perform contour processing on the binary image, that is, count the number of pixel points with a value of 255 in the binary image, and perform contour recognition processing on the obtained binary image. Based on the contour map, calculate the contour areas of each aggregate, divide the contour area of the corresponding aggregate by the number of pixel points to obtain the area contribution value of each pixel point.

[0059] Fit each pixel point into a small rectangle. The volume of the rectangle is the base area multiplied by the height. The area contribution value of the pixel point is equivalent to the base area of each pixel point. Multiply it by the corresponding height information to obtain the volume contribution value of each pixel point, and finally perform volume summation. Since the depth image can only collect the height information of the part of the aggregate exposed to the camera, the summation operation only obtains half of the volume of the aggregate. Fit the aggregate into a symmetric figure along the projection plane, and multiply the summation volume by 2 to obtain the complete volume of the aggregate.

[0060] S4. Production line application: As Figure 4 shown, based on the volume, judge the particle size information of each aggregate in the image, classify according to the particle size range, calculate the sieve residue value in each particle size range, draw the particle shape gradation curve, and judge whether this batch of aggregates is qualified.

[0061] The specific calculation method of the sieve residue value is as follows: First, calculate the sum of the volumes of the aggregates in each particle size range, compare it with the sum of the volumes of all the aggregates, calculate the cumulative sieve residue value and the cumulative sieve residue value of the aggregates in each particle size range, and draw a curve based on the cumulative sieve residue value of the aggregates in each particle size range obtained, so as to judge whether this batch of aggregates is qualified. Among them, the matplotlib library is used to draw the curve, with the particle size range of the aggregates as the horizontal axis and the cumulative sieve residue value as the vertical axis; the actual grading curve is drawn, and the actual grading curve is compared with the required grading curve to judge whether this batch of aggregates is qualified.

[0062] As Figure 1 shown, the present invention also provides a coarse aggregate grading detection device based on multi-modal fusion object detection, including:

[0063] Production line 1, on which a rolling belt is arranged for transporting aggregates;

[0064] RGB line array camera 2, which is used to collect RGB images of aggregates;

[0065] Depth camera 3, which is used to collect depth images of aggregates;

[0066] Light source 4, which is arranged above the production line 1 to provide illumination light for the aggregates;

[0067] Industrial control computer 5, which is electrically connected to the RGB line array camera 2 and the depth camera 3 to detect the grading of coarse aggregates.

[0068] It should be understood that the embodiments disclosed in the present invention are not limited to the specific processing steps or materials disclosed herein, but should extend to equivalent alternatives of such features understood by those of ordinary skill in the relevant art. It should also be understood that the terms used herein are only for the purpose of describing specific embodiments and do not mean to limit.

[0069] The "embodiments" mentioned in the specification mean that the specific features or characteristics described in connection with the embodiments are included in at least one embodiment of the present invention. Therefore, the phrase "embodiments" that appears throughout the specification does not necessarily refer to the same embodiment.

[0070] In addition, the described features or characteristics can be combined into one or more embodiments in any other suitable way. In the above description, some specific details, such as thickness, quantity, etc., are provided to provide a comprehensive understanding of the embodiments of the present invention. However, those skilled in the relevant art will understand that the present invention can be implemented without one or more of the above specific details or can also be implemented using other methods, components, materials, etc.

Claims

1. A coarse aggregate gradation detection method based on multi-modal fusion object detection, characterized in that, Including: S1. Aggregate pretreatment: Use a permeable sieve to screen out aggregates of different particle sizes. Collect RGB images and depth images for each particle size of aggregates for sampling. After establishing a dataset, perform data augmentation, and input the augmented dataset into a designed multi-modal fusion object detection network; S2. Establish a multi-modal fusion object detection network: The multi-modal fusion object detection network includes an image feature extraction module, three Transformer fusion modules with pyramid structures, and an interpretation module. Input the RGB image and depth image datasets into the image feature extraction module to extract features from the data. Input the two modal feature maps obtained into the three Transformer fusion modules with pyramid structures. Each Transformer fusion module will obtain a spliced hierarchical feature image. Decode the three spliced hierarchical feature maps and input them into the yolo detection head, and the yolo detection head can output the particle size information of the corresponding level finally; The three Transformer fusion modules with pyramid structures specifically include the Transformer1 fusion module, the Transformer2 fusion module, and the Transformer3 fusion module. Before inputting into the Transformer fusion module, an average pooling operation needs to be performed on the two modal feature maps obtained in step S2. After the average pooling operation, serialize the two modal feature maps and splice them along the channel direction to obtain a long sequence; Input the long sequence into the Transformer1 fusion module. After being processed by the multi-head attention and self-attention in the module, decode to obtain a feature map jointly concerned by the two modalities. Add the jointly concerned feature map to the RGB image feature map and the depth image feature map respectively to obtain two feature images feature1_rgb and feature1_dep that have both common features and individual modal features; For the feature images feature1_rgb and feature1_dep, perform convolution and neck module processing respectively to extract high-level features. Input the processed results into the Transformer2 fusion module for fusion again to obtain a higher-level fusion feature map. Add the obtained higher-level fusion feature map to the feature images feature1_rgb and feature1_dep respectively to obtain higher-level feature images feature2_rgb and feature2_dep. Similarly, after performing the fusion addition step on the higher-level feature images feature2_rgb and feature2_dep, obtain even higher-level feature images feature3_rgb and feature3_dep; S3. Volume calculation: Collect RGB images and depth images of the aggregate raw materials passing through the production line. Process the depth image for contour detection, detect the particle size information of each aggregate within the current view, and calculate their respective volumes; S4. Production line application: Based on the volume, determine the particle size information of each aggregate in the image, classify and process them according to the particle size range, calculate the residue values in each particle size range, draw the particle shape gradation curve, and determine whether this batch of aggregates is qualified.

2. The coarse aggregate gradation detection method based on multi-modal fusion object detection according to claim 1, wherein In step S1, the specific operation of data augmentation is to use the copypast algorithm to simultaneously extract the aggregates from the RGB image and the depth image, perform the same transformation and deformation, and splice the aggregates in different particle size ranges onto a new background image to obtain the depth image and RGB image containing aggregates of different particle sizes as the dataset for the multi-modal fusion object detection network in step S2.

3. The coarse aggregate gradation detection method based on multi-modal fusion object detection according to claim 2, wherein In step S2, the image feature extraction module specifically includes a channel attention module, a convolution module, and a neck module. After the datasets of the RGB image and the depth image are input into the image feature extraction module, they first enter the channel attention module, where the pictures are split and spliced along the channel direction to achieve the purpose of downsampling and improving the model's representation ability. Then, the data is input into the convolution module for feature extraction, and finally, it is input into the neck module to reduce the computational amount of the high-dimensional feature map and improve the model's representation ability. The RGB image feature map is obtained after the RGB image passes through the image feature extraction module, and the depth image feature map is obtained after the depth image passes through the image feature extraction module.

4. The coarse aggregate gradation detection method based on multi-modal fusion object detection according to claim 1, wherein In step S3, perform binary processing on the collected depth image. The depth image is essentially a two-dimensional matrix, and the numerical points in the matrix are the distances from the depth camera to the shooting point. First, measure the distance from the depth camera to the rolling belt plane, set this value as a threshold, and determine whether the numerical points in the depth image are greater than this threshold. If greater, set the value of this point to 255; if less, set the value of this point to 0 to form a binary image.

5. The coarse aggregate gradation detection method based on multi-modal fusion object detection according to claim 4, wherein Perform contour processing on the binary image, that is, count the number of pixel points with a value of 255 in the binary image, and perform contour recognition processing on the obtained binary image. Based on the contour image, calculate the contour area of each aggregate, and divide the contour area of the corresponding aggregate by the number of pixel points to obtain the area contribution value of each pixel point.

6. The coarse aggregate gradation detection method based on multi-modal fusion object detection according to claim 5, characterized in that, Fit each pixel point into a small rectangle. The volume of the rectangle is the base area multiplied by the height. The area contribution value of the pixel point is equivalent to the base area of each pixel point, and multiplying by the corresponding height information gives the volume contribution value of each pixel point. Finally, perform volume summation. Since the depth image can only collect the height information of the part of the aggregate exposed to the camera, the summation operation only obtains half of the volume of the aggregate. Fit the aggregate into a symmetric figure along the projection plane, and multiply the summation volume by 2 to obtain the complete volume of the aggregate.

7. The coarse aggregate grading detection device based on multi-modal fusion object detection is characterized in that Apply the coarse aggregate gradation detection method based on multi-modal fusion object detection according to any one of claims 1-6. The device includes: A production line provided with a rolling belt for transporting aggregates; An RGB line array camera for collecting RGB images of aggregates; A depth camera for collecting depth images of aggregates; A light source arranged above the production line to provide illumination light for the aggregates; An industrial control computer electrically connected to the RGB line array camera and the depth camera to detect the coarse aggregate gradation.

Citation Information

Patent Citations

  • Blast furnace sintered ore particle size detection method and system based on RGB and laser feature fusion

    CN113870341A

  • Aggregate volume calculation method and device based on visual image detection

    CN119180855A