Machine vision-based aggregate intelligent detection method

By employing 3D point cloud filtering, adaptive clustering segmentation, multi-scale CNN feature fusion, and physical constraint loss calibration, combined with the DINOv3 vision Transformer and near-infrared spectral sensor, we have achieved full-indicator, high-precision, and fully automated aggregate detection. This solves the problems of long detection cycles and large errors in existing technologies, thereby improving detection efficiency and accuracy.

CN122492613APending Publication Date: 2026-07-31ANHUI KAIYUAN HIGHWAY & BRIDGE +3
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610620942.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-05-08
Publication Date
2026-07-31

AI Technical Summary

Technical Problem

Existing aggregate testing technologies suffer from problems such as long testing cycles, large human errors, poor sampling representativeness, and inability to provide real-time feedback. Furthermore, visual inspection systems fail to achieve comprehensive and high-precision aggregate testing, especially in mud content detection, which is easily affected by light, dust, and color interference. Moisture detection does not differentiate between material characteristics and has poor versatility.

Method used

The system employs a 3D line laser camera, a high-resolution RGB line array camera, and a near-infrared spectral sensor to simultaneously acquire aggregate images and spectral information. Combined with 3D point cloud dynamic filtering, adaptive clustering segmentation, multi-scale CNN feature fusion, and physical constraint loss calibration, it achieves high-precision measurement of crushed stone volume, equivalent particle size, and length/width/thickness dimensions. Furthermore, it uses the DINOv3 vision Transformer for mud-containing semantic segmentation and near-infrared moisture prediction.

Benefits of technology

It achieves full-index, high-precision, and fully automated aggregate testing, overcomes interference under complex working conditions, improves the consistency and reliability of test results, significantly shortens the testing cycle, and reduces manpower and laboratory operating costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122492613A_ABST
    Figure CN122492613A_ABST
Patent Text Reader

Abstract

This invention discloses a machine vision-based intelligent aggregate detection method, belonging to the field of aggregate detection technology. The invention includes the following steps: synchronously acquiring aggregate images and spectral information from multiple sensors; denoising, segmenting, and projecting 3D point clouds to generate standardized depth maps; outputting single-particle volume, equivalent particle size, and length / width / height using a depth map convolutional neural network to complete gradation statistics and needle-like / flaky particle determination; performing pixel-level segmentation on RGB images to accurately identify mud-block regions and calculate mud content; automatically switching near-infrared moisture models according to material type to output high-precision moisture content; and integrating all indicators to generate a traceable detection report. This invention achieves high-precision measurement of crushed stone volume, equivalent particle size, and length / width / thickness dimensions through 3D point cloud dynamic filtering, adaptive clustering segmentation, multi-scale CNN feature fusion, and physical constraint loss calibration; improving the accuracy of needle-like / flaky particle identification, mud content identification, and moisture content detection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of aggregate testing technology, and in particular relates to an online intelligent testing method for all indicators of aggregates (crushed stone and medium sand) based on the fusion of machine vision, multispectral sensing and deep learning. It realizes real-time, continuous and full-sample testing of crushed stone geometry, gradation distribution, needle-like and flaky content, mud content and medium sand moisture content, which greatly improves the testing accuracy, efficiency and objectivity, and provides technical support for the closed-loop quality of the entire concrete production process. Background Technology

[0002] In engineering fields such as concrete production, ready-mixed mortar, and road construction materials, the geometric gradation, flaky and needle-like particle content, mud content, and moisture content of aggregates (crushed stone, manufactured sand, and natural sand) are core indicators that determine the strength, workability, and durability of concrete. Traditional aggregate testing relies on methods such as manual sampling, laboratory sieving, drying, and caliper measurement, which suffer from problems such as long testing cycles, large human errors, poor sampling representativeness, and inability to provide real-time feedback, making it difficult to meet the quality control requirements of intelligent construction, green construction, and continuous production.

[0003] In existing technologies, some visual inspection systems only achieve single-size detection and do not integrate 3D point cloud, hyperspectral and deep learning algorithms; they are easily affected by light, dust and color interference in mud detection, resulting in insufficient recognition accuracy; and they do not distinguish between crushed stone and medium sand material characteristics in moisture detection, resulting in poor versatility.

[0004] Therefore, the industry urgently needs a comprehensive, high-precision, fully automated, and rapidly deployable intelligent aggregate detection method. Summary of the Invention

[0005] The purpose of this invention is to provide a machine vision-based intelligent aggregate detection method. Through technologies such as 3D point cloud dynamic filtering, adaptive clustering segmentation, multi-scale CNN feature fusion, and physical constraint loss calibration, it achieves high-precision measurement of crushed stone volume, equivalent particle size, and length / width / thickness dimensions, solving the problems of inaccurate, large error, and long cycle of existing aggregate mud content detection.

[0006] To solve the above-mentioned technical problems, the present invention is achieved through the following technical solution:

[0007] This invention relates to a machine vision-based intelligent material detection method, comprising the following steps:

[0008] Step S1, Multi-sensor synchronous acquisition: Simultaneous acquisition of aggregate images and spectral information using a 3D line laser camera, a high-resolution RGB line array camera, and a near-infrared spectral sensor;

[0009] Step S2, Preprocessing of gravel point cloud and conversion of depth map: Denoise, segment, and project the 3D point cloud to generate a standardized depth map;

[0010] Step S3: Intelligent analysis of gravel volume and shape based on CNN: Output the volume of a single stone, equivalent particle size, length / width / height through a deep graph convolutional neural network to complete the gradation statistics and needle-like / flaky determination;

[0011] Step S4: Semantic segmentation of mud based on DINOv3 visual Transformer: Pixel-level segmentation of RGB image to accurately identify mud block regions and calculate mud content;

[0012] Step S5, Dual-calibration moisture prediction model for crushed stone / medium sand: Automatically switches to the near-infrared moisture model according to the material type, and outputs high-precision moisture content;

[0013] Step S6, Data Fusion and Report Output: Integrate all indicators to generate a traceable testing report.

[0014] As a preferred technical solution, the specific process of multi-sensor synchronous data acquisition in step S1 is as follows:

[0015] Step S11: Using the hard-triggered synchronization signal as a unified timing reference, the 3D line laser camera, high-resolution RGB line array camera, and near-infrared spectral sensor are coaxially deployed along the material conveying direction.

[0016] Step S12: When the aggregate on the conveyor belt passes through the unified detection area, the photoelectric switch triggers a pulse signal, simultaneously starting 3D point cloud scanning, RGB linear array imaging and near-infrared spectrum acquisition;

[0017] Step S13: During the acquisition process, the 3D camera and the RGB camera maintain the same line frequency and pixel resolution, and the near-infrared sensor maintains the same sampling rate to ensure that the data of the three devices correspond one-to-one in spatial location and time series.

[0018] As a preferred technical solution, the specific process of preprocessing the gravel point cloud and converting the depth map in step S2 is as follows:

[0019] Step S21, Raw Point Cloud Acquisition and Format Standardization: Raw point cloud data of a single layer of paved gravel on the conveyor belt is acquired using a line laser 3D camera (resolution ≥ 1280 pixels, sampling frequency ≥ 10kHz). The output format is PCD (Point Cloud Data), and the number of point clouds per frame is controlled within a specified range. There are points, and the coordinates of each point cloud are defined as follows: In the formula, In three-dimensional space coordinates, These are the RGB color values ​​collected synchronously.

[0020] Step S22, Multi-dimensional hierarchical filtering and denoising: Traverse each point cloud, count the number of points in its neighborhood, and remove isolated noise points with less than 5 neighborhood points.

[0021] Step S23, Adaptive Euclidean Clustering for Individual Crushed Stones: Calculate the Z-axis variance of the filtered point cloud to characterize the undulation of the aggregate surface, and set an adaptive distance threshold to perform Euclidean clustering;

[0022] Step S24, Point Cloud Projection and Depth Map Generation: Project the single gravel point cloud orthogonally along the Z-axis onto the XOY plane, establish the mapping relationship between the pixel coordinate system and the spatial coordinate system, and take the maximum Z value of the corresponding spatial position as the depth value for each pixel in the pixel coordinate system.

[0023] Step S25, Depth Map Normalization and Enhancement: Map depth values ​​to... The grayscale range is used, and the Sobel operator is used to extract the edges of the depth map, which are then superimposed onto the normalized depth map.

[0024] As a preferred technical solution, in step S24, the point cloud of a single gravel particle is orthogonally projected along the Z-axis onto the XOY plane to establish a pixel coordinate system. With spatial coordinate system The mapping relationship is as follows:

[0025] , ;

[0026] In the formula, The depth map resolution is in pixels. , The minimum X / Y coordinates of the point cloud of a single gravel;

[0027] For each pixel in the pixel coordinate system The maximum Z value at the corresponding spatial location is taken as the depth value, and the specific formula is as follows:

[0028] ;

[0029] The holes in the depth map are repaired using bilateral filtering interpolation, with the specific formula as follows:

[0030] ;

[0031] In the formula, For neighborhood windows, For depth similarity weights, Using spatial distance weights, it both repairs the voids and preserves the edge features of the rubble.

[0032] As a preferred technical solution, the specific process of intelligent analysis of gravel volume and shape based on CNN in step S3 is as follows:

[0033] Step S31, Depth Map Preprocessing and Dataset Construction: The enhanced depth map generated in step S2... Scaling to uniform Pixels were analyzed and mean-variance normalization was performed. 100,000 sets of depth maps and measured parameter samples were collected, covering [the data / databases]. Across the entire particle size range, the depth map of small-sized crushed stone is stretched to scale, while large-sized crushed stone is locally trimmed and spliced.

[0034] Step S32: Construction of a multi-scale feature fusion CNN network: Dynamic weights are assigned to the feature maps of the three branches. Pooling is used to reduce the dimensionality of the fused feature map to a 1024-dimensional feature vector. Four parallel fully connected layers are designed to output the volume respectively. ,length ,width and thickness ;

[0035] Step S33, Physical constraint loss function training: Design a physical constraint hybrid loss function to balance numerical accuracy and physical rationality;

[0036] Step S34, Model Inference and Dynamic Calibration: Input the preprocessed depth map into the trained MSF-CNN and output the initial prediction value. , , , The calibration coefficients are preset according to the particle size range;

[0037] Step S35, Grading Statistics and Needle-like / Flake-like Determination: Calculate the equivalent particle size based on the calibrated volume, count the proportion of crushed stone in each particle size range, and generate a gradation curve.

[0038] As a preferred technical solution, in step S33, the physical constraint hybrid loss function is as follows:

[0039] ;

[0040] In the formula, For mean square error loss, For physical volume constraint loss, Loss due to size sorting constraints;

[0041] The formula for calculating the mean squared error loss is as follows: ;

[0042] The formula for calculating the physical volume constraint loss is as follows: ;

[0043] The formula for calculating the size sorting constraint loss is as follows: ;

[0044] In the formula, , , , For the first Predicted values ​​for volume, length, width, and thickness. , , , For the first Measured values ​​of volume, length, width, and thickness.

[0045] As a preferred technical solution, the formula for calculating the equivalent particle size in step S35 is as follows: The percentage of crushed stone in each particle size range is calculated based on the sieve aperture size to generate a gradation curve. In the formula, Screen aperture size The pass rate For equivalent particle size smaller than The amount of gravel.

[0046] As a preferred technical solution, the specific process of mud-containing semantic segmentation based on DINOv3 visual Transformer in step S4 is as follows:

[0047] Step S41, RGB image acquisition and scene-based preprocessing: Use a linear RGB camera to acquire images of gravel and mud, and simultaneously record the light intensity at the time of acquisition for adaptive light correction.

[0048] Step S42, Adaptive DINOv3 Model Reconstruction for Aggregate Scene: Insert an aggregate texture attention module into the Transformer Encoder of DINOv3. This module learns the texture differences between mud (loose, porous) and gravel (dense, smooth), strengthens the mud feature weights, simplifies the multi-class output of general segmentation to 3 classes (gravel body, mud, background), and adds a boundary refinement head to perform pixel-level correction on the edges of the segmentation mask.

[0049] Step S43, Pixel-level segmentation and accurate calculation of mud content: Input the preprocessed enhanced image into the trained DINOv3 model, and output a pixel-level segmentation mask. Perform connected component analysis on the segmentation mask and calculate the area of ​​each connected component. The minimum mud block area threshold is used to fill in the small mud blocks, and the mud content is calculated by counting the number of mud block pixels in the effective area and the total number of pixels.

[0050] As a preferred technical solution, in step S43, a minimum mud block area threshold is set. Regarding area However, isolated pixels that match the texture features of the mud are identified as small mud lumps and added to the mask. The formula is as follows:

[0051] ;

[0052] Count the number of mud block pixels within the valid area Total number of pixels Introducing pixel-to-actual area calibration coefficient Convert pixel percentage to actual area percentage:

[0053] ;

[0054] In the formula, The number of background pixels. This represents the final mud content.

[0055] As a preferred technical solution, the specific process of automatically switching the near-infrared moisture model according to the material type and outputting high-precision moisture content in step S5 is as follows:

[0056] Step S51, Near-infrared spectral acquisition and multi-dimensional preprocessing: Based on the near-infrared absorption characteristics of moisture in crushed stone / medium sand, three core characteristic wavelengths were selected (1300nm (first-order overtone absorption of water OH bonds), 1450nm (coupling absorption of crushed stone / medium sand matrix and moisture), and 1940nm (strong second-order overtone absorption of water OH bonds)). A narrow-band near-infrared sensor (bandwidth ±5nm) was used to acquire the spectral absorbance, and the raw absorbance sequence was output. ( (t is the data acquisition time).

[0057] Step S52, Intelligent Material Identification and Model Switching Trigger: Extract material texture features from RGB images (crushed stone: sharp edges and discrete particles; medium sand: fine particles and uniform texture), calculate the absorbance ratio of three feature wavelengths, construct a lightweight classifier, and output material type labels by fusing visual and spectral features;

[0058] The formula for calculating the absorbance ratio of the three characteristic wavelengths is:

[0059] ;

[0060] Step S53, Dual-calibration model construction: Divide the model into two sub-models according to the moisture content range. The model formula is as follows:

[0061] ;

[0062] In the formula, , , These are the interval weighting coefficients. , This is the offset. This refers to the moisture content of the crushed stone. It is a near-infrared characteristic wavelength. The near-infrared absorbance value at a certain wavelength after smoothing and noise reduction;

[0063] Step S54, Moisture Content Output and Accuracy Verification: Based on the material type, call the corresponding model to output the initial moisture content, take the moving average of the predicted values ​​for 10 consecutive frames, and output the final moisture content.

[0064] The present invention has the following beneficial effects:

[0065] (1) This invention achieves high-precision measurement of crushed stone volume, equivalent particle size, length / width / thickness by using 3D point cloud dynamic filtering, adaptive clustering segmentation, multi-scale CNN feature fusion, physical constraint loss calibration and other technologies; and improves the accuracy of needle-like identification, mud content identification and moisture content detection.

[0066] (2) This invention integrates 3D vision, RGB linear array imaging, near-infrared spectroscopy and DINOv3Transformer model, and through innovative designs such as adaptive illumination correction, aggregate texture attention enhancement, and environmental temperature and humidity / distance compensation, effectively overcomes interference from on-site dust, illumination changes, material adhesion, particle size differences, etc., and ensures stable operation under complex working conditions.

[0067] (3) This invention establishes a segmented moisture model for crushed stone and a moisture compensation model for medium sand matrix, and automatically identifies material type through visual + spectral dual-modal identification, achieving one-click switching without manual intervention or hardware replacement. The same system can complete the detection of all indicators of crushed stone and the moisture content of medium sand, significantly improving equipment utilization and scene adaptability.

[0068] (4) This invention introduces a physical constraint loss function into the CNN volume shape analysis to force the size and volume to meet the physical logic, thereby avoiding problems such as size contradictions and volume distortion; at the same time, combined with dynamic particle size interval calibration, it greatly reduces system error and improves the consistency and reliability of detection results.

[0069] (5) The present invention realizes full automation of image acquisition, point cloud processing, semantic segmentation, index calculation and report generation, eliminating the need for manual sampling, sieving, weighing, drying and recording. The detection cycle is shortened from hours to seconds, significantly improving quality inspection efficiency and reducing manpower and laboratory operating costs.

[0070] Of course, any product implementing this invention does not necessarily need to achieve all of the advantages described above at the same time. Attached Figure Description

[0071] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0072] Figure 1 This is a schematic diagram of the overall system structure of the present invention;

[0073] Figure 2 Flowchart for converting 3D point cloud of gravel into depth map;

[0074] Figure 3 Here is a flowchart of semantic segmentation of muddy gravel based on DINOv3;

[0075] Figure 4 A framework diagram of a dual-calibrated moisture prediction model for crushed stone and medium sand;

[0076] Figure 5 This is a schematic diagram of the sensor layout. Detailed Implementation

[0077] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0078] Furthermore, the technical features involved in the various embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.

[0079] To make the purpose, technical solution, and advantages of this application clearer, the following description is provided in conjunction with the appendix. Figure 1-5 The present application will be further described in detail below with reference to embodiments. It should be understood that the specific embodiments described herein are for illustrative purposes only and are not intended to limit the scope of the application.

[0080] Please see Figure 1 As shown, this invention is a machine vision-based intelligent material detection method, comprising the following steps:

[0081] Step S1, Multi-sensor synchronous acquisition: Simultaneous acquisition of aggregate images and spectral information using a 3D line laser camera, a high-resolution RGB line array camera, and a near-infrared spectral sensor;

[0082] Step S2, Preprocessing of gravel point cloud and conversion of depth map: Denoise, segment, and project the 3D point cloud to generate a standardized depth map;

[0083] Step S3: Intelligent analysis of gravel volume and shape based on CNN: Output the volume of a single stone, equivalent particle size, length / width / height through a deep graph convolutional neural network to complete the gradation statistics and needle-like / flaky determination;

[0084] Step S4: Semantic segmentation of mud based on DINOv3 visual Transformer: Pixel-level segmentation of RGB image to accurately identify mud block regions and calculate mud content;

[0085] Step S5, Dual-calibration moisture prediction model for crushed stone / medium sand: Automatically switches to the near-infrared moisture model according to the material type, and outputs high-precision moisture content;

[0086] Step S6, Data Fusion and Report Output: Integrate all indicators to generate a traceable testing report.

[0087] Please see Figure 5 As shown, the specific process of multi-sensor synchronous data acquisition in step S1 is as follows:

[0088] Step S11: Using the hard-triggered synchronization signal as a unified timing reference, the 3D line laser camera, high-resolution RGB line array camera, and near-infrared spectral sensor are coaxially deployed along the material conveying direction.

[0089] Step S12: When the aggregate on the conveyor belt passes through the unified detection area, the photoelectric switch triggers a pulse signal, simultaneously starting 3D point cloud scanning, RGB linear array imaging and near-infrared spectrum acquisition;

[0090] Step S13: During the acquisition process, the 3D camera and the RGB camera maintain the same line frequency and pixel resolution, and the near-infrared sensor maintains the same sampling rate to ensure that the data of the three devices correspond one-to-one in spatial location and time series.

[0091] Specifically, a 3D line laser camera, an RGB line array camera, and a near-infrared spectral probe are installed directly above the aggregate conveyor belt, coaxially arranged along the conveying direction, with their optical axes converging in the same detection area. The detection area is 600mm wide and 400mm high. A through-beam photoelectric switch is installed at the front end of the detection area, with a response time ≤1ms, outputting a TTL 5V pulse signal, which is connected to the external trigger interfaces of the 3D camera, RGB camera, and spectral sensor, respectively. The line frequency of the 3D camera and RGB camera is set to 10kHz, with a single image resolution of 2048 pixels. The sampling rate of the near-infrared spectral sensor is set to 1000Hz, matching the conveyor belt speed of 1.5m / s. When crushed stone / medium sand passes through the detection area, it blocks the optical path of the photoelectric switch, immediately outputting a trigger pulse. All three start acquiring data simultaneously without time deviation. The system reads the timestamps of the three data points with an error ≤0.1ms. Through coordinate mapping, the 3D point cloud, RGB image, and spectral data are accurately matched to form multi-source fusion data of a single aggregate particle, which is used for subsequent volume analysis, mud content segmentation, and moisture prediction.

[0092] In practice, a top-vertical layout with a coaxial field of view is adopted, integrating a 3D line laser camera, an RGB line array camera, and a near-infrared spectral sensor directly above the aggregate conveyor belt. The centers of the three fields of view coincide and their acquisition areas completely overlap, ensuring that the same aggregate is detected by all three sensors simultaneously, achieving synchronous acquisition of geometric, texture, and composition information.

[0093] During installation, all sensors should be installed directly above the conveyor belt, collecting data vertically downwards, with the installation height at a distance from the material surface. (Optimal detection range); Installation method: shared bracket, coaxial, same field of view;

[0094] The 3D line laser camera is installed on the far left to emit a line laser to scan the three-dimensional point cloud of the crushed stone, and the laser covers the entire width of the conveyor belt.

[0095] A high-resolution RGB line scan camera is mounted in the center (main view) to capture the surface texture of gravel for mud segmentation, and its field of view is completely overlapped with that of the 3D camera;

[0096] The near-infrared spectral sensor is located on the far right, with the center of the light spot aligned with the center of the 3D and RGB fields of view, used to collect the moisture spectrum of materials in the same area.

[0097] Please see Figure 2 As shown, the specific process of gravel point cloud preprocessing and depth map conversion in step S2 is as follows:

[0098] Step S21, Raw Point Cloud Acquisition and Format Standardization: Raw point cloud data of a single layer of paved gravel on the conveyor belt is acquired using a line laser 3D camera (resolution ≥ 1280 pixels, sampling frequency ≥ 10kHz). The output format is PCD (Point Cloud Data), and the number of point clouds per frame is controlled within a specified range. There are points, and the coordinates of each point cloud are defined as follows: In the formula, In three-dimensional space coordinates, These are the RGB color values ​​collected synchronously.

[0099] Step S22, Multi-dimensional hierarchical filtering and denoising: Traverse each point cloud, count the number of points in its neighborhood, and remove isolated noise points with less than 5 neighborhood points.

[0100] Specifically, traditional Euclidean clustering uses a fixed distance threshold, while this invention innovatively proposes an "adaptive distance threshold based on particle size range" to achieve accurate segmentation of adhering gravel.

[0101] The specific process is as follows:

[0102] Step S221: Calculate the Z-axis variance of the filtered point cloud. Characterizes the degree of surface undulation of aggregates;

[0103] Step S222: Adaptive distance threshold calculation: ;

[0104] Step S223: with Euclidean clustering is performed on a threshold to divide the point cloud into multiple subsets. Each subset corresponds to a single piece of gravel;

[0105] Step S224: Remove tiny subsets (impurities / debris) with less than 100 points, and retain valid single-gravel point clouds. .

[0106] Step S23, Adaptive Euclidean Clustering for Individual Crushed Stones: Calculate the Z-axis variance of the filtered point cloud to characterize the undulation of the aggregate surface, and set an adaptive distance threshold to perform Euclidean clustering;

[0107] Step S24, Point Cloud Projection and Depth Map Generation: Project the single gravel point cloud orthogonally along the Z-axis onto the XOY plane, establish the mapping relationship between the pixel coordinate system and the spatial coordinate system, and take the maximum Z value of the corresponding spatial position as the depth value for each pixel in the pixel coordinate system.

[0108] Step S25, Depth Map Normalization and Enhancement: Map depth values ​​to... The grayscale range is used, and the Sobel operator is used to extract the edges of the depth map, which are then superimposed onto the normalized depth map.

[0109] In step S24, the point cloud of a single gravel particle is orthogonally projected along the Z-axis onto the XOY plane to establish a pixel coordinate system. With spatial coordinate system The mapping relationship is as follows:

[0110] , ;

[0111] In the formula, The depth map resolution is in pixels. , The minimum X / Y coordinates of the point cloud of a single gravel;

[0112] For each pixel in the pixel coordinate system The maximum Z value at the corresponding spatial location is taken as the depth value to avoid pixel holes. The specific formula is as follows:

[0113] ;

[0114] Holes in the depth map are repaired using bilateral filtering interpolation, with the specific formula as follows:

[0115] ;

[0116] In the formula, For neighborhood windows, For depth similarity weights, Using spatial distance weights, it both repairs the voids and preserves the edge features of the rubble.

[0117] In step S25, the depth value is mapped to... Grayscale range, enhancing visual features:

[0118] ;

[0119] In the formula, The maximum and minimum depth values ​​of a single piece of gravel are used to extract the edges of the depth map using the Sobel operator. The overlaid values ​​are then normalized to highlight the gravel outline. The formula is as follows:

[0120] , ;

[0121] ;

[0122] In the formula, This represents the gradient value of the depth map in the horizontal direction. This represents the gradient value of the depth map in the vertical direction. This is the normalized depth map. For the edge-enhanced depth map in pixel coordinates The grayscale value at that location;

[0123] Edge enhancement preserves key features of the gravel shape, providing more accurate input for subsequent CNN analysis of volume / shape.

[0124] In practical implementation, taking the point cloud preprocessing and depth map conversion of 16mm diameter gravel as an example:

[0125] The parameters are set as follows: conveyor belt speed Single frame acquisition time average particle size of aggregate Z-axis effective range , Depth map pixel resolution .

[0126] The execution process is as follows:

[0127] 1. Collect raw point clouds After coordinate calibration: (Eliminate 10mm conveyor belt offset);

[0128] 2. Use pass-through filtering to remove point clouds with a Z-axis diameter <5mm or >25mm, retaining only valid point clouds. ;

[0129] 3. Calculate the dynamic neighborhood radius Radius filtering removes noise points with fewer than 5 neighboring points, resulting in... ;

[0130] 4. Calculate the variance of the Z-axis. Adaptive clustering threshold Euclidean clustering segmented 12 individual fragment point clouds;

[0131] 5. For one of the gravel pieces ( , Projection generates pixel coordinate system: The initial depth map is obtained by filling the depth value. ;

[0132] 6. Two-sided filtering interpolation was used to repair three holes, and the normalized depth values ​​were within the specified range. ;

[0133] 7. The Sobel operator is used for edge enhancement, ultimately generating a gravel depth map with a resolution of 200×150 pixels.

[0134] This innovative approach integrates three core design elements: multi-scale feature fusion CNN (MSF-CNN), physically constrained loss function, and dynamic size calibration. It addresses industry pain points such as poor generalization, lack of physical meaning, and large cross-range errors in aggregate volume / size prediction, achieving an upgrade from "data fitting" to a "physical + data dual-driven" approach. The specific process is as follows:

[0135] In step S3, the specific process of intelligent analysis of gravel volume and shape based on CNN is as follows:

[0136] Step S31, Depth Map Preprocessing and Dataset Construction: The enhanced depth map generated in step S2... Scaling to uniform Pixels were collected and mean-variance normalized. 100,000 sets of depth maps and measured parameter samples were acquired (measured parameters: volume measured by drainage method, length / width / height measured by calipers), covering... For the entire particle size range, the depth map of small-sized gravel is scaled and the large-sized gravel is locally cropped and spliced ​​to solve the problem of insufficient small-sized samples. Then, the training set, validation set and test set are divided according to 7:2:1.

[0137] Specifically, the mean-variance normalization equation is:

[0138] ;

[0139] In the formula, This is the mean of the depth map. This represents the standard deviation of the depth map.

[0140] Step S32: Construction of a multi-scale feature fusion CNN network: Dynamic weights are assigned to the feature maps of the three branches. Pooling is used to reduce the dimensionality of the fused feature map to a 1024-dimensional feature vector. Four parallel fully connected layers are designed to output the volume respectively. ,length ,width and thickness ;

[0141] Specifically, traditional CNNs tend to lose fine-grained features of depth maps (such as gravel edges and depressions). The innovation of the MSF-CNN designed in this invention lies in "hierarchical feature extraction + cross-scale fusion", and the network structure is as follows:

[0142] 1. Input layer: Receives a 224×224×1 depth map (single-channel grayscale image);

[0143] 2. Multi-scale feature extraction module:

[0144] Branch 1 (small scale): 3×3 convolution kernel, stride 1, extracts fine-grained features such as edge and texture of gravel; Branch 2 (medium scale): 5×5 convolution kernel, stride 1, extracts overall contour features of gravel; Branch 3 (large scale): 7×7 convolution kernel, stride 2, extracts global spatial features of gravel.

[0145] 3. Feature Fusion Layer (Creative Design): Employs attention-weighted fusion to assign dynamic weights to the feature maps of the three branches (adaptively adjusted based on feature contribution).

[0146] ;

[0147] In the formula, , Learned through an attention mechanism, it prioritizes retaining features that are strongly correlated with volume / size;

[0148] 4. Dimensionality Reduction Layer: The fused feature map is reduced to a 1024-dimensional feature vector through pooling and 1×1 convolution;

[0149] 5. Regression Head: Design 4 parallel fully connected layers to output volume respectively. ,length ,width ,thickness (Avoid error coupling from a single output head).

[0150] Step S33, Physical constraint loss function training: Design a physical constraint hybrid loss function to balance numerical accuracy and physical rationality;

[0151] Step S34, Model Inference and Dynamic Calibration: Input the preprocessed depth map into the trained MSF-CNN and output the initial prediction value. , , , The calibration coefficients are preset according to the particle size range;

[0152] Specifically, according to particle size range ( , , , Preset calibration coefficient , , , ( (for particle size range);

[0153] First, through the initial equivalent particle size Determine the particle size range, then perform calibration:

[0154] ;

[0155] ;

[0156] The calibration coefficients are obtained by fitting the error of measured samples within the interval, which solves the systematic error across particle size intervals, filters outliers, and removes predicted values ​​that obviously do not conform to physical laws (such as volume < 0, length / width / height ratio > 10), thus ensuring the validity of the output.

[0157] Step S35, Grading Statistics and Needle-like / Flake-like Determination: Calculate the equivalent particle size based on the calibrated volume, count the proportion of crushed stone in each particle size range, and generate a gradation curve.

[0158] Traditional MSE loss only focuses on the difference between the predicted value and the true value, without physical constraints, and is prone to "logical contradictions in length / width / height" (such as thickness > width). Therefore, step S33 designs a hybrid loss function with physical constraints as follows:

[0159] ;

[0160] In the formula, For mean square error loss, For physical volume constraint loss, Loss due to size sorting constraints;

[0161] The formula for calculating the mean squared error loss (to ensure numerical accuracy) is as follows: ;

[0162] The formula for calculating physical volume constraint loss is: ; used to force the predicted size to satisfy the volume physical relationship (V≈L×W×T×η, where η is the shape coefficient of the crushed stone, taken as η). ).

[0163] Size sorting constraint loss avoidance, logical contradiction avoidance (forced) The calculation formula is: ;

[0164] In the formula, , , , For the first Predicted values ​​for volume, length, width, and thickness. , , , For the first Measured values ​​of volume, length, width, and thickness;

[0165] Reset the weights: =0.6, =0.3, =0.1, balancing numerical accuracy with physical rationality.

[0166] In step S35, the formula for calculating the equivalent particle size is as follows: The proportion of crushed stone in each particle size range was statistically analyzed according to sieve aperture sizes (2.36mm, 4.75mm, 9.5mm, 16mm, 19mm, 26.5mm, 31.5mm) to generate gradation curves. In the formula, Screen aperture size The pass rate For equivalent particle size smaller than The amount of gravel.

[0167] Needle-like or flaky particles are determined as follows: Based on the calibrated length / width / height, and according to national standards, needle-like particles are considered... ; flaky particles are The percentage of needle-like and flaky particles in the total particles is calculated, and the needle-like and flaky particle content is output.

[0168] In specific implementation, Taking the volume and shape analysis of crushed stone by particle size as an example:

[0169] Parameter settings: MSF-CNN network: Input 224×224 depth map, 3 feature branches, attention fusion weights Physical constraint loss weights: Shape factor ; Particle size range calibration coefficient: .

[0170] The execution process is as follows:

[0171] 1. Take the product generated in step S2 The depth map is enhanced by gravel, and then normalized to obtain the following results. (mean) Standard deviation );

[0172] 2. Input MSF-CNN, multi-scale feature branches extract edge, contour, and global features, and after attention fusion, output the initial prediction value: ;

[0173] 3. Calculate the initial equivalent particle size Determined to belong to Perform dynamic calibration within the interval:

[0174] ; ;

[0175] 4. Verify physical constraints: ,and The error is only 0.8%;

[0176] 5. Calculate the equivalent particle size: ;

[0177] 6. Determination of needle-like and flaky appearance: It was determined to be needle-shaped particles;

[0178] 7. Grading statistics: The equivalent particle size of this crushed stone is 17.3mm. The passing rate through a 16mm sieve is calculated. If the total number of particles is 1000, and 450 of them have an equivalent particle size <16mm, then the passing rate through a 16mm sieve is 45%.

[0179] The experimental results are as follows: The measured volume of the crushed stone is 0.000178 m³, with a length of 28 mm, a width of 18 mm, and a thickness of 9 mm. The errors between the MSF-CNN predicted value and the measured value are: volume 1.0%, length 1.0%, and width 1.0%, which is far better than traditional CNN (volume error 5%+, size error 3%+). The needle-like and flaky shape determination results are consistent with the manual caliper determination, and the deviation between the gradation statistics and the laboratory sieving results is <2%.

[0180] Please see Figure 3 As shown, the specific process of mud-containing semantic segmentation based on DINOv3 visual Transformer in step S4 is as follows:

[0181] Step S41, RGB Image Acquisition and Scene Preprocessing: A linear RGB camera (4096×2048 resolution, 30fps) is used to acquire images of gravel and mud, and the light intensity at the time of acquisition is recorded simultaneously. Perform adaptive correction for illumination;

[0182] Specifically, the output high-fidelity image format is an RGB three-channel matrix. Perform adaptive illumination correction:

[0183] Calculate the average brightness of the image Introducing a light compensation coefficient ( (Based on standard illumination intensity), the image is corrected pixel by pixel:

[0184] ;

[0185] in, The standard average brightness is used to solve the segmentation failure problem caused by strong light / backlight;

[0186] Guided filtering is used to preserve the edges of mud / stone while removing noise, avoiding the blurring of textures by traditional Gaussian filtering; histogram equalization is performed on the R channel (the channel with the most significant color difference between mud and gravel) to enhance the texture features of mud.

[0187] Step S42, Adaptive DINOv3 Model Reconstruction for Aggregate Scene: Insert an aggregate texture attention module into the Transformer Encoder of DINOv3. This module learns the texture differences between mud (loose, porous) and gravel (dense, smooth), strengthens the mud feature weights, simplifies the multi-class output of general segmentation to 3 classes (gravel body, mud, background), and adds a boundary refinement head to perform pixel-level correction on the edges of the segmentation mask.

[0188] Specifically, the model input resolution was increased from 224×224 to 512×512, preserving the detailed features of small mud lumps (<5mm); an aggregate texture attention module was inserted into the Transformer Encoder of DINOv3. This module strengthens the feature weights of mud lumps by learning the texture differences between mud lumps (loose, porous) and gravel (dense, smooth).

[0189] ;

[0190] In the formula, Texture mask (obtained by offline learning of aggregate texture features). To achieve element-wise multiplication, a hierarchical feature fusion method is employed: the four-layer feature maps output by the Encoder (downsampled by 1 / 2 / 4 / 8 times) are upsampled to the same resolution and then fused to address the feature loss issue of small mud clumps. This simplifies the multi-class output of general segmentation to three classes (gravel body, mud clump, and background), and adds a boundary refinement head to perform pixel-level correction on the edges of the segmentation mask. The formula is:

[0191] ;

[0192] In the formula, This is the original segmentation mask. The output of the edge detection operator ensures that the segmentation boundary is aligned with the actual mud / stone edge.

[0193] Step S43, Pixel-level segmentation and accurate calculation of mud content: Input the preprocessed enhanced image into the trained DINOv3 model, and output a pixel-level segmentation mask. (Values: 0 = background, 1 = gravel, 2 = mud), perform connected component analysis on the segmentation mask and calculate the area of ​​each connected component. The minimum mud block area threshold is used to fill in the small mud blocks, and the mud content is calculated by counting the number of mud block pixels in the effective area and the total number of pixels.

[0194] In step S43, a minimum mud block area threshold is set. (corresponding to actual) ), for area However, isolated pixels that match the texture features of the mud are identified as small mud lumps and added to the mask. The formula is as follows:

[0195] It solved the problem of underestimated mud content caused by missed detection of small mud lumps;

[0196] Count the number of mud pixels within the valid region (excluding background). Total number of pixels Introducing pixel-to-actual area calibration coefficient (Through the calibration board, Convert pixel percentage to actual area percentage:

[0197] ;

[0198] In the formula, The number of background pixels. This represents the final mud content.

[0199] In practical implementation, taking the semantic segmentation of mud-bearing gravel with a particle size of 16mm as an example:

[0200] The parameters are set as follows: Standard light intensity Standard brightness average DINOv3 model input resolution Texture attention module mask Offline learning from 10,000 aggregate texture samples; minimum mud block area threshold. Pixels, pixel-to-actual-area calibration factor ;

[0201] The execution process is as follows:

[0202] ① Collect RGB images of mud and gravel, and measure the actual light intensity. (Backlit scene), average image brightness ;

[0203] ②Light correction: , After correction, the image brightness was restored to the standard level;

[0204] ③Texture Enhancement: Histogram equalization of the R channel enhances the texture difference between mud (dark brown) and gravel (light gray);

[0205] ④ Input the reconstructed DINOv3 model, the Encoder enhances the mud texture features through the GTA module, and outputs the original segmentation mask. ;

[0206] ⑤ Boundary refinement: for Edge correction aligns the mud / stone dividing boundary with the actual edge;

[0207] ⑥ Small mud block completion: Three isolated mud block pixels with an area of ​​less than 5 pixels were detected. The texture features were matched to the mud blocks, and the completion was made to the mask. ;

[0208] ⑦ Mud content calculation: Number of background pixels Total number of pixels ; Number of pixels in mud Total number of pixels in the effective area mud content .

[0209] The implementation effect was verified: the mud content measured manually was 3.4%, and the model calculation error was only 0.03%; the traditional DINOv3 missed small mud lumps under the same lighting conditions, and the calculated mud content was 2.1%, with an error of 1.3%; the scene adaptive design of the present invention improved the segmentation accuracy from the traditional 82% to 99%, which fully meets the engineering detection accuracy requirements (error < 0.5%).

[0210] Please see Figure 4 As shown, the specific process for automatically switching the near-infrared moisture model according to the material type and outputting high-precision moisture content in step S5 is as follows:

[0211] Step S51, Near-infrared spectral acquisition and multi-dimensional preprocessing: Based on the near-infrared absorption characteristics of moisture in crushed stone / medium sand, three core characteristic wavelengths were selected (1300nm (first-order overtone absorption of water OH bonds), 1450nm (coupling absorption of crushed stone / medium sand matrix and moisture), and 1940nm (strong second-order overtone absorption of water OH bonds)). A narrow-band near-infrared sensor (bandwidth ±5nm) was used to acquire the spectral absorbance, and the raw absorbance sequence was output. ( (t is the data acquisition time).

[0212] Specifically, in step S51, after near-infrared spectral acquisition, environmental interference correction is required to improve acquisition accuracy.

[0213] For example, temperature correction: synchronously collect material temperature. Introducing a temperature compensation coefficient Correcting the absorbance shift due to temperature:

[0214] ;

[0215] In the formula, C, Fitting was achieved through offline experiments;

[0216] For example, dust / distance calibration: collecting the distance between the sensor and the material. Correcting the decrease in light absorption caused by material spreading thickness and dust blockage:

[0217] ;

[0218] In the formula, Standard distance, This is the attenuation coefficient.

[0219] Step S52, Intelligent Material Identification and Model Switching Trigger: Extract material texture features from RGB images (crushed stone: sharp edges and discrete particles; medium sand: fine particles and uniform texture), calculate the absorbance ratio of three feature wavelengths, construct a lightweight classifier, and output material type labels by fusing visual and spectral features;

[0220] The formula for calculating the absorbance ratio of the three characteristic wavelengths is:

[0221] ;

[0222] Build a lightweight classifier that combines visual and spectral features to output material type labels. ;

[0223] ;

[0224] in, This is the weight matrix. For bias, Using the Sigmoid activation function; recognition accuracy Once the identification is complete, the moisture calibration model for the corresponding material will be automatically triggered.

[0225] Step S53, Dual-calibration model construction: Divide the model into two sub-models according to the moisture content range. The model formula is as follows:

[0226] ;

[0227] In the formula, , , These are the interval weighting coefficients. , This is the offset. This refers to the moisture content of the crushed stone. It is a near-infrared characteristic wavelength. The near-infrared absorbance value at a certain wavelength after smoothing and noise reduction;

[0228] Specifically, the model is divided into two sub-models based on moisture content range:

[0229] Low moisture content range ( ): Linear weighted model, adapted to linear absorbance response under low moisture conditions;

[0230] High moisture content range ( ): Nonlinear weighted model, adapted to saturated absorption characteristics under high moisture conditions.

[0231] Due to the large specific surface area of ​​medium sand, its water adsorption characteristics differ significantly from those of crushed stone. A matrix compensation calibration model is designed as follows: This eliminates the interference of the matrix on moisture detection;

[0232] Step S54, Moisture Content Output and Accuracy Verification: Based on material type Call the corresponding model to output the initial moisture content The moving average of the predicted values ​​for 10 consecutive frames is taken to output the final moisture content. The specific formula is as follows: This reduces numerical fluctuations caused by instantaneous interference; samples are randomly selected from each batch of materials, and the model's predicted values ​​are compared with the actual values ​​measured by the drying method. If the error is greater than 0.5%, the model is automatically fine-tuned.

[0233] In practical implementation, taking the detection of moisture content in medium sand as an example, the parameters are set as follows:

[0234] Standard temperature Standard distance medium sand temperature compensation coefficient ℃, attenuation coefficient Medium sand calibration coefficient: , , , Matrix compensation coefficient Dynamic weight learning rate .

[0235] The specific execution process is as follows:

[0236] ① Collect near-infrared absorbance of medium sand: , , ;

[0237] ② Environmental correction:

[0238] Measured material temperature T=30℃, temperature correction: Similarly , ;

[0239] Actual sensor distance measurement Distance correction: Similarly , ;

[0240] ③After SG smoothing filtering:

[0241] , , ;

[0242] ④ Dual-modal material recognition:

[0243] Visual features Determined to be a medium sand texture, based on spectral characteristics. Matching sand features, output Automatically switch between medium sand models;

[0244] ⑤ Calculation of medium sand model:

[0245] Matrix absorbance ;

[0246] initial moisture content ;

[0247] ⑥ Moving average: The predicted value of 10 consecutive frames After averaging ;

[0248] ⑦ Accuracy verification: The measured moisture content by the drying method was 6.8%, and the model error was only 0.04%, which is far better than the traditional single model (error 0.5%+).

[0249] The experimental results are as follows: If medium sand is mistakenly input into the crushed stone model, the predicted value is 5.2%, with an error of 1.6%; while the fully automatic material identification + dedicated calibration model of this invention always has an error of <0.5%, meeting the high precision requirements of engineering testing; and after the model has been running for 3 months, through dynamic weight updates, the error is still stable at around 0.05%, with no drift phenomenon.

[0250] It is worth noting that the various units included in the above system embodiments are only divided according to functional logic, but are not limited to the above division, as long as the corresponding functions can be achieved; in addition, the specific names of each functional unit are only for easy differentiation and are not used to limit the scope of protection of the present invention.

[0251] Furthermore, those skilled in the art will understand that all or part of the steps in the methods of the above embodiments can be implemented by a program instructing related hardware, and the corresponding program can be stored in a computer-readable storage medium.

[0252] The preferred embodiments of the present invention disclosed above are merely illustrative of the invention. These preferred embodiments do not exhaustively describe all details, nor do they limit the invention to the specific implementations described. Clearly, many modifications and variations can be made based on the content of this specification. This specification selects and specifically describes these embodiments to better explain the principles and practical applications of the invention, thereby enabling those skilled in the art to better understand and utilize the invention. The invention is limited only by the claims and their full scope and equivalents.

Claims

1. A machine vision-based aggregate intelligent detection method, characterized in that, Includes the following steps: Step S1, Multi-sensor synchronous acquisition: Simultaneous acquisition of aggregate images and spectral information using a 3D line laser camera, a high-resolution RGB line array camera, and a near-infrared spectral sensor; Step S2, Preprocessing of gravel point cloud and conversion of depth map: Denoise, segment, and project the 3D point cloud to generate a standardized depth map; Step S3: Intelligent analysis of gravel volume and shape based on CNN: Output the volume of a single stone, equivalent particle size, length / width / height through a deep graph convolutional neural network to complete the gradation statistics and needle-like / flaky determination; Step S4: Semantic segmentation of mud based on DINOv3 visual Transformer: Pixel-level segmentation of RGB image to accurately identify mud block regions and calculate mud content; Step S5, Dual-calibration moisture prediction model for crushed stone / medium sand: Automatically switches to the near-infrared moisture model according to the material type, and outputs high-precision moisture content; Step S6, Data Fusion and Report Output: Integrate all indicators to generate a traceable testing report. 2.The machine vision-based aggregate intelligent detection method according to claim 1, characterized in that, In step S1, the specific process of multi-sensor synchronous data acquisition is as follows: Step S11: Using the hard-triggered synchronization signal as a unified timing reference, the 3D line laser camera, high-resolution RGB line array camera, and near-infrared spectral sensor are coaxially deployed along the material conveying direction. Step S12: When the aggregate on the conveyor belt passes through the unified detection area, the photoelectric switch triggers a pulse signal, simultaneously starting 3D point cloud scanning, RGB linear array imaging and near-infrared spectrum acquisition; Step S13: During the acquisition process, the 3D camera and the RGB camera maintain the same line frequency and pixel resolution, and the near-infrared sensor maintains the same sampling rate to ensure that the data of the three devices correspond one-to-one in spatial location and time series.

3. The intelligent material detection method based on machine vision according to claim 1, characterized in that, In step S2, the specific process of preprocessing the gravel point cloud and converting the depth map is as follows: Step S21, Raw Point Cloud Acquisition and Format Standardization: Raw point cloud data of a single layer of paved gravel on the conveyor belt is acquired using a line laser 3D camera. The output format is PCD, and the number of point clouds per frame is controlled within [specific parameters]. There are points, and the coordinates of each point cloud are defined as follows: In the formula, In three-dimensional space coordinates, These are the RGB color values ​​collected synchronously. Step S22, Multi-dimensional hierarchical filtering and denoising: Traverse each point cloud, count the number of points in its neighborhood, and remove isolated noise points with less than 5 neighborhood points. Step S23, Adaptive Euclidean Clustering for Individual Crushed Stones: Calculate the Z-axis variance of the filtered point cloud to characterize the undulation of the aggregate surface, and set an adaptive distance threshold to perform Euclidean clustering; Step S24, Point Cloud Projection and Depth Map Generation: Project the single gravel point cloud orthogonally along the Z-axis onto the XOY plane, establish the mapping relationship between the pixel coordinate system and the spatial coordinate system, and take the maximum Z value of the corresponding spatial position as the depth value for each pixel in the pixel coordinate system. Step S25, Depth Map Normalization and Enhancement: Map depth values ​​to... The grayscale range is used, and the Sobel operator is used to extract the edges of the depth map, which are then superimposed onto the normalized depth map.

4. The intelligent material detection method based on machine vision according to claim 3, characterized in that, In step S24, the point cloud of a single gravel particle is orthogonally projected along the Z-axis onto the XOY plane to establish a pixel coordinate system. With spatial coordinate system The mapping relationship is as follows: , ; In the formula, The depth map resolution is in pixels. , The minimum X / Y coordinates of the point cloud of a single gravel; For each pixel in the pixel coordinate system The maximum Z value at the corresponding spatial location is taken as the depth value, and the specific formula is as follows: ; The holes in the depth map are repaired using bilateral filtering interpolation, with the specific formula as follows: ; In the formula, For neighborhood windows, For depth similarity weights, This represents the spatial distance weight.

5. The intelligent material detection method based on machine vision according to claim 1, characterized in that, In step S3, the specific process of intelligent analysis of gravel volume and shape based on CNN is as follows: Step S31, Depth Map Preprocessing and Dataset Construction: The enhanced depth map generated in step S2 is preprocessed... Scaling to uniform Pixels were collected and mean-variance normalized. 100,000 sets of depth maps and measured parameter samples were acquired, covering [the area / region]. For the entire particle size range, the depth map of small-sized crushed stone is stretched to scale, and the large-sized particles are locally trimmed and spliced. Step S32: Construction of a multi-scale feature fusion CNN network: Dynamic weights are assigned to the feature maps of the three branches. Pooling is used to reduce the dimensionality of the fused feature map to a 1024-dimensional feature vector. Four parallel fully connected layers are designed to output the volume respectively. ,length ,width and thickness ; Step S33, Physical constraint loss function training: Design a physical constraint hybrid loss function to balance numerical accuracy and physical rationality; Step S34, Model Inference and Dynamic Calibration: Input the preprocessed depth map into the trained MSF-CNN and output the initial prediction value. , , , The calibration coefficients are preset according to the particle size range; Step S35, Grading Statistics and Needle-like / Flake-like Determination: Calculate the equivalent particle size based on the calibrated volume, count the proportion of crushed stone in each particle size range, and generate a gradation curve.

6. The intelligent material detection method based on machine vision according to claim 5, characterized in that, In step S33, the physical constraint hybrid loss function is as follows: ; In the formula, For mean square error loss, For physical volume constraint loss, Loss due to size sorting constraints; The formula for calculating the mean squared error loss is as follows: ; The formula for calculating the physical volume constraint loss is as follows: ; The formula for calculating the size sorting constraint loss is as follows: ; In the formula, , , , For the first Predicted values ​​for volume, length, width, and thickness. , , , For the first Measured values ​​of volume, length, width, and thickness.

7. The intelligent material detection method based on machine vision according to claim 5, characterized in that, In step S35, the formula for calculating the equivalent particle size is as follows: The percentage of crushed stone in each particle size range is calculated based on the sieve aperture size to generate a gradation curve. In the formula, Screen aperture size The pass rate For equivalent particle size smaller than The amount of gravel.

8. The intelligent material detection method based on machine vision according to claim 1, characterized in that, In step S4, the specific process of mud-containing semantic segmentation based on DINOv3 visual Transformer is as follows: Step S41, RGB image acquisition and scene-based preprocessing: Use a linear RGB camera to acquire images of gravel and mud, and simultaneously record the light intensity at the time of acquisition for adaptive light correction. Step S42, Adaptive DINOv3 Model Reconstruction for Aggregate Scene: Insert an aggregate texture attention module into the Transformer Encoder of DINOv3. This aggregate texture attention module strengthens the mud feature weight by learning the texture differences between mud and gravel, simplifies the multi-class output of general segmentation to 3 classes, and adds a boundary refinement head to perform pixel-level correction on the edges of the segmentation mask. Step S43, Pixel-level segmentation and accurate calculation of mud content: Input the preprocessed enhanced image into the trained DINOv3 model, and output a pixel-level segmentation mask. Perform connected component analysis on the segmentation mask and calculate the area of ​​each connected component. The minimum mud block area threshold is used to fill in the small mud blocks, and the mud content is calculated by counting the number of mud block pixels in the effective area and the total number of pixels.

9. The intelligent material detection method based on machine vision according to claim 8, characterized in that, In step S43, a minimum mud block area threshold is set. Regarding area However, isolated pixels that match the texture features of the mud are identified as small mud lumps and added to the mask. The formula is as follows: ; Count the number of mud block pixels within the valid area Total number of pixels Introducing pixel-to-actual area calibration coefficient Convert pixel percentage to actual area percentage: ; In the formula, The number of background pixels. This represents the final mud content.

10. The intelligent material detection method based on machine vision according to claim 1, characterized in that, In step S5, the specific process for automatically switching the near-infrared moisture model according to the material type and outputting high-precision moisture content is as follows: Step S51, Near-infrared spectral acquisition and multi-dimensional preprocessing: Based on the near-infrared absorption characteristics of moisture in crushed stone / medium sand, three core characteristic wavelengths are selected, and a narrow-band near-infrared sensor (bandwidth ±5nm) is used to acquire spectral absorbance and output the original absorbance sequence. Step S52, Intelligent Material Identification and Model Switching Trigger: Extract material texture features from RGB images, calculate the absorbance ratio of three feature wavelengths, construct a lightweight classifier, and output material type labels by fusing visual and spectral features; Step S53, Dual-calibration model construction: Divide the model into two sub-models according to the moisture content range. The model formula is as follows: ; In the formula, , , These are the interval weighting coefficients. , This is the offset. This refers to the moisture content of the crushed stone. It is a near-infrared characteristic wavelength. The near-infrared absorbance value at a certain wavelength after smoothing and noise reduction; Step S54, Moisture Content Output and Accuracy Verification: Based on the material type, call the corresponding model to output the initial moisture content, take the moving average of the predicted values ​​for 10 consecutive frames, and output the final moisture content.