Lightweight recycled aggregate crushing value multi-modal real-time detection method and system
Patent Information
- Application Number
- CN202610808899.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-05
- Publication Date
- 2026-09-01
AI Technical Summary
[0004]本发明的目的在于提供一种面向移动端部署的基于多模态决策级融合的再生骨料压碎值智能检测方法及系统,以克服传统破坏性试验流程繁琐、时效性差,以及现有单模态图像检测方法无法全面表征骨料细微观物理属性的局限性
(1)本发明通过图像模态与物理模态的决策级堆叠融合,在独立测试集上取得R2达0.9601、MAPE为4.32%的预测性能,显著优于单一模态模型,且克服了特征级融合中的跨模态优化冲突问题;
Smart Images

Figure CN122676294A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of intelligent testing technology for building materials, and in particular to a method and system for real-time multimodal detection of crushing value of lightweight recycled aggregate. Background Technology
[0002] In recent years, intelligent detection methods based on image recognition technology have provided new insights into the performance evaluation of recycled aggregates. Existing research has constructed recycled aggregate image datasets and utilized deep learning models to extract macroscopic morphological features of the aggregates, achieving rapid prediction of crushing values and overcoming, to some extent, the inefficiency of traditional manual detection. However, these methods are essentially limited to a single image modality. In fact, the macroscopic mechanical behavior of recycled aggregates is the result of cross-scale coupling. While two-dimensional images can capture morphological information such as particle shape, angularity, and surface texture, they cannot directly reflect the material's microscopic properties, such as the overall density and pore connectivity of the aggregate, as measured by apparent density and water absorption rate. Existing research shows that apparent density is significantly negatively correlated with crushing values, while water absorption rate is strongly positively correlated with crushing values. Because a single image modality cannot perceive these key physical properties, it often exhibits significant prediction bias when dealing with complex aggregate samples with high internal porosity and microcrack development. Furthermore, practical engineering applications face the real challenge of "modal absence." At construction sites, sometimes only aggregate images can be captured, making it impossible to measure physical parameters in a timely manner. In laboratory environments, sometimes only physical parameter records exist without corresponding image data. Existing single-modal models cannot flexibly adapt to scenarios with dynamically changing data availability, significantly limiting the implementation and promotion of intelligent detection technology on engineering sites. Simultaneously, existing deep learning models have a large number of parameters and high computational complexity, making real-time inference difficult on resource-constrained mobile or edge devices. Moreover, model iteration and updates typically require redeployment and installation, resulting in low management efficiency.
[0003] Therefore, there is an urgent need for a high-precision intelligent detection method and system for the crushing value of recycled aggregate that can integrate multi-source heterogeneous information, flexibly respond to modality-deficient scenarios, and be suitable for mobile edge computing environments. Summary of the Invention
[0004] The purpose of this invention is to provide an intelligent detection method and system for crushing value of recycled aggregate based on multimodal decision-level fusion for mobile deployment, so as to overcome the limitations of traditional destructive testing process, which is cumbersome and has poor timeliness, and existing single-modal image detection methods, which cannot fully characterize the fine physical properties of aggregate.
[0005] To achieve the above objectives, this invention provides a multimodal real-time detection method for the crushing value of lightweight recycled aggregate, comprising the following steps: S1. Data Acquisition: Collect images of recycled aggregates with different types and gradations of recycled aggregate parent materials, and conduct standard crushing value tests on them to measure the precise crushing value corresponding to each recycled aggregate image. Assign a corresponding precise crushing value label to each recycled aggregate image to form a benchmark image dataset; collect physical property data of recycled aggregates that correspond one-to-one with the recycled aggregate image samples, including data collected in the laboratory and supplemented by data extracted from published literature, and construct a structured dataset together with the benchmark image dataset; S2. Data Augmentation: The baseline image dataset is subjected to two-stage data augmentation processing. The first stage is basic data augmentation, which introduces image transformation operations, and each original image is processed by image transformation operations 4 times. The second stage is random occlusion augmentation, which randomly generates light debris occlusion regions on the images obtained in the first stage. Finally, the data-augmented multimodal dataset is obtained. S3. Single-modal basis learner training: Construct image modal basis learners and physical modal basis learners, and train them separately; S4. Dual-path cross-validation and prediction generation: Data is partitioned and an asymmetric dual-path cross-validation mechanism is constructed. Five-fold cross-validation is performed on the image modality base learner and the physical modality base learner respectively to generate a sequence of prediction values that are completely independent at the sample level. S5. Construction of Multimodal Decision-Level Fusion Model: A decision-level fusion strategy is adopted, and a stacked generalization method is introduced. The meta-learner performs nonlinear combination of the output of the base learner, and finally outputs the prediction result of the sample crushing value.
[0006] Furthermore, in step S1, the types of recycled aggregate parent materials include natural aggregates, recycled concrete aggregates, recycled brick aggregates, and recycled mixtures. The characteristic fields of the structured dataset include longitude, latitude, pressure load, crushing value, apparent density, and water absorption rate.
[0007] Furthermore, in step S2, the basic data enhancement specifically involves introducing image transformation operations including random rotation, random flipping, random cropping, brightness jitter, saturation jitter, contrast jitter, hue jitter, random sharpness adjustment, RGB offset, and Gaussian noise. Each original image is processed by the image transformation operations four times. The random occlusion enhancement specifically involves randomly generating lightweight debris occlusion areas on the image obtained in the first stage; the lightweight debris includes leaves, pieces of paper, plastic, and film, and their positions, sizes, and angles in the image are all generated using a random generation strategy.
[0008] Furthermore, step S3 specifically includes: S31. The image modality base learner uses a lightweight convolutional neural network model, including the ShuffleNet V2 lightweight convolutional neural network, as the backbone network. Image modality base learner training: The augmented recycled aggregate image dataset was uniformly adjusted to a size of 224×224 pixels. The model parameters were initialized using weights pre-trained on the ImageNet dataset, and fine-tuned by transfer learning on the recycled aggregate image dataset. The Adam optimizer was used with an initial learning rate of 0.001 and a batch size of 512. The training lasted for 100 epochs, and an early stopping mechanism was introduced to automatically terminate training when the validation set loss did not decrease within 20 consecutive epochs.
[0009] S32. The physical modality base learner adopts a machine learning model including a random forest ensemble learning model. The input features include numerical variables and categorical variables, and the crushing value of recycled aggregate is used as the target variable for model output. The numerical variables include longitude, latitude, pressure load, apparent density and water absorption rate, and the categorical variables are the types of recycled aggregate parent materials. Physical modality base learner training: Numerical features are standardized, and categorical features are encoded using one-hot encoding. Key hyperparameters of the random forest are tuned using a grid search combined with five-fold cross-validation. The hyperparameter search range includes: the number of decision trees ({50, 100, 200, 300}), the maximum depth ({5, 10, 15, None}), and the minimum number of split samples ({2, 5, 10}). Using root mean square error as the optimization objective, the optimal hyperparameter combination is selected after 20 random searches using RandomizedSearchCV. The model's generalization ability and stability are then evaluated on an independent test set.
[0010] Furthermore, step S4 specifically includes: S41. Data partitioning: The physical property datasets that correspond one-to-one with the recycled aggregate image samples collected by the laboratory in step S1 are denoted as dataset A, and the physical property datasets that are supplemented by extracting from published literature are denoted as dataset B. Based on the types of recycled aggregate parent materials, stratified random sampling was performed on dataset A, and 20% of the samples were divided into independent test sets, while the remaining 80% of the samples were recorded as cross-validation data. S42, Physical Modality Cross-Validation Path: The cross-validation data is randomly divided into 5 equal-sized subsets; in each fold cross-validation, 4 subsets are selected and together with all samples from dataset B to form the training set for the current fold, and the remaining 1 subset forms the validation set; the machine learning model after hyperparameter tuning in step S32 is used to train on the training set for the current fold, and a physical modality prediction value is generated for each set of samples in the validation set for the current fold. S43, Image Modality Cross-Validation Path: The cross-validation data is divided into 5 folds according to the same subset partitioning method as the physical modality path. In each fold of cross-validation, all recycled aggregate images associated with 4 subsets are selected to form the training set for the current fold, and all images associated with the remaining 1 subset are selected to form the validation set. The lightweight convolutional neural network fine-tuned by transfer learning in step S31 is used to train on the training set for the current fold, and an image modality prediction value is generated for each set of samples in the validation set for the current fold. For the case where a single sample contains multiple images, the average of the prediction values of all images corresponding to the same validation sample is taken as the image modality prediction value of that sample. S44. Predicted value concatenation and meta-feature set construction: The physical modality predicted value and the image modality predicted value are paired and concatenated according to samples to form a two-dimensional prediction feature matrix; Each row of the two-dimensional prediction feature matrix corresponds to a sample, containing two input features: the image modality prediction value and the physical modality prediction value of the sample. The two-dimensional prediction feature matrix is used as the input variable, and the true crushing value of the sample is used as the target variable to form a meta-feature set for training the meta-learner. Furthermore, step S5 specifically includes: Decision-level fusion is performed at the output level of the model prediction, specifically by introducing a meta-learner to learn the mapping relationship between the base model predictions and the true target. The base learners are the image model and physical feature model trained in step S3, denoted as follows: and For each sample in the training set in steps S42 and S43, two initial prediction values are generated by the base learner. and , and the true value Together they constitute a new meta-feature set , For the samples in the training set, The total number of samples in the training set; The dataset is divided into a meta-training set and a meta-validation set for training and tuning the meta-learner. During training, the meta-learner aims to minimize the mean squared error and learns the combined weights and interaction effects of the predictions from the two base models. Finally, for new samples, the predicted values are fused. The calculation process is shown in the following formula: ; ; ; in, For a trained meta-learner.
[0011] The present invention also provides a lightweight recycled aggregate crushing value multimodal real-time detection system, comprising: a mobile terminal acquisition and interaction module, a cloud model service module, and a mobile terminal application terminal; The mobile terminal data acquisition and interaction module is built on the Android platform and installed on the handheld mobile devices of engineers at the construction site. It is responsible for data acquisition, local preprocessing, result display and intelligent classification. The cloud model service module is deployed on a cloud server and is responsible for the storage, management and inference calculation of the core inference model; the model weight file of the image modality base learner trained in step S3, the model file of the physical modality base learner, and the meta-learner model file trained in step S5 are uniformly deployed on the cloud server. Furthermore, the mobile terminal data acquisition and interaction module specifically includes the following functions: Automated collection of geographic location information: Integrating an open-source geographic information system engine, it allows field engineers to accurately anchor the coordinates of recycled aggregate production sites through interactive maps, realize the automated mapping and recording of longitude and latitude parameters, and provide spatial location feature input for physical modality base learners; Image acquisition and local preprocessing: The mobile device's camera is used to capture images of recycled aggregates; before the images are uploaded to the cloud server, the mobile device calls the underlying image processing library to perform local preprocessing and compression on the high-resolution images. The preprocessing operations include: downsampling the images to a suitable resolution using bilinear interpolation, and converting the images into lightweight data streams using 50% quality JPEG encoding. Multimodal data upload: The pre-processed image data, geographic location information, and on-site physical parameters are encapsulated into a unified inference request and sent to the cloud model service module via wireless network; Results reception and intelligent classification display: Receive the crushing value prediction results returned from the cloud and display them visually on the mobile interface; The mobile terminal has a built-in intelligent classification module, which automatically classifies the recycled aggregate into three levels: Class I, Class II and Class III, based on the predicted crushing value, which correspond to the aggregate quality grades applicable to different types of engineering projects.
[0012] Furthermore, the cloud-based model service module specifically includes the following functions: Model hot update mechanism: A model hot update module is built in the cloud management console; project managers do not need to interrupt the server operation, but only need to dynamically switch the currently effective model version in the drop-down menu of the visual interface to achieve seamless and smooth upgrades of different sections' adapted models or algorithm iteration versions; Online Inference Service: The cloud server receives inference requests uploaded by mobile devices and automatically calls the corresponding model for inference calculation based on the modality type identifier contained in the request. When the request contains both image data and physical parameters, the complete multimodal decision-level fusion model is invoked, and the final predicted value is output by the image modality base learner, physical modality base learner, and meta-learner in sequence. When the request contains only image data, the image modality base learner is automatically called for inference. When the request contains only physical parameters, the physical modality base learner is automatically called for inference. After the inference is completed, the prediction result is returned to the mobile device in real time.
[0013] Furthermore, the workflow of the detection system is as follows: Training phase: Import the augmented multimodal dataset obtained in step S2 into the cloud server, complete the training of the single-modal base learner and the multimodal decision-level fusion model according to the process of steps S3 to S5, and deploy all trained model files to the cloud model service module. Testing Phase: Using a mobile device, the aggregate origin is first located via an interactive map to obtain latitude and longitude information; then, images of the recycled aggregate batch to be tested are captured, and physical parameters such as apparent density and water absorption rate of the batch are input when conditions permit; the mobile device performs local preprocessing and compression of the images, encapsulates the multimodal data into an inference request, and uploads it to the cloud; the cloud automatically selects the corresponding model for inference based on modal integrity, and returns the crushing value prediction results and aggregate quality grade to the mobile device for display; Model update phase: When it is necessary to fine-tune the model or iterate the algorithm for the aggregate characteristics of a specific mining area, the model version can be dynamically switched through the hot update module in the cloud management console. All mobile devices can use the latest model synchronously without reinstalling the application.
[0014] The present invention employs the above-mentioned multimodal real-time detection method and system for the crushing value of lightweight recycled aggregate, and its beneficial effects are as follows: (1) This invention achieves R on an independent test set by stacking and fusing decision-level data of image modalities and physical modalities. 2 With a prediction performance of 0.9601 and MAPE of 4.32%, it significantly outperforms the single-modal model and overcomes the cross-modal optimization conflict problem in feature-level fusion. (2) The system of the present invention can automatically select fusion prediction, pure image prediction or pure physical parameter prediction according to the modal integrity of the uploaded data, which is suitable for complex scenarios such as only being able to acquire images or only being able to measure physical parameters at the construction site; the mobile terminal reduces the transmission volume by an order of magnitude through local image preprocessing, meeting the low latency and low power consumption edge computing requirements of "instant measurement"; (3) This invention supports dynamic switching of model versions in the cloud management console, and mobile devices can be updated synchronously without reinstalling the application, realizing seamless upgrades and efficient operation and maintenance of model iteration. Attached Figure Description
[0015] Figure 1 This is a flowchart of a method for real-time multimodal detection of crushing value of lightweight recycled aggregate according to an embodiment of the present invention; Figure 2 This is a schematic diagram of the two-stage enhancement process of a real-time multimodal detection method for the crushing value of lightweight recycled aggregate according to an embodiment of the present invention; Figure 3 This is a flowchart of the asymmetric dual-path cross-validation mechanism of a multimodal real-time detection method for crushing value of lightweight recycled aggregate according to an embodiment of the present invention; Figure 4 This is a comparison chart of the prediction results and the actual values of the multimodal decision-level fusion model of a real-time multimodal detection method for crushing value of lightweight recycled aggregate according to an embodiment of the present invention; wherein, (a) is the prediction result of the random forest model, (b) is the prediction result of the ShuffleNet V2 model, and (c) is the prediction result of the decision-level fusion model. Figure 5 This is a Grad-CAM heatmap generated by a trained ShuffleNet V2 model of a real-time multimodal detection method for crushing values of lightweight recycled aggregates according to an embodiment of the present invention. Figure 6 This is a regenerated heatmap of a lightweight recycled aggregate crushing value multimodal real-time detection method according to an embodiment of the present invention, which involves occluding the original image; wherein, (a) is a single-region local occlusion, (b) is a multi-region dispersed occlusion, (c) is a vertical strip occlusion, and (d) is a single-region local occlusion. Figure 7 This is the SHAP value of the crushing value prediction result of each input feature in all samples of a multimodal real-time detection method for crushing value of lightweight recycled aggregate using a random forest model, according to an embodiment of the present invention. Figure 8 This is a SHAP feature dependency graph of the random forest model on water absorption rate in a multimodal real-time detection method for crushing value of lightweight recycled aggregate according to an embodiment of the present invention. Figure 9 This is a structural diagram of a multimodal real-time detection system for the crushing value of lightweight recycled aggregate according to an embodiment of the present invention; Figure 10 This is a schematic diagram of the cloud model service module of a real-time multimodal detection system for crushing value of lightweight recycled aggregate according to an embodiment of the present invention; Figure 11 This is a schematic diagram of the mobile terminal acquisition and interaction module of a real-time multimodal detection system for crushing value of lightweight recycled aggregate according to an embodiment of the present invention; wherein, (a) is data acquisition, (b) is data processing, and (c) is result display; Detailed Implementation
[0016] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to specific embodiments and the accompanying drawings. It should be understood that these descriptions are merely exemplary and not intended to limit the scope of the invention. Furthermore, descriptions of well-known structures and techniques are omitted in the following description to avoid unnecessarily obscuring the concept of the invention.
[0017] Example 1, such as Figure 1 As shown in the figure, an embodiment of the present invention provides a multimodal real-time detection method for the crushing value of lightweight recycled aggregate, comprising the following steps: S1. Data Acquisition: Collect images of recycled aggregates of different types and gradations (in this embodiment, the types include natural aggregates, recycled concrete aggregates, recycled brick aggregates, and recycled mixtures; actual engineering scenarios may not be limited to the above four types). Simultaneously, conduct standard crushing value tests on the recycled aggregates to measure the precise crushing value corresponding to each recycled aggregate image. Assign a corresponding precise crushing value label to each recycled aggregate image to form a benchmark image dataset (the differences in construction waste in different regions are significant, therefore this step is a necessary condition for the realization of this invention).
[0018] We collected physical property data of recycled aggregates that corresponded one-to-one with the image samples of recycled aggregates. This included data collected in the laboratory and supplementary data extracted from published literature. Together with the baseline image dataset, we constructed a structured dataset.
[0019] The dataset in this embodiment includes six feature fields: longitude, latitude, pressure load, crushing value, apparent density, and water absorption rate. It covers recycled aggregate samples from multiple provinces and cities in eastern and central my country (longitude: 104.04°E~126.50°E, latitude: 22.28°N~43.81°N). Among them, the pressure load is mainly concentrated at 200kN, the crushing value varies from 4.20% to 42.30%, and the apparent density ranges from 1.98g / cm³. 3 ~2.86g / cm 3 The water absorption rate ranged from 0.40% to 24.50%, with all indicators exhibiting significant fluctuations, providing rich data support for constructing a multimodal prediction model. Detailed statistical information is shown in Table 1. Table 1 Statistical information of detailed data
[0020] S2. Data Augmentation: To meet the practical engineering needs of "instant measurement" on mobile devices and enable the model to adapt to the complex and ever-changing environmental conditions of construction sites, a two-stage data augmentation process is performed on the baseline image dataset from step S1, tailored to mobile application scenarios (e.g., ...). Figure 2 (As shown).
[0021] S21. Basic Data Enhancement: Introducing image transformation operations including random rotation, random flipping, random cropping, brightness jitter, saturation jitter, contrast jitter, hue jitter, random sharpness adjustment, RGB offset, and Gaussian noise to simulate the complex and varied lighting conditions, shooting angles, and different aggregate accumulation patterns at the construction site; each original image is processed 4 times by the image transformation operations.
[0022] S22, Random Occlusion Enhancement: Randomly generate lightweight debris occlusion areas on the enhanced image obtained in the first stage to simulate irregular debris occlusion scenarios on the aggregate surface in actual engineering.
[0023] Lightweight debris includes four typical types of construction site disturbances: leaves, paper scraps, plastic, and film. Their positions, sizes, and angles in the images are all generated using a randomization strategy.
[0024] After data augmentation, the number of images increased fivefold, forming a multimodal image dataset, which provides a sufficient sample base for the training of subsequent lightweight image models.
[0025] S3. Single-modal base learner training: To meet the stringent requirements of low latency and low power consumption for mobile deployment, while ensuring high-precision prediction of the crushing value of recycled aggregate, this section constructs independent base learners for the image modality and the physical modality, respectively. The two base learners are trained and optimized independently in their respective data domains.
[0026] S31. Image Modality Base Learner Construction: The image modality base learner uses a lightweight convolutional neural network model as the backbone network. In this embodiment, ShuffleNet V2 lightweight convolutional neural network is used as an example. ShuffleNet V2 uses actual running speed as the direct optimization indicator. Through the channel segmentation mechanism, the input features are divided into two branches: identity mapping and convolution processing. While ensuring feature reuse and model capacity, it significantly reduces hardware execution time. Its parameter size is only 2.3M, which meets the strict requirements of mobile edge inference for model size and computational efficiency.
[0027] Image modality base learner training: The regenerated aggregate image dataset after data augmentation in step S2 was uniformly adjusted to a size of 224×224 pixels. The model parameters were initialized using weights pre-trained on the ImageNet dataset, and fine-tuned by transfer learning on the regenerated aggregate image dataset. The Adam optimizer was used with an initial learning rate of 0.001 and a batch size of 512. The training lasted for 100 epochs, and an early stopping mechanism was introduced. The training was automatically terminated when the validation set loss did not decrease within 20 consecutive epochs to prevent overfitting.
[0028] After training, ShuffleNet V2 achieved a determination coefficient R on the image test set. 2 The predictive performance of 0.9435, RMSE of 0.0118, MAE of 0.0086, and MAPE of 5.50% indicates that the lightweight image model has a robust fitting ability for the apparent texture features of aggregates.
[0029] S32. Construction of Physical Modal Base Learner: A machine learning model is adopted. In this embodiment, a random forest ensemble learning model is used. In the context of research on small sample structured datasets, random forest can effectively suppress the risk of overfitting through a parallel variance reduction strategy of bootstrapping and random feature selection. It is suitable for establishing a nonlinear mapping relationship between the physical parameters of recycled aggregate and the crushing value.
[0030] The input features of the physical modal basis learner include numerical variables and categorical variables. The numerical variables include longitude, latitude, pressure load, apparent density, and water absorption rate, while the categorical variables are the types of recycled aggregate parent materials. The numerical features are standardized, and the categorical features are encoded using unique thermal encoding. The crushing value of the recycled aggregate is used as the target variable for model output.
[0031] Physical modality base learner training: To fully utilize the model's performance, a grid search combined with five-fold cross-validation was used to fine-tune the key hyperparameters of the random forest. The hyperparameter search range included: the number of decision trees (selected from {50, 100, 200, 300}), the maximum depth (selected from {5, 10, 15, None}), and the minimum number of split samples (selected from {2, 5, 10}). Using root mean square error as the optimization objective, the optimal hyperparameter combination was selected after 20 random searches using RandomizedSearchCV, and the model's generalization ability and stability were evaluated on an independent test set.
[0032] After training and optimization, the random forest model achieved a coefficient of determination R on the physical feature test set. 2The predictive performance was 0.9476, with a root mean square error (RMSE) of 1.1024, a mean absolute error (MAE) of 0.8774, and a mean absolute percentage error (MAPE) of 6.33%. These results indicate that physical parameters such as apparent density and water absorption can effectively explain the main variations in aggregate crushing values, and the random forest model demonstrates strong generalization ability and fitting ability for complex nonlinear relationships.
[0033] S4. Dual-path cross-validation and prediction generation: To generate a high-quality prediction feature matrix for training the meta-learner, while ensuring an unbiased estimate of the generalization ability of the fusion model, this step constructs an asymmetric dual-path cross-validation mechanism (e.g., ...). Figure 3 As shown in the diagram, five-fold cross-validation is performed on both the image modality base learner and the physical modality base learner to generate completely independent prediction sequences at the sample level. The specific steps are as follows: S41. Data partitioning: The physical property datasets that correspond one-to-one with the recycled aggregate image samples collected in the laboratory in step S1 are denoted as dataset A, and the physical property datasets extracted from published literature are denoted as dataset B. Dataset B and dataset A are completely independent at the sample level, but they have statistical homogeneity in the crushing value index and can be used together for training and validation of the physical modality base learner.
[0034] Based on the types of recycled aggregate parent materials, stratified random sampling was performed on dataset A. 20% of the samples were divided into an independent test set. This independent test set was completely isolated in all subsequent model training, cross-validation, and meta-learner training stages and was used only for the generalization performance evaluation of the final fusion model. The remaining 80% of the samples were recorded as cross-validation data, which served as the data basis for cross-validation of the single-modal model.
[0035] S42, Physical Modality Cross-Validation Path: The cross-validation data is randomly divided into 5 equal-sized subsets; in each fold cross-validation, 4 subsets are selected and together with all samples from dataset B to form the training set for the current fold, and the remaining 1 subset forms the validation set; the random forest model after hyperparameter tuning in step S32 is used to train on the training set for the current fold, and a physical modality prediction value is generated for each set of samples in the validation set for the current fold.
[0036] Since the predictions of physical modes on each fold validation set are generated by models that were not trained on that fold, all physical mode predictions are unbiased estimates.
[0037] S43, Image Modality Cross-Validation Path: The cross-validation data is divided into 5 folds according to the same subset partitioning method as the physical modality path. In each fold of cross-validation, all recycled aggregate images associated with 4 subsets are selected to form the training set for the current fold, and all images associated with the remaining 1 subset are selected to form the validation set. The ShuffleNet V2 lightweight convolutional neural network, fine-tuned by transfer learning in step S31, is used to train on the training set for the current fold and generates an image modality prediction value for each set of samples in the validation set for the current fold. For cases where a single sample contains multiple images, the average of the prediction values of all images corresponding to the same validation sample is taken as the image modality prediction value of that sample.
[0038] Since the predictions on each fold validation set are generated by a model that did not participate in the training of that fold, the unbiasedness of the image modality predictions is ensured.
[0039] S44. Predicted Value Concatenation and Meta-Feature Set Construction: After all 5-fold cross-validation is completed, the physical modality predictions and image modality predictions are paired and concatenated on a sample-by-sample basis to form a two-dimensional prediction feature matrix. Each row of this two-dimensional prediction feature matrix corresponds to a sample, containing two input features: the image modality prediction and the physical modality prediction. Using this two-dimensional prediction feature matrix as the input variable and the true crushing value of the sample as the target variable, a meta-feature set is constructed for training the meta-learner. Since dataset A contains a total of 46 cross-validation samples, the meta-feature set contains 46 samples, each containing two predicted features and a corresponding true crushing value label.
[0040] This dual-path cross-validation mechanism ensures that the training data of the meta-learner and the final performance evaluation data are completely isolated at the sample level. Furthermore, the two paths adopt the same cross-validation subset partitioning method, which guarantees the strict alignment of the two modality predictions in terms of sample labels. This lays a solid data foundation for the training of the decision-level fusion model based on the stacked generalization strategy in the subsequent step S5.
[0041] S5. Construction of a Multimodal Decision-Level Fusion Model: Single-modal models still have inherent limitations when dealing with complex engineering scenarios. Image models struggle to perceive the fine physical properties of aggregates, while machine learning models based on physical features cannot utilize rich visual morphological information. To construct a more comprehensive and robust crushing value prediction system, this invention employs a decision-level fusion strategy, introducing a stacked generalization method. This involves nonlinearly combining the outputs of the base model through a meta-learner, thereby fully exploring the complementary value of cross-modal information.
[0042] Decision-level fusion occurs at the output level of model predictions, and its core advantages lie in maintaining the independence, training flexibility, and interpretability of sub-models. Specifically, it introduces a meta-learner to learn the mapping relationship between the base model's predicted values and the true target, thereby capturing the complementary patterns of prediction errors from different modalities and overcoming the limitation of linear weighting in modeling complex interactions.
[0043] Wherein, the base learners are the image model and physical feature model trained in step S3, denoted as and respectively. and For each sample in the training set in steps S42 and S43, two initial prediction values are generated by the base learner. and , and the true value Together they constitute a new meta-feature set , For the samples in the training set, The total number of samples in the training set; The dataset is divided into a meta-training set and a meta-validation set for training and tuning the meta-learner. During training, the meta-learner aims to minimize the mean squared error and learns the combined weights and interaction effects of the predictions from the two base models. Finally, for new samples, the predicted values are fused. The calculation process is shown in the following formula: ; ; ; in, For a trained meta-learner.
[0044] It should be noted that stacked generalization is essentially an implementation of multimodal fusion. It integrates high-level semantic information from different modal models through a meta-learner, which retains the modular advantages of decision-level fusion and enhances the nonlinear expressive power of the fusion model. Therefore, it can be regarded as a "multimodal decision-level fusion model based on stacked generalization".
[0045] Model prediction results are as follows Figure 4 As shown, the fusion model using random forest as the meta-learner achieved the best prediction performance. The evaluation metrics of this fusion model on completely isolated independent test sets are as follows: Coefficient of Determination R0 2 The coefficient of determination (R²) reached 0.9601, with a root mean square error (RMSE) of 0.9669, a mean absolute error (MAE) of 0.6943, and a mean absolute percentage error (MAPE) of 4.32%. Compared to a single image modality basis learner, the R² value was significantly higher. 2The improvement was approximately 1.66 percentage points; compared to a single physical modality base learner, all error metrics were significantly reduced. Furthermore, compared to feature-level fusion strategies that directly concatenate deep semantic features and physical feature vectors along the channel dimension for end-to-end joint optimization, the decision-level fusion model proposed in this invention demonstrates significant advantages in prediction accuracy and local bias correction capabilities. In particular, it effectively avoids the "step-like" overestimation of predicted values caused by cross-modal optimization conflicts in feature-level fusion, especially in the transition range of moderate crushing values.
[0046] S6. Model interpretability verification: To verify the rationality and physical reliability of the decision-making logic of the image modal base learners and physical modal base learners trained in steps S31 and S32 when predicting the crushing value of recycled aggregate, this step uses the gradient weighted class activation mapping method and the SHAP framework based on game theory to perform visualization and quantitative interpretability analysis on the two base learners.
[0047] S61. Interpretability analysis of image modality basis learners: We employ the Grad-CAM gradient-weighted class activation mapping method to visualize and analyze the decision-making mechanism of the ShuffleNet V2 image model. This method uses the gradient information generated in the final convolutional layer by the backpropagation crushing value prediction task to perform weighted summation on the feature channels, generating a spatial heatmap with the same size as the input image, thereby achieving spatial localization of key decision-making regions of the model.
[0048] Input the data-enhanced recycled aggregate image from step S2 into the trained ShuffleNet V2 model to generate the corresponding Grad-CAM heatmap (e.g., ...). Figure 5 (As shown). In the heatmap, the red to yellow areas represent high-activation regions that contribute significantly to the model's predicted crushing values, while the blue to cool-colored areas represent low-activation regions that contribute little.
[0049] Visual analysis results show that the model exhibits a highly robust and physically interpretable attention mechanism during feature extraction. Firstly, for distractors in the image with high color saturation and complex textures but semantically impurities, including plastic pieces, fallen leaves, and paper scraps, the heatmap shows a consistent low-response state with cool colors. This indicates that the model is not dominated by visual saliency cues, but effectively suppresses noise information unrelated to aggregate mechanical properties through deep semantic feature encoding. Secondly, the model's highly activated regions spatially correspond closely to the recycled aggregate regions, particularly focusing on the angular areas of particles, surface textures, and the void structures formed by particle stacking. These highly activated regions are key structural features determining the interlocking state of aggregates and the crushing mechanical response, indicating that the model spontaneously learns physical representations strongly correlated with crushing values. Thirdly, even with large-area occlusions, including fallen leaves or paper scraps occupying the main part of the image, the model can still forcibly shift attention to the unoccluded aggregate regions without significant deviation, confirming that the model does not rely on local texture shortcuts for prediction, but rather possesses selective perception capabilities of the aggregate semantic regions.
[0050] To further verify the robustness of the model under extreme conditions (occlusion scenarios), occlusion sensitivity experiments were conducted on the original images (e.g., Figure 6 (As shown). High-contrast artificial masks were used to accurately cover the core visual regions with the strongest Grad-CAM response in the original image, including three occlusion modes: single-region local occlusion, multi-region scattered occlusion, and vertical strip occlusion. The occluded image was then re-input into the model to generate a heatmap again, and the changes in attention distribution were observed.
[0051] Experimental results show that when high-weight activation regions are occluded, the model's attention does not collapse due to the sudden loss of local key information, nor is it misled by the high-frequency edge noise introduced by the mask itself; the activation values in the masked region all approach zero. Instead, the weights quickly and smoothly shift to aggregate edges and gaps in the image where the activation intensity was originally low, or to other unoccluded aggregate regions. This dynamic attention redirection phenomenon reveals that the model's prediction of aggregate crushing values does not rely on mechanical memory of local single particles, but is based on a global understanding of distributed physical properties such as the spatial topology of the aggregate group, overall particle shape, and surface texture. It possesses strong anti-occlusion capabilities and internal feature redundancy, meeting the high-reliability prediction requirements under complex interferences such as lens contamination, material overlap, or localized lighting distortion in real-world industrial scenarios.
[0052] S62, Interpretability Analysis of Physical Modality Base Learners: The decision-making logic of the random forest model is quantitatively explained using the game theory-based SHAP framework. The SHAP method calculates the marginal contribution of each input feature under different feature combinations, assigns a SHAP value to each feature of each sample, and quantifies the direction and magnitude of the feature's influence on the model's prediction results.
[0053] Perform global feature importance analysis on the random forest model trained in step S32. Calculate the SHAP value (e.g., the sum of squared impacts of each input feature on the crushing value prediction result) for all samples. Figure 7 As shown in the figure, the features were summarized and sorted. The analysis results show that apparent density and water absorption rate rank first and second in the importance of features, far higher than spatial features such as longitude and latitude and aggregate parent material type. Specifically, samples with high apparent density are completely enriched in the negative SHAP value range, indicating that the higher the apparent density of the aggregate, the smaller the crushing value predicted by the model; samples with high water absorption rate are significantly concentrated in the positive SHAP value range, indicating that the higher the water absorption rate of the aggregate, the greater the crushing value predicted by the model. This feature contribution law, which is entirely driven by data, is highly consistent with the basic law in materials mechanics that "increased water absorption rate indicates an increase in micropores and initial microcracks inside the aggregate, leading to the deterioration of the macroscopic mechanical load-bearing skeleton; increased apparent density represents the densification of the matrix structure, enhancing the compressive and shear resistance of the material."
[0054] Further plotting of SHAP feature dependency graphs (e.g.) Figure 8 As shown in the figure, this study delves into the nonlinear response of the random forest model to the key feature of water absorption. Polynomial fitting of the scatter plot reveals that the fitted curve exhibits a slope decay point at approximately 16.25% water absorption. This decay point divides the mechanical response of the aggregate into two physical stages: during the sensitive period of rapid change when the water absorption is below this threshold, the SHAP value increases steeply with increasing water absorption, indicating that an increase in porosity within this range will rapidly destroy the structural integrity of the aggregate, leading to a drastic decrease in macroscopic strength; when the water absorption exceeds this threshold and enters a period of gradual saturation, the slope of the SHAP value increases significantly, indicating that for aggregates already in a highly porous and deteriorated state, further increases in water absorption may still lead to a continued decrease in actual strength, but the model's response to this feature value has tended to saturate. Furthermore, in the high water absorption range, the data points all exhibit a cool color tone that represents low apparent density, indicating that even without explicit physical equation constraints, machine learning models can still autonomously discover and reconstruct the underlying physical law that "high porosity is necessarily accompanied by low apparent density" by leveraging the feature expression capabilities of high-dimensional space.
[0055] In summary, the dual interpretability analysis of Grad-CAM and SHAP fully confirms that the decision-making mechanisms of both the image modal basis learner and the physical modal basis learner conform to the basic evolution laws of materials mechanics, and their prediction results have a solid physical basis, providing a reliable single-modal foundation for the construction and application of the multimodal decision-level fusion model proposed in this invention.
[0056] Example 2: A real-time multimodal detection system for the crushing value of lightweight recycled aggregate according to an embodiment of the present invention, such as... Figure 9 As shown. The multimodal decision-level fusion model trained in step S5 is practically applied to the real-time detection of crushing values of recycled aggregates at construction sites. This embodiment designs and constructs a prototype detection system based on a collaborative architecture of "cloud-based hot update + mobile edge inference". The system adopts a client-server topology and consists of two parts: a mobile terminal acquisition and interaction module and a cloud-based model service module. It realizes rapid, non-destructive, and intelligent grading of the quality of recycled aggregates entering the site. It includes: a mobile terminal acquisition and interaction module, a cloud-based model service module, and a mobile application terminal.
[0057] 1. Cloud Model Service Module: Deployed on a cloud server, responsible for the storage, management, and inference computation of the core inference model. Specifically, it includes the following functions (such as...). Figure 10 (as shown) (1) Model storage and management: The image modality base learner ShuffleNet V2 model weight file, the physical modality base learner random forest model file trained in step S3, and the meta-learner model file trained in step S5 are uniformly deployed on the cloud server. Among them, the deep learning model is stored in .pth format and the machine learning model is stored in .pkl format.
[0058] (2) Model hot update mechanism: A model hot update module is built in the cloud management console. Engineering managers do not need to interrupt server operation. They only need to dynamically switch the currently effective model version in the drop-down menu of the visual interface to achieve seamless and smooth upgrades of different sections' adapted models or algorithm iteration versions. This mechanism avoids encapsulating and integrating the huge inference model in the mobile application, effectively reducing the memory usage of the mobile application to 14.4MB, while ensuring efficient management of multiple sections and multiple versions of models.
[0059] (3) Online Inference Service: The cloud server receives inference requests uploaded by the mobile device and automatically calls the corresponding model for inference calculation based on the modality type identifier contained in the request. When the request contains both image data and physical parameters, the complete multimodal decision-level fusion model is called, and the final predicted value is output by the image modality base learner, the physical modality base learner, and the meta-learner in sequence. When the request contains only image data, the image modality base learner is automatically called for inference. When the request contains only physical parameters, the physical modality base learner is automatically called for inference. After the inference is completed, the prediction result is returned to the mobile device in real time.
[0060] 2. Mobile Data Acquisition and Interaction Module: Built on the Android platform, this module is installed on the handheld mobile devices of engineers at the construction site. It is responsible for data acquisition, local preprocessing, result display, and intelligent classification. Specifically, it includes the following functions (such as...). Figure 11 (as shown) (1) Automated collection of geographic location information: The open-source geographic information system engine is integrated, which allows on-site engineers to accurately anchor the coordinates of the recycled aggregate production site through interactive maps, realize the automated mapping and recording of longitude and latitude parameters, and provide spatial location feature input for the physical modality base learner.
[0061] (2) Image Acquisition and Local Preprocessing: Images of recycled aggregates were captured using the mobile device's camera. To mitigate the impact of network bandwidth fluctuations at the construction site on data transmission, the mobile device used a low-level image processing library to perform local preprocessing and compression on the high-resolution images before uploading them to the cloud server. Preprocessing operations included: bilinear interpolation downsampling the images to a suitable resolution, and converting the images into lightweight data streams using 50% quality JPEG encoding. This strategy reduced the data transmission volume by approximately an order of magnitude while preserving sufficient texture features for ShuffleNet V2 extraction, significantly suppressing timeout anomalies caused by network latency.
[0062] (3) Multimodal data upload: The preprocessed image data, geographical location information and on-site physical parameters (such as apparent density, water absorption rate, etc., when the measurement conditions are available) are packaged into a unified inference request and sent to the cloud model service module via wireless network.
[0063] (4) Result Reception and Intelligent Classification Display: The system receives the crushing value prediction results returned from the cloud and displays them visually on the mobile interface. The mobile terminal has a built-in intelligent classification module based on the current national standard GB / T 25177. This module automatically classifies recycled aggregates into three levels: Class I, Class II, and Class III, based on the predicted crushing value. These levels correspond to the aggregate quality grades applicable to different types of engineering projects, providing on-site engineers with immediate and effective decision support.
[0064] 3. The complete workflow of the detection system is as follows: (1) Training phase: Import the enhanced multimodal dataset constructed in steps S1 and S2 into the cloud server, complete the training of the single-modal base learner and the training of the multimodal decision-level fusion model according to the process from steps S3 to S5, and deploy all trained model files to the cloud model service module.
[0065] (2) Testing phase: On-site engineers use mobile devices to first locate the aggregate production area through an interactive map and obtain latitude and longitude information; then take pictures of the recycled aggregate batch to be tested, and input physical parameters such as apparent density and water absorption rate of the batch when conditions permit; after the mobile device performs local preprocessing and compression of the images, it encapsulates the multimodal data into an inference request and uploads it to the cloud; the cloud automatically selects the corresponding model for inference based on modal integrity, and returns the crushing value prediction results and aggregate quality grade to the mobile device for display.
[0066] (3) Model update stage: When it is necessary to fine-tune the model or iterate the algorithm for the aggregate characteristics of a specific mining area, the engineering management personnel can dynamically switch the model version through the hot update module in the cloud management console. All mobile devices can use the latest model synchronously without reinstalling the application, realizing zero-downtime deployment of model iteration.
[0067] Therefore, this invention employs a lightweight, multimodal real-time detection method and system for crushing values of recycled aggregates. Through decision-level stacking and fusion of image and physical modes, it significantly outperforms single-modal models and overcomes the cross-modal optimization conflict problem in feature-level fusion. The system can automatically select fusion prediction, pure image prediction, or pure physical parameter prediction based on the modal integrity of the uploaded data, making it suitable for complex scenarios such as construction sites where only images can be acquired or only physical parameters can be measured. The mobile terminal reduces the transmission volume by an order of magnitude through local image preprocessing, meeting the low-latency and low-power edge computing requirements of "instant measurement." It supports dynamic switching of model versions in the cloud management console, allowing mobile terminals to update synchronously without reinstalling the application, achieving seamless upgrades and efficient operation and maintenance of model iterations.
[0068] It is worth noting that all contents not described in detail in this invention are existing technologies and are well known to those skilled in the art.
[0069] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can still be made to the technical solutions of the present invention, and these modifications or equivalent substitutions cannot cause the modified technical solutions to deviate from the spirit and scope of the technical solutions of the present invention.
Claims
1. A method for real-time multimodal detection of crushing value of lightweight recycled aggregate, characterized in that, Includes the following steps: S1. Data Acquisition: Collect images of recycled aggregates with different types and gradations of recycled aggregate parent materials, and conduct standard crushing value tests on them to measure the precise crushing value corresponding to each recycled aggregate image. Assign a corresponding precise crushing value label to each recycled aggregate image to form a benchmark image dataset; collect physical property data of recycled aggregates that correspond one-to-one with the recycled aggregate image samples, including data collected in the laboratory and supplemented by data extracted from published literature, and construct a structured dataset together with the benchmark image dataset; S2. Data Augmentation: Perform two-stage data augmentation processing on the benchmark image dataset; The first stage is basic data augmentation, which introduces image transformation operations. Each original image is processed by image transformation operations four times. The second stage is random occlusion enhancement, where lightweight objects are randomly generated to occlude regions on the image obtained in the first stage; the final result is a data-enhanced multimodal dataset. S3. Single-modal basis learner training: Construct image modal basis learners and physical modal basis learners, and train them separately; S4. Dual-path cross-validation and prediction generation: Data is partitioned and an asymmetric dual-path cross-validation mechanism is constructed. Five-fold cross-validation is performed on the image modality base learner and the physical modality base learner respectively to generate a sequence of prediction values that are completely independent at the sample level. S5. Construction of Multimodal Decision-Level Fusion Model: A decision-level fusion strategy is adopted, and a stacked generalization method is introduced. The meta-learner performs nonlinear combination of the output of the base learner, and finally outputs the prediction result of the sample crushing value.
2. The method for real-time multimodal detection of crushing value of lightweight recycled aggregate according to claim 1, characterized in that, In step S1, the types of recycled aggregate parent materials include natural aggregates, recycled concrete aggregates, recycled brick aggregates, and recycled mixtures; The characteristic fields of the structured dataset include longitude, latitude, pressure load, crushing value, apparent density, and water absorption rate.
3. The method for real-time multimodal detection of crushing value of lightweight recycled aggregate according to claim 1, characterized in that, In step S2, the basic data enhancement specifically involves introducing image transformation operations including random rotation, random flipping, random cropping, brightness jitter, saturation jitter, contrast jitter, hue jitter, random sharpness adjustment, RGB offset, and Gaussian noise. Each original image is processed by the image transformation operations four times. The random occlusion enhancement specifically involves randomly generating lightweight debris occlusion areas on the image obtained in the first stage; the lightweight debris includes leaves, pieces of paper, plastic, and film, and their positions, sizes, and angles in the image are all generated using a random generation strategy.
4. The method for real-time multimodal detection of crushing value of lightweight recycled aggregate according to claim 1, characterized in that, Step S3 specifically includes: S31. The image modality base learner uses a lightweight convolutional neural network model, including the ShuffleNet V2 lightweight convolutional neural network, as the backbone network. Image modality base learner training: The augmented recycled aggregate image dataset was uniformly adjusted to 224×224 pixels. The model parameters were initialized using weights pre-trained on the ImageNet dataset, and fine-tuned by transfer learning on the recycled aggregate image dataset. The Adam optimizer was used with an initial learning rate of 0.001 and a batch size of 512. The training lasted for 100 epochs, and an early stopping mechanism was introduced to automatically terminate training when the validation set loss did not decrease within 20 consecutive epochs. S32. The physical modality base learner adopts a machine learning model including a random forest ensemble learning model. The input features include numerical variables and categorical variables, and the crushing value of recycled aggregate is used as the target variable for model output. The numerical variables include longitude, latitude, pressure load, apparent density and water absorption rate, and the categorical variables are the types of recycled aggregate parent materials. Physical modality base learner training: Numerical features are standardized, and categorical features are encoded using one-hot encoding. Key hyperparameters of the random forest are tuned using a grid search combined with five-fold cross-validation. The hyperparameter search range includes: the number of decision trees ({50, 100, 200, 300}), the maximum depth ({5, 10, 15, None}), and the minimum number of split samples ({2, 5, 10}). Using root mean square error as the optimization objective, the optimal hyperparameter combination is selected after 20 random searches using RandomizedSearchCV. The model's generalization ability and stability are then evaluated on an independent test set.
5. The method for real-time multimodal detection of crushing value of lightweight recycled aggregate according to claim 4, characterized in that, Step S4 specifically includes: S41. Data partitioning: The physical property datasets that correspond one-to-one with the recycled aggregate image samples collected by the laboratory in step S1 are denoted as dataset A, and the physical property datasets that are supplemented by extracting from published literature are denoted as dataset B. Based on the types of recycled aggregate parent materials, stratified random sampling was performed on dataset A, and 20% of the samples were divided into independent test sets, while the remaining 80% of the samples were recorded as cross-validation data. S42, Physical Modality Cross-Validation Path: The cross-validation data is randomly divided into 5 equal-sized subsets; in each fold cross-validation, 4 subsets are selected and together with all samples from dataset B to form the training set for the current fold, and the remaining 1 subset forms the validation set; the machine learning model after hyperparameter tuning in step S32 is used to train on the training set for the current fold, and a physical modality prediction value is generated for each set of samples in the validation set for the current fold. S43, Image Modality Cross-Validation Path: The cross-validation data is divided into 5 folds according to the same subset partitioning method as the physical modality path. In each fold of cross-validation, all recycled aggregate images associated with 4 subsets are selected to form the training set for the current fold, and all images associated with the remaining 1 subset are selected to form the validation set. The lightweight convolutional neural network fine-tuned by transfer learning in step S31 is used to train on the training set for the current fold, and an image modality prediction value is generated for each set of samples in the validation set for the current fold. For the case where a single sample contains multiple images, the average of the prediction values of all images corresponding to the same validation sample is taken as the image modality prediction value of that sample. S44. Predicted value concatenation and meta-feature set construction: The physical modality predicted value and the image modality predicted value are paired and concatenated according to samples to form a two-dimensional prediction feature matrix; Each row of the two-dimensional prediction feature matrix corresponds to a sample, containing two input features: the image modality prediction value and the physical modality prediction value of the sample. The two-dimensional prediction feature matrix is used as the input variable, and the true crushing value of the sample is used as the target variable to form a meta-feature set for training the meta-learner.
6. The method for real-time multimodal detection of crushing value of lightweight recycled aggregate according to claim 5, characterized in that, Step S5 specifically involves: Decision-level fusion is performed at the output level of the model prediction, specifically by introducing a meta-learner to learn the mapping relationship between the base model predictions and the true target. The base learners are the image model and physical feature model trained in step S3, denoted as follows: and ; For each sample in the training set in steps S42 and S43, two initial prediction values are generated by the base learner. and , and the true value Together they constitute a new meta-feature set , For the samples in the training set, The total number of samples in the training set; It is divided into a meta-training set and a meta-validation set for training and tuning the meta-learner; During training, the meta-learner aims to minimize the mean squared error by learning the combined weights and interaction effects of the predictions from the two base models; finally, for new samples, it fuses the predicted values. The calculation process is shown in the following formula: ; ; ; in, For a trained meta-learner.
7. A real-time multimodal detection system for the crushing value of lightweight recycled aggregate, employing the real-time multimodal detection method for the crushing value of lightweight recycled aggregate as described in any one of claims 1-6, characterized in that, include: Mobile data acquisition and interaction module, cloud model service module, mobile application terminal; The mobile terminal data acquisition and interaction module is built on the Android platform and installed on the handheld mobile devices of engineers at the construction site. It is responsible for data acquisition, local preprocessing, result display and intelligent classification. The cloud model service module is deployed on a cloud server and is responsible for the storage, management and inference calculation of the core inference model. The model weight file of the image modality base learner trained in step S3, the model file of the physical modality base learner, and the meta-learner model file trained in step S5 are all deployed on the cloud server.
8. The multimodal real-time detection system for the crushing value of lightweight recycled aggregate according to claim 7, characterized in that, The mobile terminal data acquisition and interaction module specifically includes the following functions: Automated collection of geographic location information: Integrating an open-source geographic information system engine, it allows field engineers to accurately anchor the coordinates of recycled aggregate production sites through interactive maps, realize the automated mapping and recording of longitude and latitude parameters, and provide spatial location feature input for physical modality base learners; Image acquisition and local preprocessing: Capture images of recycled aggregates using the mobile device's camera; Before the image is uploaded to the cloud server, the mobile device calls the underlying image processing library to perform local preprocessing and compression on the high-resolution image. The preprocessing operations include: downsampling the image to a suitable resolution using bilinear interpolation, and converting the image into a lightweight data stream using JPEG encoding at 50% quality. Multimodal data upload: The pre-processed image data, geographic location information, and on-site physical parameters are encapsulated into a unified inference request and sent to the cloud model service module via wireless network; Results reception and intelligent classification display: Receive the crushing value prediction results returned from the cloud and display them visually on the mobile interface; The mobile terminal has a built-in intelligent classification module, which automatically classifies the recycled aggregate into three levels: Class I, Class II and Class III, based on the predicted crushing value, which correspond to the aggregate quality grades applicable to different types of engineering projects.
9. A multimodal real-time detection system for the crushing value of lightweight recycled aggregate according to claim 7, characterized in that, The cloud-based model service module specifically includes the following functions: Model hot update mechanism: A model hot update module is built in the cloud management console; project managers do not need to interrupt the server operation, but only need to dynamically switch the currently effective model version in the drop-down menu of the visual interface to achieve seamless and smooth upgrades of different sections' adapted models or algorithm iteration versions; Online Inference Service: The cloud server receives inference requests uploaded by mobile devices and automatically calls the corresponding model for inference calculation based on the modality type identifier contained in the request. When the request contains both image data and physical parameters, the complete multimodal decision-level fusion model is invoked, and the final predicted value is output by the image modality base learner, physical modality base learner, and meta-learner in sequence. When the request contains only image data, the image modality base learner is automatically called for inference. When the request contains only physical parameters, the physical modality base learner is automatically called for inference. After the inference is completed, the prediction result is returned to the mobile device in real time.
10. A multimodal real-time detection system for the crushing value of lightweight recycled aggregate according to claim 7, characterized in that, The workflow of the detection system is as follows: Training phase: Import the augmented multimodal dataset obtained in step S2 into the cloud server, complete the training of the single-modal base learner and the multimodal decision-level fusion model according to the process of steps S3 to S5, and deploy all trained model files to the cloud model service module. Testing phase: Using a mobile device, the aggregate production location is first located through an interactive map to obtain latitude and longitude information; then, images of the recycled aggregate batch to be tested are taken, and physical parameters such as apparent density and water absorption rate of the batch are input when conditions permit; after the mobile device performs local preprocessing and compression of the images, the multimodal data is encapsulated into an inference request and uploaded to the cloud; The cloud automatically selects the corresponding model for inference based on modal integrity, and returns the crushing value prediction results and aggregate quality grade to the mobile device for display. Model update phase: When it is necessary to fine-tune the model or iterate the algorithm for the aggregate characteristics of a specific mining area, the model version can be dynamically switched through the hot update module in the cloud management console. All mobile devices can use the latest model synchronously without reinstalling the application.