Concrete workability measurement method and system based on multimodal visual large model

Through the combination of multimodal visual large model and rheological algorithm, non-contact rapid measurement and optimization of concrete working performance parameters is achieved, and the problems of low efficiency and low accuracy in traditional methods are solved, and the construction quality and intelligence level are improved.

CN120355715BActive Publication Date: 2025-08-22SHENZHEN UNIV
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510846225.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-24
Publication Date
2025-08-22
Estimated Expiration
2045-06-24

AI Technical Summary

Technical Problem

Traditional concrete working performance measurement methods are inefficient and have low accuracy, and rely on manual operations, so real-time monitoring and comprehensive evaluation cannot be achieved.

Method used

Using a method based on multimodal vision big model, the contactless rapid measurement and optimization of concrete working performance parameters is achieved through image acquisition, feature matching and three-dimensional reconstruction, combined with rheological algorithms and neural networks.

Benefits of technology

It realizes accurate identification and verification of concrete working performance parameters, improves measurement efficiency and accuracy, optimizes material costs, and enhances construction performance and the level of intelligence of the entire process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120355715B_ABST
    Figure CN120355715B_ABST
Patent Text Reader

Abstract

The present application relates to a method and system for measuring the working performance of concrete based on a large multimodal visual model, which solves the problem of the relatively troublesome rapid detection of concrete working performance parameters. The method includes: based on the spatial semantic understanding capability of the large multimodal visual model, combined with multi-view stereo vision and structured light scanning technology, reconstructing the three-dimensional geometric structure of the concrete paste, extracting morphological characteristic parameters, and forming a feature vector; inputting the feature vector into a pre-trained multi-task neural network, fusing spatiotemporal features and combining with rheological algorithms to identify various working performance parameters; using ensemble learning to fuse at least three recognition results, and verifying the parameters based on the basic equations of fluid dynamics through a fluid simulation platform; based on the verification results, generating a mix ratio optimization suggestion including material components. The present application has the following effects: realizing non-contact rapid measurement of concrete working performance parameters, improving accuracy and efficiency, and providing mix ratio optimization suggestions.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of construction engineering detection, and in particular to a concrete working performance measurement method and system based on a multimodal visual large model. Background Art

[0002] As one of the most commonly used materials in construction, concrete's workability properties (such as rheology, fluidity, and viscosity) directly impact construction quality and structural durability. Traditional methods for measuring concrete workability, such as slump and flow tests, require specialized equipment and manual labor, resulting in low efficiency and subjectivity.

[0003] Traditional testing requires on-site sampling and the use of professional equipment for testing. Although it has high accuracy, the detection efficiency is low.

[0004] Regarding the above-mentioned related technologies, the following problems exist: the dedicated equipment is complex to operate and has limited on-site application; the detection lag makes it difficult to achieve real-time monitoring; a single parameter cannot fully evaluate workability; and the results rely on manual experience, resulting in a lack of objectivity. Summary of the Invention

[0005] In order to achieve non-contact rapid measurement of concrete working performance parameters, improve accuracy and efficiency, and provide mix ratio optimization suggestions, this application provides a concrete working performance measurement method and system based on a multimodal visual large model.

[0006] In a first aspect, the present application provides a method for measuring concrete working performance based on a multimodal visual large model, which adopts the following technical solutions:

[0007] A method for measuring concrete workability based on a multimodal visual large model, comprising:

[0008] Collect multi-angle images, continuous video sequences, and environmental parameters of concrete paste at different construction stages to construct a multi-dimensional spatiotemporal dataset;

[0009] The pre-trained multimodal vision model is used to identify standard reference objects in the image, and the feature matching algorithm is used to establish the proportional mapping relationship between the image pixel coordinate system and the physical world coordinate system.

[0010] Based on the spatial semantic understanding capability of the multimodal visual large model, combined with multi-view stereo vision and structured light scanning technology, the 3D geometric structure of the concrete paste is reconstructed, morphological feature parameters are extracted, and feature vectors are formed.

[0011] The feature vectors are input into a pre-trained multi-task neural network, which fuses the spatiotemporal features and combines them with rheological algorithms to identify various working performance parameters.

[0012] Ensemble learning is used to fuse at least three identification results, and the parameters are verified based on the basic equations of fluid dynamics through a fluid simulation platform;

[0013] Based on the verification results, a non-dominated sorting genetic algorithm is used to generate mix ratio optimization suggestions including material components with construction performance and material cost as multi-objective optimization functions.

[0014] By adopting the above technical solution, this method can achieve accurate identification and verification of concrete working performance parameters through multimodal data fusion, three-dimensional reconstruction and intelligent algorithms, and generate mix ratio recommendations in combination with multi-objective optimization, which can improve measurement efficiency and accuracy, optimize material costs, ensure construction performance, and enhance the intelligence level of the entire process.

[0015] Optionally, a pre-trained multimodal vision model is used to identify standard reference objects in the image, and a feature matching algorithm is used to establish a proportional mapping relationship between the image pixel coordinate system and the physical world coordinate system, including:

[0016] Define a reference object with a standard size preset in the imaging area, use the target detection module of the large model to identify the outline of the reference object, separate the reference object from the background through the instance segmentation algorithm, extract the pixel coordinate range and calculate its spatial position relationship;

[0017] Perform SIFT feature extraction on the key feature points of the reference object, perform feature matching with the pre-stored standard model, optimize the matching results based on the RANSAC algorithm, and calculate the pixel-physical size ratio coefficient;

[0018] The Zhang calibration method is used to calibrate the internal parameters of the imaging system. The three-dimensional reconstruction model is established by combining the spatial constraint relationship of multiple reference objects. The sub-pixel measurement of concrete parameters is achieved based on the pixel-physical size ratio coefficient.

[0019] By adopting the above technical solutions, this method accurately identifies reference objects through large models, combines SIFT feature matching with RANSAC optimization, establishes a highly reliable pixel-to-physical mapping, and uses Zhang calibration and multi-reference object constraints to improve measurement accuracy to the sub-pixel level, providing a solid spatial positioning foundation for subsequent three-dimensional reconstruction and parameter analysis, and enhancing the accuracy and stability of system measurements.

[0020] Optionally, in a standard slump test scenario, a pre-trained multimodal visual model is used to identify standard reference objects in the image, and a feature matching algorithm is used to establish a proportional mapping relationship between the image pixel coordinate system and the physical world coordinate system, including:

[0021] A standard slump bucket with a preset height is defined as a geometric reference object placed in the imaging area. The vertical distance between the center of its bottom circle and the imaging plane is preset to a fixed value. A circular marking point with a diameter of d is set on the surface of the bucket as a feature positioning benchmark. The target detection module of the multimodal vision large model is used to identify the outline of the slump bucket. The bucket body and the concrete slump body are separated by the instance segmentation algorithm, and the pixel coordinate range of the bucket body and the two-dimensional pixel coordinates of the marking point are extracted. ;

[0022] Based on the pre-trained monocular vision depth estimation model, the depth value of the slumping bucket marker is predicted, and the 3D physical coordinates of the marker are calculated by combining the camera intrinsic parameter matrix. , construct the barrel space coordinate system;

[0023] SIFT features are extracted from the edge points of the top surface of the collapsed barrel and matched with the pre-stored standard barrel 3D model. The mismatched points are eliminated based on the RANSAC algorithm, and the proportional coefficient k between the pixel distance and the actual physical distance is calculated.

[0024] By adopting the above technical solution, this method takes the slump bucket as the benchmark, accurately segments the target through a large model, combines monocular depth estimation with three-dimensional model matching, constructs the bucket coordinate system and calculates the scale coefficient k, realizes high-precision mapping of pixels to physical coordinates, provides a sub-pixel measurement benchmark for slump testing, and improves the geometric positioning accuracy and scene adaptability of concrete working performance parameter measurements.

[0025] Optionally, in a standard slump test scenario, the extracted morphological characteristic parameters include:

[0026] By pre-training the spatial semantic understanding capabilities of a large multimodal visual model, the pixel coordinates of the highest point of the concrete collapse are located. Based on the proportional coefficient k between the pixel distance and the actual physical distance, combined with the bucket's spatial coordinate system, the vertical distance between the highest point of the collapse and the top surface of the collapse bucket is calculated to obtain the slump value.

[0027] The instance segmentation algorithm is used to segment the collapsed concrete edge at the pixel level and identify the edge contour of the concrete diffusion area. Based on the scale factor k, the concrete edge contour in the pixel coordinate system is converted into a contour curve in the physical coordinate system. The maximum diameter of the concrete diffusion in the direction perpendicular to the axis of the collapse bucket is calculated.

[0028] The surface of the concrete collapse body is reconstructed at the sub-pixel level in 3D. The normal vector of each surface point is calculated based on the differential geometry method. By presetting the curvature threshold, the key curvature feature points of the collapse body surface are extracted, the mean curvature and Gaussian curvature are calculated, and the degree of morphological change of the collapse body surface is quantified.

[0029] The image semantic segmentation capability of the multimodal visual large model is used to identify defective areas on the surface of the collapsed body. By calculating the area ratio and distribution density of the defective areas, the uniformity and molding quality of the concrete are evaluated.

[0030] By adopting the above technical solution, the method uses large-model semantic understanding and visual algorithms to accurately extract parameters such as slump and diffusion diameter, combines three-dimensional reconstruction to quantify surface morphology, identify defects and evaluate uniformity, and realize automatic measurement of multi-dimensional parameters of slump testing, thereby improving detection efficiency and accuracy and providing comprehensive data support for concrete workability analysis.

[0031] Optionally, during the concrete mixing, transportation, or pouring process, based on the spatial semantic understanding capabilities of the multimodal visual large model, combined with multi-view stereo vision and structured light scanning technology, the 3D geometric structure of the concrete paste is reconstructed, morphological feature parameters are extracted, and feature vectors are formed, including:

[0032] Based on multi-view stereo vision technology, the segmented concrete paste is 3D reconstructed to construct a point cloud model of its surface;

[0033] The high-precision texture information of the concrete surface is obtained through structured light scanning technology and integrated with the point cloud model to obtain a more accurate three-dimensional geometric structure.

[0034] Leveraging the spatial semantic understanding capabilities of a multimodal visual large model, the 3D geometric structure is analyzed to identify the morphological characteristics of concrete paste at different locations.

[0035] If the morphological feature is a continuous curve change on the surface, the curve is modeled by mathematical fitting methods to extract the curvature and slope of the curve to quantify the degree of change of the curve;

[0036] If the morphological feature is an arc shape at the corner, a method based on a geometric model is used to accurately describe it and calculate the arc radius, center position, and arc length;

[0037] The extracted morphological characteristic parameters are combined and screened to form a characteristic vector related to the rheological properties of concrete.

[0038] By adopting the above technical solution, this method integrates multi-view stereo vision and structured light scanning, combined with large-model semantic analysis, to achieve high-precision reconstruction of the three-dimensional structure of concrete slurry. It uses adaptive extraction methods for different morphological features to form rheologically related feature vectors, providing multi-dimensional and precise data for the identification of working performance parameters, and improving the comprehensiveness and accuracy of the analysis.

[0039] Optionally, the feature vectors are fed into a pre-trained multi-task neural network, which fuses the spatiotemporal features and combines them with rheological algorithms to identify various performance parameters, including:

[0040] A recurrent neural network is used to process time series data of feature vectors to extract the dynamic characteristics of concrete during mixing, transportation, and pouring. Convolutional neural networks are also used to perform secondary extraction of the spatial characteristics of concrete's three-dimensional geometric structure. The time dimension feature vectors and the space dimension feature vectors are tensor-concatenated to construct a composite feature representation that includes spatiotemporal information.

[0041] The constitutive equations of the Bingham model and the Herschel-Bulkley model are introduced as physical constraints in the fully connected layer of the multi-task neural network to map the spatiotemporal composite feature vector to the rheological parameter space.

[0042] Through the multi-branch structure of the neural network output layer, the viscosity, shear stress, yield strength, and thixotropy index of concrete are calculated in parallel. Combined with real-time environmental parameters, a correction function is constructed to compensate for the initial recognition results and output the final accurate measurement values ​​of the working performance parameters.

[0043] By adopting the above technical solution, this method fuses spatiotemporal features through cyclic and convolutional neural networks, combines the physical constraints of the constitutive equation and multi-branch output, accurately extracts dynamic changes and spatial features, and introduces environmental parameter correction, which can improve the comprehensiveness, accuracy and environmental adaptability of work performance parameter measurements.

[0044] In a second aspect, the present application provides a concrete workability measurement system based on a multimodal visual large model, which adopts the following technical solutions:

[0045] A concrete work performance measurement system based on a multimodal visual large model includes a memory, a processor, and a program stored in the memory and executable on the processor. The program, when loaded and executed by the processor, can implement the concrete work performance measurement method based on the multimodal visual large model as described in the first aspect. BRIEF DESCRIPTION OF THE DRAWINGS

[0046] Figure 1 This is a flow chart of a method for measuring concrete working performance based on a multimodal visual large model according to an embodiment of the present application.

[0047] Figure 2 This is a flowchart of another embodiment of the present application, which uses a pre-trained multimodal visual large model to identify standard reference objects in an image and uses a feature matching algorithm to establish a proportional mapping relationship between the image pixel coordinate system and the physical world coordinate system. DETAILED DESCRIPTION

[0048] The present application is further described in detail below with reference to the accompanying drawings.

[0049] Reference Figure 1, a concrete working performance measurement method based on a multimodal visual large model disclosed in this application, comprising:

[0050] Step S100: Acquire multi-angle images, continuous video sequences, and environmental parameters of concrete slurry at different construction stages to construct a multi-dimensional spatiotemporal dataset.

[0051] Multi-angle images: Static images captured from multiple angles (such as the front, side, and top) of the concrete slurry, used to capture morphological characteristics from different perspectives. Continuous video sequences: Continuous dynamic images captured during the concrete construction process (mixing, transportation, pouring, and other stages), used to analyze the dynamic characteristics of the slurry over time. Environmental parameters: These include, but are not limited to, external environmental data that affects the workability of concrete, such as temperature, humidity, and light intensity.

[0052] Image and video acquisition: Industrial cameras, high-speed cameras, and other equipment deployed on construction equipment (such as mixers and conveying pumps) or at fixed locations are used to collect images and videos during each stage of concrete construction (mixing, transportation, and pouring).

[0053] Environmental parameter acquisition: Real-time monitoring and recording of construction environment data through temperature and humidity sensors, light sensors and other equipment.

[0054] The general process is as follows: During the concrete construction process, image / video acquisition equipment and environmental sensors are used to synchronously obtain multi-angle images, continuous video sequences and corresponding environmental parameters of the concrete slurry at different stages, and integrate them to form a multidimensional data set containing spatial (image / video) and temporal (continuous sequence / environmental changes) dimensions, providing basic data support for subsequent analysis.

[0055] In step S200, the standard reference objects in the image are identified by a pre-trained multimodal visual model, and a feature matching algorithm is used to establish a proportional mapping relationship between the image pixel coordinate system and the physical world coordinate system.

[0056] Among them, the pre-trained multimodal visual large model: a deep learning model pre-trained based on massive image / video data, with target detection (identifying object location) and instance segmentation (separating objects from background) capabilities, such as the YOLOv8+MaskR-CNN fusion model. Standard reference objects: calibration objects with known fixed physical dimensions, such as a standard slump bucket with a bottom diameter of 100mm and a height of 300mm, or a geometric component with a 5mm diameter circular marking point on the surface. Feature matching algorithm: The SIFT algorithm is used to extract key points of objects in the image (such as edge corners), and the RANSAC algorithm is used to eliminate mismatched points and retain correctly matched pairs. Pixel-to-physical ratio coefficient: The actual physical length corresponding to 1 pixel in the image (such as 1 pixel = 0.5mm) is calculated by combining the actual size of the reference object with the pixel size.

[0057] The general process is described in detail in steps S210 to S230 and will not be repeated here.

[0058] In addition, before using the pre-trained multimodal vision model to identify standard reference objects in the image, consider adding an adaptive perception module to the multimodal vision model to evaluate image quality and feature expression capabilities in real time and dynamically adjust the feature extraction strategy to improve the system's adaptability to different lighting conditions, viewing angle changes, concrete types, and construction scenarios. The specific steps are as follows:

[0059] Step 101: introduce the adaptive perception module of the multimodal visual large model to analyze the quality characteristics of the input image in real time, including but not limited to contrast, sharpness, noise level, texture complexity and color distribution.

[0060] Among them, the adaptive perception module is an intelligent component that evaluates image quality in real time and adjusts feature extraction strategies.

[0061] Image quality characteristics: metrics that quantify the visual properties of an image, including: Contrast: the difference between bright and dark areas of an image (e.g., contrast ratio, CR). Sharpness: edge clarity (measured by the variance of the Laplace operator). Noise level: the degree of random interference in the image (e.g., the standard deviation of Gaussian noise).

[0062] Texture complexity: the richness of image details (assessed by the local binary pattern (LBP) histogram entropy value).

[0063] Color distribution: color space statistical characteristics (such as HSV color gamut coverage).

[0064] The general process is as follows: 1. Image preprocessing: Convert the input image to the HSV color space and separate the brightness (V channel) and color information. 2. Quality feature extraction: Contrast: Calculate the Michelson contrast of the V channel with a threshold range of [0,1] (recommended threshold ≥0.3). Sharpness: Apply the Laplacian operator to the V channel and calculate the gradient variance (recommended threshold ≥50). Noise level: Perform a median filter on the V channel and calculate the residual standard deviation of the original image and the filtered image (recommended threshold ≤10). Texture complexity: Calculate the entropy of the LBP histogram (recommended range 2.0-7.0). Color distribution: Count the distribution density of pixels in the HSV space and calculate the color gamut coverage (recommended threshold ≥70%). 3. Feature normalization: Normalize each quality indicator to the range of [0,1] to facilitate subsequent decision-making.

[0065] Step 102 : Evaluate the applicability of the current feature extraction strategy based on the image quality features, and establish a mapping relationship between the quality features and the feature extraction strategy by comparing the feature extraction effects of the model under different image quality conditions.

[0066] Feature extraction strategy: The specific method used by the model to extract image features (such as convolution kernel size and attention mechanism weights). Mapping relationship: Establishing a functional relationship between image quality features and the optimal feature extraction strategy. Applicability evaluation: Using quantitative indicators to determine the effectiveness of the current strategy under specific image quality conditions.

[0067] Feature consistency: Calculates the cosine similarity of feature vectors under different quality conditions (threshold ≥ 0.85). Key region response: Measure activation values ​​in key regions such as the edge of the collapsed barrel (visually evaluated using Grad-CAM). Semantic segmentation accuracy: Uses mIoU (mean intersection over union) to assess the accuracy of reference object recognition.

[0068] The general process is as follows: 1. Dataset Construction: Collect over 1,000 concrete images under varying quality conditions (covering variations in lighting, viewing angle, and material type). Five different feature extraction strategies are applied to each image, and the corresponding quality characteristics and evaluation metrics are recorded. 2. Offline Training: Use a decision tree model to fit the relationship between quality characteristics and the optimal strategy. Cross-validation (k=5) is performed to ensure model generalization (accuracy ≥90%). 3. Online Mapping: Input the quality feature vectors of a live image. Use the decision tree to predict the optimal feature extraction strategy.

[0069] Step 103: Dynamically generate a feature extraction optimization plan, use a reinforcement learning algorithm to optimize the feature extraction process, adjust the model's attention mechanism, and enhance the perception of key feature areas;

[0070] Reinforcement learning (RL): An optimization framework for maximizing cumulative rewards through interaction between an agent and its environment. Feature extraction optimization: A strategy for dynamically adjusting model parameters (such as attention weights and convolution kernel configurations). Reward function: A metric that quantifies the effectiveness of feature extraction (such as reference object recognition accuracy and feature vector stability).

[0071] RL framework: Proximal Policy Optimization (PPO) algorithm (dominant Actor-Critic architecture). State space: Image quality feature vector (contrast, sharpness, noise, texture, color) from step S101. Action space: Discrete action set, including: Adjusting attention weight (±10%), Switching convolution kernel size (3×3 / 5×5 / 7×7), Enabling / disabling specific feature channels. Reward function:

[0072] .

[0073] The general process is as follows: 1. Environment Construction: Encapsulate the large multimodal visual model into an interactive environment that receives actions and returns rewards. 2. Policy Training: Use the PPO algorithm to train the agent, learning the optimal policy through over 10,000 simulated interactions. Experience Replay (ReplayBuffer) and entropy regularization are used to improve training stability. 3. Online Optimization: Input the current image quality features in real time (step S101).

[0074] The agent outputs the optimal action sequence (e.g., increase attention weight → switch to a 5×5 convolution kernel). 4. Model update: After processing every 100 images, the RL strategy is fine-tuned using the new data to adapt to the scene changes.

[0075] Step 104 , adaptively adjust the feature extraction layer of the multimodal vision large model according to the optimization scheme, including but not limited to adjusting the convolution kernel size, the number of feature channels, and the attention weight, to enhance the model's ability to express features for different construction scenarios.

[0076] Feature extraction layers are the neural network layers (such as convolutional layers and Transformer layers) responsible for extracting image features within large multimodal vision models. Adaptive adjustment dynamically modifies model parameters (such as convolution kernel weights and attention mechanisms) based on image quality. Attention weights control the coefficients by which the model focuses on different areas, enhancing the perception of key features.

[0077] Adjustments: 1. Convolution kernel parameters: Dynamically adjust the kernel size (3×3 / 5×5 / 7×7) and stride (1 / 2). 2. Channel attention: Recalibrate feature channel weights using the Squeeze-and-Excitation (SE) module. 3. Spatial attention: Enhance key region responses using CBAM.

[0078] Adjustment rules: 1. When the image is blurry (sharpness < 0.4), increase the convolution kernel size (e.g., from 3×3 to 5×5). 2. When the noise is high (noise level > 0.6), reduce the number of feature channels (e.g., by 20%). 3. When the contrast is low (contrast < 0.3), increase the spatial attention weight (e.g., by 30%).

[0079] The general process is as follows: 1. Optimization solution analysis: Receive the action sequence generated in step S103 (e.g., increase attention weight → switch to a 5×5 convolution kernel). 2. Parameter mapping: Convert the RL action to a specific model parameter adjustment (e.g., "increase attention weight by 0.15" corresponds to a SE module scaling factor of ×1.15). 3. Model modification: Dynamically modify the model's parameters during forward propagation using a hook function. 4. Validation and adjustment: Use a validation set (e.g., 10 representative images) to confirm that the adjusted model's feature extraction performance has improved.

[0080] Step 105 , re-execute the feature extraction operation, analyze the image using the adaptively adjusted model, and obtain a more accurate feature vector for subsequent work performance parameter identification.

[0081] The general process is as follows: 1. Model loading: Load the multimodal visual model dynamically adjusted in step S104. 2. Forward propagation: Input the preprocessed image (e.g., normalized to 0,1). Generate a feature map using the modified convolutional layer and attention module. 3. Feature vector generation: Globally pool the feature map to generate a fixed-length feature vector. 4. Quality verification: Calculate the L2 norm and entropy of the feature vector to determine whether it meets the threshold requirements. 5. Output and storage: Store the verified feature vector as a numpy array or JSON format for subsequent parameter identification.

[0082] In step S300, based on the spatial semantic understanding capability of the multimodal visual large model, multi-view stereo vision and structured light scanning technology are combined to reconstruct the three-dimensional geometric structure of the concrete slurry, extract morphological feature parameters, and form a feature vector.

[0083] Among them, spatial semantic understanding capability: the ability of a large multimodal visual model to understand the spatial relationship (such as position, shape, and topological structure) of objects in an image, such as identifying the relative position of a concrete slump and a slump bucket. Multi-view stereo vision technology: a technology that reconstructs three-dimensional structures through multi-perspective images, such as generating point clouds by using multiple images taken by a binocular camera or a single camera. Structured light scanning technology: a high-precision scanning method that projects structured light (such as stripes, coded patterns) onto the surface of an object and calculates three-dimensional coordinates through the deformed pattern. Morphological feature parameters: quantitative indicators that characterize the geometric morphology of concrete, such as slump value, diffusion diameter, surface curvature, defect area ratio, etc. Feature vector: a numerical vector formed by combining multiple morphological feature parameters in a specific order, used for machine learning model input.

[0084] 3D reconstruction algorithm: Adopts multi-view stereo vision (MVS) + structured light scanning fusion algorithm, first generates point cloud through MVS, and then uses structured light data to optimize texture details.

[0085] Curvature calculation: Based on the differential geometry method, the mean curvature and Gaussian curvature are calculated through the surface point cloud normal vector.

[0086] Defect recognition model: Use the pre-trained U-Net semantic segmentation model to identify defect areas such as honeycombs and segregation.

[0087] The general process is referred to step S3A0 to step S3D0 and will not be described in detail here.

[0088] In step S400, the feature vector is input into a pre-trained multi-task neural network, and the spatiotemporal features are integrated with the rheological algorithm to identify various working performance parameters.

[0089] Among them, pre-trained multi-task neural network: a deep learning model that has been trained in advance and can output multiple work performance parameters at the same time, such as the convolutional-recurrent neural network (CNN-RNN) that combines spatiotemporal features.

[0090] Spatiotemporal features: Composite features that integrate the temporal dimension (such as the dynamic changes during the mixing phase) and the spatial dimension (such as three-dimensional geometric structure). Rheological algorithms: Mathematical models based on fluid mechanics theory that describe the flow characteristics of concrete, such as the Herschel-Bulkley model (which characterizes the relationship between shear stress and shear rate). Performance parameters: Indicators that measure the construction performance of concrete, including viscosity, shear stress, yield strength, and thixotropy index.

[0091] For the specific process, please refer to steps S410 to S430.

[0092] In step S500, ensemble learning is used to fuse at least three recognition results, and parameters are verified based on the basic equations of fluid dynamics through a fluid simulation platform.

[0093] Ensemble learning: A method that improves accuracy by combining the predictions of multiple models. Here, a voting method is used to fuse at least three independent identification results (e.g., taking the average of three neural network inference results or a majority vote). Fluid simulation platform: A software tool that simulates the flow behavior of concrete based on fluid dynamics equations, such as OpenFOAM (an open-source computational fluid dynamics platform). Fundamental fluid dynamics equations: The Navier-Stokes equations (NS equations) describe fluid motion and are used to verify the rationality of concrete performance parameters (e.g., verifying slump values ​​through simulated slump processes).

[0094] Integration strategy: Use simple averaging to fuse multiple recognition results (such as taking the arithmetic average of three viscosity measurements).

[0095] Simulation verification rules: A physical model of concrete collapse / flow was established in OpenFOAM. After inputting the identification parameters, the simulation results (such as diffusion diameter and flow velocity distribution) were compared with the actual image / video data. The error threshold was set to ±5%.

[0096] Ensemble learning results: The performance parameters (e.g., viscosity, yield strength) output from step S400 are independently inferred at least three times using the same model, or predicted using three differently initialized neural network models. The ensemble results are calculated using a simple averaging method to reduce the chance of errors from a single model.

[0097] Fluid simulation verification: A simulation domain for the concrete paste is constructed in OpenFOAM. Integrated parameters are input and flow processes (such as slumping and pumping) are simulated using the NS equations. Simulation results (such as the final slump shape and flow velocity curve) are compared with actual image / video data, and the absolute and relative errors of key indicators (such as slump and diffusion diameter) are calculated.

[0098] If the error is within the range of ±5%, the parameter is deemed valid; if it exceeds the threshold, return to step S300 or S400 to re-extract features or adjust model parameters.

[0099] Step S600: Based on the verification results, a non-dominated sorting genetic algorithm is used to generate a mix ratio optimization suggestion including material components with construction performance and material cost as the multi-objective optimization function.

[0100] Among them, the non-dominated sorting genetic algorithm (NSGA-II) is a multi-objective optimization algorithm that generates a Pareto front solution set by simulating the biological evolution process. It is used here to balance the construction performance of concrete (such as slump, strength) and material cost. The multi-objective optimization function takes the construction performance (such as the slump value that meets the standard requirements) and the material cost (the total price of raw materials such as cement and aggregate) as the optimization objectives, and constructs a mathematical expression (such as Mix ratio optimization suggestions: Adjustment plans based on material components (cement, sand, stone, admixtures, etc.) to meet performance requirements and optimize costs.

[0101] The general process is as follows:

[0102] Data preparation: Collect historical mix ratio data (e.g., cement usage 300 kg / m³, sand ratio 40%), corresponding construction performance parameters (slump 180 mm, compressive strength 30 MPa), material prices (cement 500 yuan / ton, sand 100 yuan / ton), and environmental parameters (temperature 25°C, humidity 60%) to build a database.

[0103] Impact weight analysis: Use the random forest algorithm to train the model to determine the impact weight of each material component (such as cement and fly ash content) on performance (such as strength) and cost (for example, for every 10kg / m³ increase in cement, the strength increases by 2MPa and the cost increases by 5 yuan / m³).

[0104] Constraint generation: Set performance constraints (slump 180-220mm) based on the current construction scenario (such as pumping construction), and update the cost constraint boundary (target cost ≤ 400 yuan / m³) based on real-time material prices (such as a 10% increase in admixture prices).

[0105] Multi-objective optimization: Initialize the population (e.g., randomly generate 50 mix ratio schemes), iteratively optimize (crossover, mutation, selection) through the NSGA-II algorithm, and generate a Pareto front solution set (e.g., 10 solutions that balance performance and cost).

[0106] Scheme screening and recommendation: Based on the Shapley value analysis, the contribution of the material components in each group of schemes is analyzed (for example, replacing cement with fly ash can reduce costs but may affect early strength). Combined with user preferences (such as prioritizing cost control), the hierarchical analysis method is used to sort and recommend the optimal scheme (for example, cement 280kg / m³ + fly ash 70kg / m³, cost 380 yuan / m³, slump 190mm).

[0107] Before step S400, it is also possible to consider adding a real-time monitoring and compensation mechanism for the dynamic changes of environmental variables. Environmental parameters (such as temperature, humidity, and air pressure) can be modeled using the long short-term memory network (LSTM) in deep learning to predict their impact on the performance of concrete. Environmental compensation factors can be incorporated into the feature vector to improve the stability of measurement accuracy in complex environments. The specific steps are as follows:

[0108] Step S301 , collecting environmental parameters in real time through sensors, including but not limited to ambient temperature, humidity, air pressure, light intensity and wind speed.

[0109] Environmental parameters: external conditions that affect the working performance of concrete, sensor network: a monitoring system composed of multiple types of sensors distributed at the construction site.

[0110] Step S302 : Using a long short-term memory network (LSTM) to model the time series data of environmental parameters, and predicting the impact of environmental parameters at different time points on the working performance of concrete.

[0111] Among them, the Long Short-Term Memory (LSTM) network is a recurrent neural network that uses a gating mechanism to process long-term dependencies in sequential data. Time series data is chronologically arranged observations of environmental parameters (such as changes in temperature and humidity over time). The predictive model is a mathematical model that learns the changing patterns of environmental parameters based on historical data and predicts their impact on concrete performance.

[0112] General process: 1. Data preparation: Obtain the environmental parameter time series from step S301, with a 48-hour window size. 2. Label construction: Use historical concrete performance data (slump, expansion, etc.) as labels. Time alignment: Associate the environmental parameters with the concrete performance changes after 2 hours.

[0113] 3. Model Training: Use a sliding window to generate training samples (e.g., environmental parameters from time t-48 to time t predict performance at time t+2). Early Stopping Strategy: Stop training if validation set loss does not decrease for five consecutive rounds. 4. Prediction Deployment: Input the current environmental parameter sequence in real time and output a prediction of performance changes over the next two hours.

[0114] Step S303: constructing environmental compensation factors, by analyzing the correlation between environmental parameters and concrete working performance in a large amount of historical data, establishing a mapping relationship, and generating an environmental factor matrix for characteristic vector compensation.

[0115] Environmental compensation factor: A correction factor that quantifies the effect of environmental parameters on concrete performance. Correlation analysis: A statistical analysis of the correlation between environmental parameters and concrete performance.

[0116] The general process is as follows: 1. Data Collection: Integrate the environmental parameters from step S301 with historical concrete performance records (≥5000 data sets). 2. Feature Engineering: Extract statistical features of environmental parameters (such as the daily rate of temperature change and the cumulative effect of humidity). Calculate the relative changes in performance indicators (such as the percentage of slump loss). 3. Model Training: Input: Environmental parameter vectors T, H, P, I, and W. Output: Compensation factor matrix (3×3, corresponding to correction factors for slump, spread, and setting time). Cross-validation: k=5 to ensure model generalization. 4. Compensation Factor Generation: Input the LSTM prediction results (step S302) into the mapping model to generate the environmental factor matrix.

[0117] Step S304 : multiplying the original eigenvector by the environmental factor matrix element by element to obtain an environmentally compensated eigenvector, which can reflect the actual impact of environmental variables on the working performance of concrete.

[0118] Among them, element-by-element multiplication: multiply the eigenvector by the corresponding position element of the environmental factor matrix to achieve environmental compensation.

[0119] Step S305: input the environment-compensated feature vector into a pre-trained multi-task neural network.

[0120] Reference Figure 2 , through the pre-trained multimodal vision model to identify the standard reference objects in the image, and use the feature matching algorithm to establish the proportional mapping relationship between the image pixel coordinate system and the physical world coordinate system, including:

[0121] In step S210, a reference object with a standard size is defined in the imaging area, the reference object outline is identified using the target detection module of the large model, the reference object and the background are separated by the instance segmentation algorithm, the pixel coordinate range is extracted and the spatial position relationship is calculated.

[0122] Imaging area: The range covered by the camera's field of view, i.e., the physical space observable in the image. Standard size reference object: A calibration object with known precise physical dimensions, such as a square calibration plate with a side length of 100 mm or a standard slumping bucket with a height of 300 mm (in accordance with GB / T50080).

[0123] Object detection module: A functional module in a multimodal vision model used to identify the location and category of objects in an image, such as the detection head of YOLOv8. Instance segmentation algorithm: An algorithm that separates each independent object in an image from the background and generates a contour mask, such as MaskR-CNN. Pixel coordinate range: The two-dimensional coordinate boundary of the reference object in the image (such as the upper left pixel coordinates). , lower right pixel coordinates Spatial position relationship: the relative position of the reference object in the image (such as the vertical distance from the imaging plane and the horizontal offset).

[0124] The general process is as follows: 1. Place a reference object: Fix a standard slumping bucket within the camera's field of view, ensuring that its physical dimensions are known (e.g., bucket height 300mm) and that it is completely in frame. 2. Detect outlines: Use the YOLOv8n model to identify the slumping bucket and output bounding box coordinates (e.g., x1=100, y1=200, x2=300, y2=500). 3. Segment background: Use Mask R-CNN to perform instance segmentation on the slumping bucket, generate a binary mask, and separate the bucket body from the concrete. 4. Extract coordinates: Obtain the bucket body's pixel coordinate range from the mask (height = 300 pixels, width = 200 pixels) and record its spatial position (e.g., the center of the bottom circle is 500mm from the camera).

[0125] Step S220 , performing SIFT feature extraction on the key feature points of the reference object, performing feature matching with a pre-stored standard model, optimizing the matching results based on the RANSAC algorithm, and calculating the pixel-physical size ratio coefficient.

[0126] SIFT feature extraction: Scale-Invariant Feature Transform (SIFT) extracts feature descriptors of key points in an image, invariant to rotation, scale, and illumination. Feature matching: Compares the extracted SIFT features with the features of a pre-stored standard model to find similar feature point pairs. RANSAC algorithm: Random Sample Consensus (RANDOM SAMPLE CONSENSUS) iteratively removes mismatched points and retains correctly matched pairs. Pixel-to-Physical Size Ratio: The actual physical length corresponding to one pixel in an image (e.g., 1 pixel = 0.5 mm) is calculated by comparing the actual size of the reference object with the pixel size.

[0127] The general process:

[0128] 1. SIFT feature extraction: For the image of the reference object (such as the collapsed barrel) obtained by segmentation in step S210, use OpenCV's SIFT algorithm to extract key points (such as the corner points of the barrel edge) and its 128-dimensional feature descriptor. 2. Feature matching: Perform kNN matching (k=2) between the image features and the features of the pre-stored standard model, and retain matching pairs with a distance ratio less than 0.7 (such as: ifm.distance<0.7*n.distance:keep_match). 3. RANSAC optimization: Calculate the homography matrix through the RANSAC algorithm and eliminate incorrect matching points that do not conform to the geometric transformation (number of iterations = 1000). 4. Calculate the scale coefficient: Select two points with known physical distances in the standard model (such as the top surface diameter of the collapsed barrel is 100mm) and measure their pixel distance in the image (such as 200 pixels). Calculate the scale coefficient: .

[0129] In step S230, the imaging system is calibrated with the Zhang calibration method to perform internal parameter calibration. A three-dimensional reconstruction model is established in combination with the spatial constraint relationship of multiple reference objects. Sub-pixel measurement of concrete parameters is achieved based on the pixel-physical size ratio coefficient.

[0130] Zhang's calibration method, based on a planar checkerboard calibration plate, calculates camera intrinsic parameters (focal length, principal point, and distortion coefficients) from multiple images taken at different angles. A 3D reconstruction model combines camera intrinsic and extrinsic parameters with pixel-to-physical scaling to map 2D image points to a mathematical model of 3D physical space. Sub-pixel measurement technology achieves measurement accuracy exceeding pixel resolution (e.g., 0.1 pixel accuracy) and uses interpolation algorithms to achieve more precise coordinate positioning.

[0131] The general process is as follows: 1. Camera intrinsic parameter calibration: Use Zhang's calibration method to take at least 10 images of the checkerboard calibration plate at different angles. Calculate the camera intrinsic parameter matrix (focal length fx / fy, principal point cx / cy) and distortion coefficient (k1 / k2 / p1 / p2 / k3) through OpenCV. 2. Construction of a three-dimensional reconstruction model: Combined with the pixel-physical scale coefficient obtained in step S220 (such as k=0.5mm / pixel), the mapping relationship between the image pixel coordinates and the three-dimensional physical coordinates is established through the DLT algorithm. 3. Sub-pixel measurement: Perform sub-pixel positioning (accuracy 0.1 pixel) of the concrete edge points (such as the edge of the top surface of the slump bucket), and calculate the actual physical size (such as the slump value) based on the scale coefficient k.

[0132] In the standard slump test scenario, the pre-trained multimodal visual model is used to identify the standard reference objects in the image, and the feature matching algorithm is used to establish the proportional mapping relationship between the image pixel coordinate system and the physical world coordinate system, including:

[0133] Step S2A0: Define a standard slump bucket with a preset height placed in the imaging area as a geometric reference object, preset the vertical distance between the center of its bottom circle and the imaging plane as a fixed value, set a circular marking point with a diameter of d on the surface of the bucket as a feature positioning reference, use the target detection module of the multimodal vision large model to identify the slump bucket outline, separate the bucket body from the concrete slump body through the instance segmentation algorithm, and extract the pixel coordinate range of the bucket body and the two-dimensional pixel coordinates of the marking point. .

[0134] The standard slumping bucket is a truncated cone-shaped calibration object (upper diameter 100mm, lower diameter 200mm, height 300mm) that complies with GB / T50080. Ring markers are circular markers with a known diameter d (e.g., 10mm) placed on the surface of the bucket for precise positioning. The object detection module uses a YOLOv8-based object positioning model to output the bounding box coordinates of the slumping bucket. The instance segmentation algorithm uses a pixel-level segmentation model based on Mask R-CNN to separate the bucket from the concrete.

[0135] Step S2B0: Based on the pre-trained monocular vision depth estimation model, the depth value of the slumping bucket mark point is predicted, and the three-dimensional physical coordinates of the mark point are calculated in combination with the camera intrinsic parameter matrix. , construct the barrel space coordinate system.

[0136] Monocular vision depth estimation model: A deep learning model that predicts the depth of an object (the vertical distance from the camera) using only a single image, such as the pre-trained MonoDepth2 model (based on the ResNet backbone network). Depth value: The Z-axis coordinate of the marker in the camera coordinate system (unit: mm), representing the distance between the object and the camera. Camera intrinsic parameter matrix: A matrix that describes the internal geometric characteristics of the camera, including the focal length. , principal point coordinates , obtained through Zhang's calibration method. 3D physical coordinates: the coordinates of the marker point in the world coordinate system , used to construct the barrel space coordinate system.

[0137] The general process is as follows: 1. Depth value prediction: Input the two-dimensional pixel coordinates of the marker point extracted in step S2A0 into the MonoDepth2 model to predict the depth value Z of the point (such as 480mm, indicating that the marker point is 480mm away from the camera).

[0138] 3D coordinate calculation: Combined with the camera intrinsic matrix (such as ), the two-dimensional pixel coordinates are converted to Convert to three-dimensional physical coordinates , as follows: , assuming that the origin of the world coordinate system is the center of the bottom of the barrel, and the Z axis is aligned with the camera optical axis.

[0139] Constructing the barrel coordinate system: With the center of the barrel bottom as the origin, establish a world coordinate system (with the X-axis horizontally pointing to the right, the Y-axis vertically upward, and the Z-axis pointing to the camera), and use the three-dimensional coordinates of the marker points to calibrate the direction of the coordinate system.

[0140] Step S2C0: perform SIFT feature extraction on the edge points of the top surface of the collapsed barrel, match them with the pre-stored standard barrel 3D model, eliminate mismatched points based on the RANSAC algorithm, and calculate the ratio coefficient k between the pixel distance and the actual physical distance.

[0141] RANSAC algorithm: It iteratively removes mismatched points, retaining feature matches that match the barrel's geometry (e.g., top edge point correspondence). Scale factor k: The ratio of pixel distance to actual physical distance (e.g., 1 pixel = 0.5 mm), used to convert image size to actual size.

[0142] The general process is as follows: 1. Feature extraction: For the barrel image constructed in step S2B0, use the SIFT algorithm to extract the top surface edge point features (such as the 5 corner points on the circumference of the barrel mouth). Extract the SIFT features of the corresponding edge points from the pre-stored standard three-dimensional model (such as the two end points of the top surface diameter). 2. Feature matching and optimization: Use the FLANN matcher to find matching pairs of image features and model features, and retain reliable matches with a distance ratio less than 0.7 (such as 3 pairs of valid matches). Use the RANSAC algorithm to eliminate incorrect matching points to ensure that the matching pairs meet the circular geometric constraints of the barrel top surface. 3. Calculate the proportional coefficient k: Measure the pixel distance of the matching points in the image (such as 200 pixels). It is known that the top surface diameter of the standard model is 100mm, and it is calculated: .

[0143] In the standard slump test scenario, the extracted morphological characteristic parameters include:

[0144] In step S3A0, the spatial semantic understanding capability of the pre-trained multimodal visual large model is used to locate the pixel coordinates of the highest point of the concrete collapse body. Based on the proportional coefficient k between the pixel distance and the actual physical distance, combined with the bucket space coordinate system, the vertical distance between the highest point of the collapse body and the top surface of the collapse bucket is calculated to obtain the slump value.

[0145] Spatial semantic understanding capability: The ability of a pre-trained large model to analyze the three-dimensional spatial relationships (such as height and position) of objects in an image is achieved using the DETR3D model. Slump value: The vertical distance (unit: mm) between the highest point of a concrete slump and the top surface of the slump bucket, reflecting the concrete's fluidity. The bucket spatial coordinate system: A three-dimensional coordinate system with the Z axis pointing vertically upward, with the center of the slump bucket bottom as the origin (constructed in step S2B0).

[0146] The general process is as follows: 1. Positioning the highest pixel: Input the image processed in step S2C0 into the DETR3D model to obtain the three-dimensional coordinates of the collapse point cloud . Filter the points corresponding to the maximum value of Z coordinate (such as ) as the highest point. 2. Coordinate conversion and distance calculation: Convert the pixel coordinates of the highest point into physical coordinates through the barrel space coordinate system (step S2B0) . Knowing that the Z coordinate of the top surface of the slump bucket is 300mm (bucket height), calculate the vertical distance: slump value = ( -300) mm.

[0147] In step S3B0, the edge of the collapsed concrete is segmented at the pixel level using an instance segmentation algorithm to identify the edge contour of the concrete diffusion area. Based on the scale factor k, the edge contour of the concrete in the pixel coordinate system is converted into a contour curve in the physical coordinate system, and the maximum diameter of the concrete diffusion in the direction perpendicular to the axis of the collapse bucket is calculated.

[0148] Instance segmentation algorithm: This algorithm classifies each object (such as a concrete slump) in an image at the pixel level, using Mask R-CNN. Diffusion region edge contour: The irregular circular boundary formed on a horizontal surface after concrete slumps. Maximum diameter: The distance between the two furthest points on the concrete edge contour perpendicular to the axis of the slump bucket (unit: mm).

[0149] The general process is as follows: 1. Region Segmentation and Contour Extraction: The image processed in step S2C0 is fed into Mask R-CNN to generate a binary mask of the concrete collapse. The pixel coordinates of the contour points along the edge of the mask are extracted using OpenCV (e.g., [(x1, y1), (x2, y2), ...]). 2. Coordinate Conversion and Diameter Calculation: Using the scaling factor k (step S2C0), the pixel coordinates are converted to physical coordinates (e.g., x1' = x1 × k). The convex hull of the physical coordinate point set is calculated, and the distance between the two farthest points on the convex hull (e.g., 250 mm) is used as the maximum diameter.

[0150] In step S3C0, the surface of the concrete collapse body is reconstructed at the sub-pixel level in three dimensions. The normal vector of each surface point is calculated based on the differential geometry method. By presetting the curvature threshold, the key curvature feature points of the collapse body surface are extracted, the average curvature and Gaussian curvature are calculated, and the degree of morphological change of the collapse body surface is quantified.

[0151] Sub-pixel 3D reconstruction: A surface reconstruction technique with higher precision than a single pixel (e.g., 0.1 pixel), achieved using the StereoDR-GAN model. Differential geometry: A mathematical method for calculating normal vectors and curvature using the first- and second-order differential forms of a surface. Mean curvature: The average of the two principal curvatures of a surface at a point, reflecting the local curvature. Gaussian curvature: The product of the two principal curvatures of a surface at a point, characterizing the surface type (elliptical / hyperbolic / parabolic).

[0152] The general process is as follows: 1. Sub-pixel reconstruction: Input the edge contour image from step S3B0 into StereoDR-GAN to generate a dense point cloud (resolution 0.5mm). 2. Normal vector and curvature calculation: For each point (such as p1, p2, ...), use PCA to analyze the covariance matrix of its neighborhood point set and calculate the normal vector. Fitting a quadratic surface , calculate the average curvature H and Gaussian curvature K according to the surface parameters. 3. Feature point extraction: filter to meet and The points are taken as key curvature feature points (such as mutation points at the edge of the collapse body).

[0153] In step S3D0, the image semantic segmentation capability of the multimodal visual large model is used to identify defective areas on the surface of the collapsed body. The uniformity and molding quality of the concrete are evaluated by calculating the area ratio and distribution density of the defective areas.

[0154] Image semantic segmentation: This technology classifies each pixel in an image into specific semantic categories (such as honeycomb, segregation, and normal areas), using a U-Net model. Honeycomb defects: Honeycomb-like holes formed on the concrete surface due to unexpelled air bubbles. Segregation defects: Areas where aggregate and paste separate due to uneven distribution of concrete components. Area ratio: The percentage of pixels in the defective area relative to the total number of pixels in the inspection area (%). Distribution density: The number of defective areas per unit area (per mm²).

[0155] Segmentation model: U-Net (ResNet34 encoder, pre-trained on a concrete defect dataset) is used. Post-processing algorithm: Connected domain analysis is used to identify independent defect regions. Quantification rule: Area ratio = number of defective pixels / total number of pixels × 100%; Distribution density = number of defect regions / total area.

[0156] The general process is as follows: 1. Defect segmentation: U-Net classifies the image and generates a three-channel probability map of honeycomb, segregation, and normal. The threshold of 0.5 is used to determine the defect area. 2. Region extraction: Probability Figure 2 After quantization, the 8-neighborhood connectivity rule is used to mark independent defective areas. 3. Quantitative evaluation: Count the number of defective pixels and areas, and calculate the area ratio and distribution density using the proportional coefficient k (step S2C0).

[0157] During the concrete mixing, transportation, or pouring process, based on the spatial semantic understanding capability of the multimodal visual large model, combined with multi-view stereo vision and structured light scanning technology, the 3D geometric structure of the concrete paste is reconstructed, morphological feature parameters are extracted, and feature vectors are formed, including:

[0158] Step S310 : Based on the multi-view stereo vision technology, the segmented concrete slurry is 3D reconstructed to construct a point cloud model of its surface.

[0159] Among them, multi-view stereo vision technology: a method for reconstructing 3D structures from images from multiple camera perspectives, implemented using the COLMAP open-source framework. Point cloud model: a geometric model of an object's surface represented by a large number of 3D coordinate points. Feature point matching: the process of identifying the same physical point in different images, using the SIFT feature + RANSAC algorithm.

[0160] The general process is as follows:

[0161] 1. Image Acquisition: Use at least three industrial cameras (resolution 1920×1080, frame rate 30fps) to simultaneously capture the concrete paste from different angles. 2. Feature Extraction and Matching: Extract SIFT feature points from each image (e.g., 5,000 key points per image). Use the FLANN matcher to establish feature correspondences between different images, and use the RANSAC algorithm to remove mismatched points. 3. Camera Calibration and Triangulation: Calculate the camera intrinsic parameter matrix (e.g., focal length fx = 2800 pixels) using the Zhang calibration method. Based on the feature matching points and camera parameters, calculate the 3D coordinates of each feature point through triangulation. 4. Point Cloud Generation: Use COLMAP's dense reconstruction module to generate a dense point cloud of the concrete paste (point density ≥ 1,000 points / cm²) based on multi-view consistency.

[0162] In step S320 , high-precision texture information of the concrete surface is obtained through structured light scanning technology, and is integrated with the point cloud model to obtain a more accurate three-dimensional geometric structure.

[0163] Structured light scanning technology uses a Gray code structured light scheme to project a known pattern (such as sinusoidal stripes) onto an object's surface and calculate 3D coordinates based on the deformed pattern. Texture information includes surface color, grayscale, and detailed features (such as aggregate particle distribution). Point cloud fusion aligns and matches the high-precision texture coordinates acquired by structured light with the 3D coordinates of a multi-view point cloud.

[0164] Structured light coding: Projects Gray code sequence (8-bit encoding) with a decoding accuracy of 0.1mm.

[0165] Phase unwrapping: Least squares phase unwrapping algorithm is used (error ≤ 0.2π).

[0166] Point cloud registration: Align multi-source point clouds based on the ICP (Iterative Closest Point) algorithm (convergence threshold 0.05 mm).

[0167] The general process is as follows: 1. Structured light projection and image acquisition: An 8-bit Gray code fringe sequence (16 patterns in total) is projected onto the concrete slurry surface, and a monocular camera (resolution 2560×2048) synchronously captures the deformed pattern. 2. Phase calculation and decoding: The phase value of the fringe deformation for each pixel is calculated, and the absolute phase is obtained through Gray code decoding to solve the 2D texture coordinates (u, v) of the surface point. 3. 3D coordinate calculation: Based on the calibration parameters of the structured light system (the relative position of the projector and camera), the 2D texture coordinates are converted to 3D physical coordinates (X, Y, Z) with an accuracy of 0.1mm. 4. Point cloud fusion: The high-precision point cloud (texture point cloud) acquired by structured light is iteratively aligned with the multi-view point cloud in step S310 using the ICP algorithm to generate a fused 3D model (error ≤ 0.3mm).

[0168] In step S330 , the spatial semantic understanding capability of the multimodal visual large model is used to analyze the three-dimensional geometric structure and identify the morphological features of the concrete paste at different positions.

[0169] Spatial semantic understanding: The model's ability to identify distinct semantic regions within a 3D geometric structure (e.g., aggregate, paste, pores). Morphological characteristics: The geometric properties of the concrete paste surface (e.g., curvature, pore distribution, particle arrangement).

[0170] Semantic segmentation model: Use PointNet++ (MSG architecture, input point cloud is sampled to 1024 points).

[0171] Feature extraction rules: 1. Curvature threshold: 0.01 (distinguishing between flat and curved surfaces). 2. Hole detection: Radius threshold: 5mm (identifying holes with a diameter ≥ 5mm). 3. Aggregate identification: Volume threshold: 100mm³ (distinguishing between coarse and fine aggregate).

[0172] The general process is as follows: 1. Point cloud preprocessing: Downsample the fused point cloud from step S320 to 1024 points, and calculate the normal vector and local curvature of each point. 2. Semantic segmentation: Use PointNet++ to classify each point (label: slurry, coarse aggregate, fine aggregate, hole), and output a semantic mask. 3. Feature recognition: Based on the semantic mask, filter out hole regions (radius ≥ 5mm) and calculate their number, volume, and distribution density. Identify aggregate particles (volume ≥ 100mm³) and calculate their size distribution and spatial arrangement.

[0173] In step S340 , if the morphological feature is a continuous curve change of the surface, the curve is modeled by a mathematical fitting method, and the curvature and slope of the curve are extracted to quantify the degree of change of the curve.

[0174] Continuous curves: Smooth curves present on the surface of concrete paste (e.g., collapse edges, flow paths). Mathematical fitting methods: Approximate discrete points using parameterized curve models. Curvature: A metric describing the local curvature of a curve (unit: 1 / mm). Slope: The tangent of the angle between the tangent line of a curve at a specific point and the horizontal plane.

[0175] Curve fitting model: cubic B-spline curve (number of control points = 10, order = 3) was used.

[0176] Parameter Optimization Algorithm: Based on the least squares method (Levenberg-Marquardt iteration, convergence threshold 1e-6). Feature Extraction Rules: 1. Curvature Calculation: Estimated using second-order derivatives (sampling interval 5mm). 2. Slope Calculation: Tangent angles are calculated every 10mm along the curve.

[0177] General process: 1. Point cloud sampling: From the continuous curve region identified in step S330, uniformly sample 100 points (with spacing of approximately 2 mm). 2. B-spline fitting: Initialize 10 control points and iteratively optimize the B-spline curve to approximate the sampled points (error ≤ 0.5 mm) using the least squares method. 3. Curvature and slope calculation: Calculate the curvature of the fitted curve at 5 mm intervals (e.g., a maximum curvature of 0.031 / mm represents a bend with a radius of approximately 33 mm). Calculate the angle between the tangent and the horizontal plane every 10 mm along the curve and convert it to a slope (e.g., a slope of 0.5 represents an angle of approximately 26.6°).

[0178] In step S350 , if the morphological feature is an arc shape at a corner, a method based on a geometric model is used to accurately describe the shape and calculate the radius, center position, and length of the arc.

[0179] The arc shape at the corners is the arc transition area formed by the flow or collapse of the concrete paste surface. The geometric model is a parametric representation of the arc based on mathematical equations (center, radius, starting angle, and ending angle).

[0180] Arc fitting algorithm: Least squares circle fitting (Pratt algorithm, highly resistant to noise) is used. Parameter optimization: The optimal circle center (x0, y0, z0) and radius R are determined through Levenberg-Marquardt iteration. Quality assessment: The mean square error (MSE) between the point and the fitted circle is calculated, with a threshold of 0.5 mm.

[0181] The general process is as follows: 1. Arc point cloud extraction: From the corner regions identified in step S330, select a set of points (approximately 200-300 points) with a continuous curvature change of ≥0.02. 2. Arc fitting: Initialize the estimated value of the circle center (such as the geometric center of the point set) and iteratively optimize it using the Pratt algorithm to calculate the optimal arc parameters. 3. Parameter calculation: Arc radius R: Directly output the fitting result (e.g., R = 12.5mm). Circle center position: 3D coordinates (x0, y0, z0), with an accuracy of ±0.3mm. Arc length: Calculated using the start and end angles (e.g., arc length = 25.7mm).

[0182] Step S360: combining and screening the extracted morphological characteristic parameters to form a characteristic vector related to the rheological properties of concrete.

[0183] Morphological characteristic parameters are quantitative indicators extracted from the three-dimensional structure of concrete (such as curvature, radius, and arc length). Eigenvectors are numerical sequences formed by arranging multiple morphological characteristic parameters in a fixed order, used to characterize the rheological properties of concrete. Combination screening uses correlation analysis and principal component analysis to eliminate redundant parameters and retain key features.

[0184] Feature Screening Algorithm: Principal Component Analysis (PCA) was used (retaining components with a cumulative contribution rate ≥ 90%). Combination Rule: Arrange in the order of "geometric features (curvature, radius) → topological features (arc length, number of holes) → distribution features (aggregate density)". Normalization Method: Min-Max Normalization was used (scaling parameters to the [0, 1] interval).

[0185] General process: 1. Parameter Collection: Summarize all morphological features extracted in steps S330-S350 (e.g., 10 curvature values, 5 arc radii). 2. Correlation Analysis and Screening: Calculate the Pearson correlation coefficient between parameters and remove redundant parameters with correlations > 0.8 (e.g., remove two highly correlated curvature indices). 3. Principal Component Analysis: Perform PCA on the remaining parameters, retaining principal components with a cumulative contribution ≥ 90% (e.g., filter out 8 key components from 15 parameters). 4. Feature Combination and Normalization: Arrange key parameters according to preset rules to form a feature vector (e.g., [0.3, 0.7, 0.1, ...]). Perform Min-Max normalization on each parameter to eliminate dimensionality effects.

[0186] The feature vector is input into a pre-trained multi-task neural network, which integrates spatiotemporal features and combines them with rheological algorithms to identify various performance parameters including:

[0187] In step S410, a recurrent neural network is used to process the time series data of the feature vectors to extract the dynamic change characteristics of concrete during the mixing, transportation, and pouring stages. At the same time, a convolutional neural network is used to perform a secondary extraction of the spatial characteristics of the three-dimensional geometric structure of concrete, and the time dimension feature vectors and the space dimension feature vectors are tensor-joined to construct a composite feature representation containing spatiotemporal information.

[0188] Recurrent Neural Networks (RNNs): A neural network that processes sequential data (such as time series) and uses LSTMs (Long Short-Term Memory) to capture long-term dependencies. Convolutional Neural Networks (CNNs): A network that extracts features from image or spatial data, using ResNet18 as its backbone. Tensor concatenation: This concatenates feature vectors of different dimensions along specified axes to form a composite tensor containing spatiotemporal information. Spatiotemporal features: Dynamic changes in the temporal dimension (such as viscosity changes during stirring) and geometric properties in the spatial dimension (such as the curvature of the slurry surface).

[0189] Temporal feature extraction: An LSTM network (256 hidden units, stacked in two layers) was used. Spatial feature extraction: ResNet18 (pre-trained on ImageNet, parameters of the first three layers frozen) was used. Feature fusion rule: The temporal feature vector (dimension 1×128) output by the LSTM was concatenated with the spatial feature vector (dimension 512×1) output by the ResNet18 along the channel axis to form a 1×640 composite tensor.

[0190] The general process is as follows: 1. Temporal feature extraction: The feature vectors from step S360 are sequenced in chronological order (e.g., before mixing, during transportation, and during pouring). This is input into an LSTM network, which outputs a 128-dimensional temporal feature vector. 2. Secondary spatial feature extraction: The 3D concrete structure point cloud is projected into a depth map, which is input into a ResNet18 network to extract a 512-dimensional spatial feature vector. 3. Feature fusion: The temporal and spatial feature vectors are concatenated along the channel axis to form a 640-dimensional composite feature tensor containing temporal and spatial information, which serves as input to subsequent networks.

[0191] In step S420 , the constitutive equations of the Bingham model and the Herschel-Bulkley model are introduced as physical constraints in the fully connected layer of the multi-task neural network to map the spatiotemporal composite feature vector to the rheological parameter space.

[0192] Bingham model: constitutive equation describing plastic fluid ( ), including yield strength and plastic viscosity . Herschel-Bulkley model: an extension of the Bingham model ( ), increasing the power law exponent n, making it suitable for non-Newtonian fluids. Physical constraints: Embedding rheological constitutive equations into the neural network ensures that the predictions conform to physical laws. Rheological parameter space: A multidimensional space composed of parameters such as yield strength and viscosity.

[0193] Network architecture: A physical constraint layer is added after the fully connected layer, and the Bingham / Herschel-Bulkley equation is used as the regularization term of the loss function.

[0194] Constraint method: 1. Explicit constraint: add equation constraints in the output layer (such as ).

[0195] Implicit constraint: Add the residual of the equation to the loss function (such as ).

[0196] Weight distribution: physical loss weight , data loss weight .

[0197] The general process is as follows: 1. Feature mapping: The spatiotemporal composite feature vector (640 dimensions) of step S410 is mapped to the intermediate feature space (256 dimensions) through the fully connected layer. Physical constraint embedding: The predicted value of the Bingham / Herschel-Bulkley equation (such as ). 2. Compare the predicted value with the calculated value of the constitutive equation (such as ) is added to the total loss function: 3. Parameter optimization: Optimize network parameters and physical parameters simultaneously through back propagation (such as ), to ensure that the prediction results are consistent with the rheological laws.

[0198] In step S430, the viscosity, shear stress, yield strength, and thixotropy index of the concrete are calculated in parallel through the multi-branch structure of the neural network output layer. A correction function is constructed in combination with real-time environmental parameters to compensate for the initial recognition results and output the final accurate working performance parameter measurement values.

[0199] The multi-branch structure includes multiple independent branches in the neural network output layer, each predicting different performance parameters (such as viscosity and yield strength). Real-time environmental parameters include variables such as ambient temperature and humidity that affect concrete performance. Correction function is a mathematical model that compensates the initial prediction based on environmental parameters. Performance parameters reflect concrete construction performance, including viscosity, shear stress, yield strength, and thixotropy index.

[0200] Among them, the output structure: uses a four-branch fully connected network, each branch outputs a single parameter prediction value. Correction function: uses a linear regression model (formula: ), where T is temperature, H is humidity, and a, b, and c are coefficients. Compensation rule: The linear regression coefficients are calibrated based on historical data, with the error threshold set to 10%.

[0201] The general process is as follows: 1. Parallel parameter prediction: The output features (256 dimensions) of step S420 are input into a four-branch fully connected network to calculate the initial predicted values ​​of viscosity, shear stress, yield strength, and thixotropy index respectively. 2. Environmental parameter acquisition: Real-time acquisition of ambient temperature (range 5-40°C) and humidity (range 30%-90%) data. 3. Result correction: For each predicted parameter, compensation is performed using a linear regression correction function (e.g., for every 1°C increase in temperature, the viscosity decreases by 3%). 4. Output calibration: Check whether the corrected parameters exceed the reasonable range (e.g., a yield strength > 100Pa is considered abnormal). If so, the default compensation coefficient adjustment is enabled.

[0202] Based on the verification results, a non-dominated sorting genetic algorithm was used to generate the following mix ratio optimization recommendations, including the following material components, with construction performance and material cost as the multi-objective optimization function:

[0203] Step S610: Collect and obtain historical concrete mix ratio data, actual construction performance parameters, material market price fluctuation data and corresponding construction environment parameters to build a multi-dimensional material property and cost database.

[0204] Multidimensional Database: A multidimensional dataset containing material composition, construction performance, cost, and environmental parameters. Construction Performance Parameters: Indicators that characterize concrete workability, such as slump, expansion, and compressive strength. Material Market Price Fluctuation Data: Real-time or historical prices for raw materials such as cement, aggregates, and admixtures.

[0205] Historical mix ratio data: Extracted from a construction record database (e.g., Excel spreadsheets, SQL databases), containing over 1,000 mix ratio records. Construction performance parameters: Automatically collected through the real-time measurement system in step S430, or manually entered from laboratory reports. Material price data: Connect to building material supplier APIs (e.g., China Cement Network) to obtain real-time prices, or manually update a CSV file monthly. Environmental parameters: Obtain historical data such as temperature and humidity from a meteorological database (e.g., OpenWeatherMap).

[0206] Database Construction Rules: Data Cleaning: Outliers (e.g., data with slump > 300mm) were eliminated, and missing values ​​were interpolated using the mean. Dimensional Design: Material Components: Cement, sand, stone, water, admixture type and dosage. Performance Parameters: Slump, expansion, initial setting time, 28-day compressive strength. Cost Dimensions: Material unit price, transportation costs, processing fees. Environmental Parameters: Temperature, humidity, and wind speed during construction. Storage Structure: A relational database (e.g., MySQL) was used, with table-based storage.

[0207] In step S620 , principal component analysis and random forest algorithm are used to explore the influence weights of each material component on construction performance and cost, and a quantitative association model is established.

[0208] Principal Component Analysis (PCA): A dimensionality reduction technique that transforms high-dimensional data into low-dimensional principal components through linear transformation, extracting the data's primary variation characteristics. Random Forest: An ensemble learning method that constructs multiple decision trees to score the importance of input features. Impact Weight: A numerical value that quantifies the relative importance of each material component to construction performance and cost.

[0209] The general process is as follows: 1. Data preprocessing: Extract the feature matrix (material components) and target vector (construction performance, cost) from the database in step S610. Standardize continuous variables (StandardScaler) and one-hot encode categorical variables. 2. PCA dimensionality reduction: Apply PCA to the feature matrix to reduce the original dimensions (e.g., 20 material parameters) to the number of principal components (e.g., 15). Calculate the contribution rate of each principal component and determine the number of principal components to retain. 3. Random forest training: Construct random forest models for construction performance and cost separately. Obtain the importance score of each material component (range 0-1) through the feature_importances_ attribute. 4. Weight quantization: Normalize the importance scores to obtain the weight of each material component's impact on different objectives. For example, cement has a weight of 0.65 on strength and a weight of 0.78 on cost.

[0210] In step S630, performance constraints are dynamically generated according to the special requirements of different construction scenarios, and the cost constraint boundaries are updated in real time in combination with the material market price fluctuation data to form an adaptive constraint system.

[0211] Performance constraints: Lower and upper limits on concrete performance indicators (e.g., slump ≥ 180mm) set based on construction scenarios (e.g., pumping, large-volume pouring). Cost constraints: Upper limits on the total cost budget (e.g., unit cost ≤ 500 yuan) to account for material price fluctuations. Adaptive constraint system: An intelligent system that dynamically adjusts constraints based on real-time data.

[0212] The general process is as follows: 1. Construction scenario identification: The user enters information such as construction type (e.g., underwater pouring) and structural location (e.g., beams). 2. Performance constraint generation: The user matches corresponding constraints from a pre-set rule library (e.g., underwater pouring requires segregation resistance ≥ 95%). 3. Cost boundary calculation: The user obtains current material prices (step S610) and calculates the benchmark mix cost. The cost cap is dynamically adjusted based on price volatility (e.g., the average price over the past seven days ± standard deviation). 4. Constraint verification: The user checks whether historical mix data meets the constraints. If not, the constraints are relaxed by 10% and re-verified.

[0213] In step S640, an elite retention strategy and an adaptive crossover mutation operator are introduced into the non-dominated sorting genetic algorithm. During the population evolution process, the crossover mutation probability is dynamically adjusted according to the convergence degree of the current solution set. Combined with congestion calculation and Pareto front analysis, multiple groups of non-inferior solution optimization schemes covering the balance range between construction performance and cost are screened out.

[0214] Among them, the non-dominated sorting genetic algorithm (NSGA-II) is a multi-objective optimization algorithm that selects Pareto optimal solutions through non-dominated sorting and congestion distance. This algorithm takes the construction performance objective function and the material cost objective function as optimization objects, wherein: Construction performance objective function: quantified as the weighted score of each performance index (such as slump, compressive strength), predicted by the quantitative association model of step S430 and normalized to the interval [0,1]. The higher the score, the better the performance. Material cost objective function: directly calculated as the sum of the product of the amount of cement, sand, stone and other materials and the real-time unit price (unit: yuan / m³). The lower the value, the lower the cost. Multi-objective optimization function integration: convert the dual objectives into a minimization problem that can be handled by NSGA-II, that is: ,in, By taking the negative, we can turn the performance maximization problem into a minimization problem. Direct retention is the cost minimization goal. Elite retention strategy: Directly retain the best individuals in the parent generation to avoid the loss of excellent solutions during the evolution process. Adaptive crossover mutation operator: Dynamically adjust the crossover probability according to the degree of population convergence ( ) and mutation probability ( ). Pareto front: a set of non-inferior solutions where any improvement in the solution will lead to a deterioration of other objectives.

[0215] Algorithm configuration: 1. Population size: 100 individuals (each group of individuals represents a combination ratio scheme). 2. Number of iterations: 200 generations. 3. Crossover probability: Initial , as the convergence decreases to 0.6. 4. Mutation probability: Initial , as the convergence degree increases to 0.3. 4. Crowding calculation: The density of each solution in the target space is calculated using Euclidean distance.

[0216] The general process is as follows: 1. Initialize the population: randomly generate 100 mix ratio schemes (satisfying the constraints of step S630). 2. Fitness evaluation: calculate the performance score (such as slump, strength) and cost score of each scheme, and convert them into the minimization objective of NSGA-II: , add penalty items to solutions that violate the constraints (such as deducting performance points and increasing cost values). 3. Non-dominated sorting: Divide the population into non-dominated solution sets of different levels (level 1 is the optimal solution set). 4. Crowding calculation: Evaluate the distribution density of solutions in the target space through Euclidean distance, prioritize solutions in sparse areas, and maintain diversity. 5. Elite retention and evolution: Directly retain 10% of the best individuals to enter the next generation. Adaptive crossover mutation: Increase the mutation probability (up to 0.3) when the convergence degree is high, and increase the crossover probability (up to 0.8) when it is low. 6. Convergence judgment: If the optimal solution set changes by ≤5% for 10 consecutive generations, stop evolution.

[0217] Adaptive adjustment rules:

[0218] Crossover probability adjustment: ; Among them, f is the individual fitness and f' is the average fitness of the current generation.

[0219] Mutation probability adjustment: .

[0220] Output: Generates a set of non-inferior solutions (e.g., 20 mix ratios). Each mix ratio includes: 1. Material composition ratios (cement: sand: stone: water: admixtures). 2. Predicted construction performance (slump, strength, etc.). 3. Cost estimates. 4. Pareto rank and congestion value.

[0221] Step S650: Establish a feature attribution model and a quantitative association model based on Shapley values, perform interpretability analysis on each set of optimization solutions, and quantify the contribution of each material component to performance improvement and cost change.

[0222] Shapley value: A method used in game theory to fairly distribute the benefits of cooperation, quantifying the contribution of each feature to the model's predictions. Feature attribution model: A model that uses Shapley values ​​to calculate the impact of each material component on performance and cost. Contribution quantification: Normalizes each feature's Shapley value to obtain a relative importance percentage.

[0223] Calculation method: KernelSHAP algorithm is used (sampling times 500, error tolerance 0.01).

[0224] Basic model: Use the multi-task neural network in step S430 as the prediction model.

[0225] Feature range: 15 material component parameters including cement content, sand ratio, admixture dosage, etc.

[0226] The general process:

[0227] Model loading: Load the neural network model trained in step S430 (for predicting viscosity, strength, cost, etc.). 2. Data preparation: Select 10 representative mix ratio solutions from the non-inferior solutions generated in step S640. 3. Shapley value calculation: For each solution, calculate the Shapley value of each material component for each performance indicator and cost using the KernelSHAP algorithm. For example, calculate the contribution of cement content to strength. 4. Contribution normalization: Normalize the Shapley value to a percentage: 5. Result visualization: Generate radar charts or bar charts to show the contribution of each material component to different targets.

[0228] Step S660: Based on the user's different preferences for cost control and performance assurance, the optimization schemes are weighted and sorted using the analytic hierarchy process to recommend the optimal concrete mix ratio that meets the user's specific needs.

[0229] Analytic Hierarchy Process (AHP): This method breaks down complex decision-making problems into a multi-level structure and determines weights through pairwise comparison. Preferences: The user's emphasis on objectives such as cost control and strength assurance (e.g., a cost weight of 0.6, a strength weight of 0.4). Weighted Ranking: Calculates a comprehensive score based on the weight of each objective and prioritizes the options.

[0230] The general process is as follows:

[0231] Goal hierarchy: Top-level goal: Select the optimal mix ratio. Criteria level: Cost, strength, workability, durability, etc. (selected by the user). Solution level: The non-inferior solutions generated in step S640 (e.g., 20 mix ratios).

[0232] Preference matrix generation: The user inputs the target importance ratio (e.g. cost: intensity = 3:2).

[0233] Constructing a judgment matrix (example): .

[0234] Weight calculation and verification: Calculate and normalize the feature vector to obtain the weight vector (e.g., cost 0.43, strength 0.31, workability 0.26). Verify the consistency ratio (CR = 0.05 < 0.1, passed the test).

[0235] Scheme scoring and ranking: For each scheme, calculate the weighted score: Arrange in descending order of scores and recommend the top 3 solutions.

[0236] Reconstructing the 3D geometry of the concrete paste includes:

[0237] In step SA00, target detection and semantic segmentation of a large multimodal visual model are used to identify key objects and mark their semantic categories. Simultaneously, a monocular depth estimation algorithm is used to generate an initial three-dimensional structure with semantic labels. Key objects include concrete slurry and slump buckets.

[0238] Among them, multimodal vision large models: pre-trained models that fuse multiple sources of information, such as image and depth (such as CLIP and SegmentAnything). Monocular depth estimation algorithms: methods that predict scene depth from a single RGB image. Semantic labeling: assigning a category to each point in a 3D point cloud (such as concrete slurry or slump bucket).

[0239] The general process is as follows: 1. Image acquisition: Use an industrial camera (resolution ≥1920×1080) to capture multi-angle RGB images of the concrete paste. 2. Semantic segmentation: Input the image into SAM to generate an instance segmentation mask (labels: concrete paste, slump bucket, background). 3. Depth estimation: Use MiDaS to predict the depth map, adjusting the resolution to match the SAM output. 4. Point cloud generation: Combine the RGB image with the depth map and generate the initial point cloud using the pinhole imaging model. Assign a label to each point based on the SAM semantic mask (e.g., label 1 = concrete paste). 5. Quality control: Filter depth outliers (e.g., points with depth values ​​> 5m). Dilate the semantic boundary (by 3 pixels) to reduce edge classification errors.

[0240] In step SB00, a sparse point cloud of the scene is constructed using the structure-from-motion algorithm, and a dense point cloud is generated by combining multi-view stereo vision technology to obtain a high-precision geometric model.

[0241] Among them, Structure from Motion (SfM) is an algorithm that recovers the camera pose and 3D structure of a scene from a multi-angle image sequence. Sparse point cloud is a low-density 3D point set containing only feature points. Dense point cloud is a high-density 3D point set generated through stereo matching that fully expresses the surface details of an object.

[0242] The general process is as follows: 1. Image feature extraction: Extract SIFT features from the multi-angle images in step SA00 (approximately 5,000-10,000 key points per image). 2. Feature matching: Establish feature correspondences using a brute-force matcher and use the RANSAC algorithm to remove false matches (with an inlier threshold of 2px). 3. Camera pose estimation: Use COLMAP's incremental SfM to gradually add images starting from the initial two views to estimate camera internal and external parameters. 4. Sparse point cloud generation: Calculate the 3D coordinates of feature points through triangulation to generate a sparse point cloud (approximately 100,000-500,000 points). 5. Dense point cloud reconstruction: Based on the estimated camera pose, use the PMVS algorithm to generate a dense point cloud (point density ≥ 1,000 points / cm²).

[0243] In step SC00, the confidence of semantic and geometric data is calculated based on the Bayesian reasoning framework, and the large model semantic understanding and traditional algorithm geometric accuracy are adaptively integrated in different regions through a dynamic weighting strategy.

[0244] Among them, the Bayesian reasoning framework: a statistical method based on the probability model to integrate multi-source data and calculate the posterior probability Confidence calculation: A quantitative indicator (ranging from 0 to 1) that evaluates the degree of match between semantic labels and geometric data. Dynamic weighting strategy: Adaptively adjusts the fusion weight of semantic and geometric data based on regional characteristics.

[0245] The general process:

[0246] Data preparation: Input the semantic labels (e.g., concrete paste, slump bucket) from step SA00 and the dense point cloud from step SB00.

[0247] Confidence calculation: 1. Semantic confidence: Directly uses the mask quality score output by SAM (range 0-1). 2. Geometric confidence: Calculates the local curvature of the point cloud (estimated via PCA), assigning high confidence to flat areas. Evaluates point cloud density, assigning higher confidence to dense areas (threshold ≥ 800 points / cm²).

[0248] Dynamic weighted fusion: Increase the semantic weight (w≈0.7) in edge areas (such as the interface between the slurry and the slump bucket), and increase the geometric weight (w≈0.3) in smooth areas (such as the slurry surface).

[0249] Result optimization: Conditional random field (CRF) post-processing is performed on the fused labels to eliminate small area noise.

[0250] Output: Semantic point cloud with confidence: Each point contains 3D coordinates, semantic labels, and fusion confidence. Weight distribution map: Visualizes the fusion weight of semantics and geometry in each region.

[0251] In step SD00, the large model Transformer module is used to infer spatial relationships to correct geometric contradictions, optimize semantic boundaries with geometric dimensions, and iterate in both directions.

[0252] The Transformer module, a deep learning architecture based on a self-attention mechanism, is used to reason about long-range spatial relationships in point clouds. Geometric inconsistencies are detected, such as areas where semantic boundaries and geometric structure are inconsistent (e.g., semantic segmentation mislabels a plane as a curved surface). Bidirectional iterations alternately optimize semantic boundaries and geometric dimensions to gradually improve model accuracy.

[0253] The general process is as follows: 1. Point cloud preprocessing: Divide the point cloud from step SC00 into 5cm×5cm local blocks and extract geometric features (normals, curvature). 2. Spatial relationship reasoning: Use PointTransformer to calculate the relationship between points and identify the spatial constraints of the semantic region (e.g., the slump bucket should be perpendicular to the ground). 3. Contradiction detection and positioning: Compare semantic boundaries with geometric features and mark conflicting areas (e.g., the discontinuity at the junction of the slurry and the slump bucket). 4. Bidirectional optimization: Semantic correction: Adjust the semantic boundary based on the geometric structure (e.g., reclassify the surface area from the "plane" category to the "slurry" category). Geometric optimization: Fit a more accurate geometric model based on semantic information (e.g., fit the slump bucket with a cylindrical surface). 5. Iterative convergence: Repeat the above steps until the proportion of conflicting areas drops below a threshold (e.g., <5%).

[0254] Output: Corrected semantic point cloud: geometric contradictions are eliminated, and the semantic boundary is consistent with the geometric structure. Optimization parameter log: records the changes in the conflicting areas and the amount of geometric parameter adjustments in each iteration.

[0255] In step SE00, the fusion model is smoothed using the moving least squares method, and a complete triangular mesh is generated using the Poisson surface reconstruction algorithm.

[0256] Moving Least Squares (MLS): A method for smoothing point clouds using locally weighted polynomial fitting. Poisson Surface Reconstruction: A 3D surface reconstruction algorithm based on implicit function solving, generating a watertight triangulated mesh. Surface Smoothing: A filtering operation that reduces point cloud noise while preserving geometric features.

[0257] The general process is as follows: 1. Point cloud preprocessing: Downsample the point cloud of step SD00 (voxel size 0.005m) to reduce the amount of calculation. 2. Normal estimation and orientation: Estimate the normal direction of each point through PCA, and use the propagation algorithm to ensure global consistency of the normal. 3. MLS smoothing: Apply the moving least squares method, iterate 3 times, and project the point onto the local fitting surface each time. Adaptively adjust the search radius (use a smaller radius for areas with large curvature). 4. Poisson surface reconstruction: Construct an indicator function and extract the isosurface through the MarchingCubes algorithm. Crop invalid areas (such as the inside of the collapse bucket) to retain only the concrete paste surface. 5. Mesh optimization: Subdivide the generated triangular mesh (Loop subdivision) to improve surface details. Simplify the mesh (QuadricEdgeCollapse), and control the number of facets between 100k-200k.

[0258] In step SF00, the iterative closest point algorithm is used to align the model with the preset standard reference object to perform point cloud registration, and finally output a high-precision semantic 3D model.

[0259] The Iterative Closest Point (ICP) algorithm is an algorithm that iteratively finds the optimal rotation and translation matrix to achieve optimal alignment between two point clouds. Point cloud registration and alignment aligns the reconstructed model with a pre-set standard reference object to unify the coordinate system. A standard reference object is an object with a known precise 3D structure (such as a slump bucket or calibration plate) that defines the global coordinate system.

[0260] The general process is as follows: 1. Definition of standard reference object: preset a standard three-dimensional model of the slump bucket (known dimensions: upper diameter 100mm, lower diameter 200mm, height 300mm). 2. Feature extraction and coarse registration: segment the slump bucket area from the reconstructed model of step SE00. Use the FPFH (FastPointFeatureHistograms) feature descriptor to estimate the initial transformation matrix through the SAC-IA algorithm. 3. Fine registration optimization: apply Point-to-PlaneICP, with the standard slump bucket model as the target, and iteratively optimize the pose of the reconstructed model. 4. Global alignment: apply the transformation matrix calculated by ICP to the entire reconstructed model to achieve alignment with the standard coordinate system. 5. Quality assessment: calculate the average distance between the aligned model and the standard reference object (required to be ≤0.5mm). Verify the error between the geometric dimensions (such as the slump expansion diameter) and the actual measurement value (required to be ≤2%).

[0261] Based on the same inventive concept, an embodiment of the present invention provides a concrete working performance measurement system based on a multimodal visual large model, including a memory and a processor, wherein the memory stores a program that can be executed on the processor to implement the following Figures 1 to 2 Procedure for either method.

[0262] The embodiments of this specific implementation method are all preferred embodiments of the present application and are not intended to limit the scope of protection of the present application. Therefore, any equivalent changes made based on the structure, shape, and principle of the present application should be included in the scope of protection of the present application.

Claims

1. A method for measuring concrete working performance based on a multimodal visual large model, characterized in that: include: Collect multi-angle images, continuous video sequences, and environmental parameters of concrete paste at different construction stages to construct a multi-dimensional spatiotemporal dataset; The pre-trained multimodal vision model is used to identify standard reference objects in the image, and the feature matching algorithm is used to establish the proportional mapping relationship between the image pixel coordinate system and the physical world coordinate system. Based on the spatial semantic understanding capability of the multimodal visual large model, combined with multi-view stereo vision and structured light scanning technology, the 3D geometric structure of the concrete paste is reconstructed, morphological feature parameters are extracted, and feature vectors are formed. The feature vectors are input into a pre-trained multi-task neural network, which fuses the spatiotemporal features and combines them with rheological algorithms to identify various working performance parameters. Ensemble learning is used to fuse at least three identification results, and the parameters are verified based on the basic equations of fluid dynamics through a fluid simulation platform; Based on the verification results, a non-dominated sorting genetic algorithm is used to generate a mix ratio optimization suggestion including material components, taking construction performance and material cost as the multi-objective optimization function; During the concrete mixing, transportation, or pouring process, based on the spatial semantic understanding capability of the multimodal visual large model, combined with multi-view stereo vision and structured light scanning technology, the 3D geometric structure of the concrete paste is reconstructed, morphological feature parameters are extracted, and feature vectors are formed, including: Based on multi-view stereo vision technology, the segmented concrete paste is 3D reconstructed to construct a point cloud model of its surface; The high-precision texture information of the concrete surface is obtained through structured light scanning technology and integrated with the point cloud model to obtain a more accurate three-dimensional geometric structure. Leveraging the spatial semantic understanding capabilities of a multimodal visual large model, the 3D geometric structure is analyzed to identify the morphological characteristics of concrete paste at different locations. If the morphological feature is a continuous curve change on the surface, the curve is modeled by mathematical fitting methods to extract the curvature and slope of the curve to quantify the degree of change of the curve; If the morphological feature is an arc shape at the corner, a method based on a geometric model is used to accurately describe it and calculate the arc radius, center position, and arc length; The extracted morphological characteristic parameters are combined and screened to form a characteristic vector related to the rheological properties of concrete; The feature vector is input into a pre-trained multi-task neural network, which integrates spatiotemporal features and combines them with rheological algorithms to identify various performance parameters including: A recurrent neural network is used to process time series data of feature vectors to extract the dynamic characteristics of concrete during mixing, transportation, and pouring. Convolutional neural networks are also used to perform secondary extraction of the spatial characteristics of concrete's three-dimensional geometric structure. The time dimension feature vectors and the space dimension feature vectors are tensor-concatenated to construct a composite feature representation that includes spatiotemporal information. The constitutive equations of the Bingham model and the Herschel-Bulkley model are introduced as physical constraints in the fully connected layer of the multi-task neural network to map the spatiotemporal composite feature vector to the rheological parameter space. Through the multi-branch structure of the neural network output layer, the viscosity, shear stress, yield strength, and thixotropy index of concrete are calculated in parallel. Combined with real-time environmental parameters, a correction function is constructed to compensate for the initial recognition results and output the final accurate measurement values ​​of the working performance parameters.

2. The method for measuring concrete working performance based on a multimodal visual large model according to claim 1 is characterized in that: Using a pre-trained multimodal vision model to identify standard reference objects in an image, and using a feature matching algorithm to establish a proportional mapping relationship between the image pixel coordinate system and the physical world coordinate system includes: Define a reference object with a standard size preset in the imaging area, use the target detection module of the large model to identify the outline of the reference object, separate the reference object from the background through the instance segmentation algorithm, extract the pixel coordinate range and calculate its spatial position relationship; Perform SIFT feature extraction on the key feature points of the reference object, perform feature matching with the pre-stored standard model, optimize the matching results based on the RANSAC algorithm, and calculate the pixel-physical size ratio coefficient; The Zhang calibration method is used to calibrate the internal parameters of the imaging system. The three-dimensional reconstruction model is established by combining the spatial constraint relationship of multiple reference objects. The sub-pixel measurement of concrete parameters is achieved based on the pixel-physical size ratio coefficient.

3. The method for measuring concrete working performance based on a multimodal visual large model according to claim 1 is characterized in that: In the standard slump test scenario, the pre-trained multimodal visual model is used to identify the standard reference objects in the image, and the feature matching algorithm is used to establish the proportional mapping relationship between the image pixel coordinate system and the physical world coordinate system, including: A standard slump bucket with a preset height is defined as a geometric reference object placed in the imaging area. The vertical distance between the center of its bottom circle and the imaging plane is preset to a fixed value. A circular marking point with a diameter of d is set on the surface of the bucket as a feature positioning benchmark. The target detection module of the multimodal vision large model is used to identify the outline of the slump bucket. The bucket body and the concrete slump body are separated by the instance segmentation algorithm, and the pixel coordinate range of the bucket body and the two-dimensional pixel coordinates of the marking point are extracted. ; Based on the pre-trained monocular vision depth estimation model, the depth value of the slumping bucket marker is predicted, and the 3D physical coordinates of the marker are calculated by combining the camera intrinsic parameter matrix. , construct the barrel space coordinate system; SIFT features are extracted from the edge points of the top surface of the collapsed barrel and matched with the pre-stored standard barrel 3D model. The mismatched points are eliminated based on the RANSAC algorithm, and the proportional coefficient k between the pixel distance and the actual physical distance is calculated.

4. The method for measuring concrete working performance based on a multimodal visual large model according to claim 3 is characterized in that: In the standard slump test scenario, the extracted morphological characteristic parameters include: By pre-training the spatial semantic understanding capabilities of a large multimodal visual model, the pixel coordinates of the highest point of the concrete collapse are located. Based on the proportional coefficient k between the pixel distance and the actual physical distance, combined with the bucket's spatial coordinate system, the vertical distance between the highest point of the collapse and the top surface of the collapse bucket is calculated to obtain the slump value. The instance segmentation algorithm is used to segment the collapsed concrete edge at the pixel level and identify the edge contour of the concrete diffusion area. Based on the scale factor k, the concrete edge contour in the pixel coordinate system is converted into a contour curve in the physical coordinate system. The maximum diameter of the concrete diffusion in the direction perpendicular to the axis of the collapse bucket is calculated. The surface of the concrete collapse body is reconstructed at the sub-pixel level in 3D. The normal vector of each surface point is calculated based on the differential geometry method. By presetting the curvature threshold, the key curvature feature points of the collapse body surface are extracted, the mean curvature and Gaussian curvature are calculated, and the degree of morphological change of the collapse body surface is quantified. The image semantic segmentation capability of the multimodal visual large model is used to identify defective areas on the surface of the collapsed body. By calculating the area ratio and distribution density of the defective areas, the uniformity and molding quality of the concrete are evaluated.

5. The method for measuring concrete working performance based on a multimodal visual large model according to claim 1 is characterized in that Based on the verification results, a non-dominated sorting genetic algorithm is used to generate the mix ratio optimization suggestions including material components, taking construction performance and material cost as the multi-objective optimization function. Collect and obtain historical concrete mix ratio data, actual construction performance parameters, material market price fluctuation data and corresponding construction environment parameters to build a multi-dimensional material property and cost database; Using principal component analysis and random forest algorithms, we explored the impact of each material component on construction performance and cost, and established a quantitative correlation model. Dynamically generate performance constraints based on the special requirements of different construction scenarios, combine material market price fluctuation data, and update cost constraint boundaries in real time to form an adaptive constraint system; An elite retention strategy and an adaptive crossover mutation operator are introduced into the non-dominated sorting genetic algorithm. During population evolution, the crossover mutation probability is dynamically adjusted based on the convergence level of the current solution set. Combined with congestion calculation and Pareto front analysis, multiple sets of non-inferior optimization solutions covering the balance range between construction performance and cost are screened out. Establish a feature attribution model and quantitative association model based on Shapley values, conduct interpretable analysis on each set of optimization solutions, and quantify the contribution of each material component to performance improvement and cost change; Based on users' different preferences for cost control and performance assurance, the optimization schemes are weighted and ranked using the analytic hierarchy process to recommend the optimal concrete mix ratio that meets specific needs.

6. The method for measuring concrete workability based on a multimodal visual large model according to any one of claims 1 to 5, characterized in that: Reconstructing the 3D geometry of the concrete paste includes: Utilizing object detection and semantic segmentation based on a large multimodal visual model, key objects are identified and labeled with semantic categories. Simultaneously, a monocular depth estimation algorithm is used to generate an initial 3D structure with semantic labels. Key objects include concrete slurry and slumping buckets. The motion recovery structure algorithm is used to construct a sparse point cloud of the scene, and the multi-view stereo vision technology is combined to generate a dense point cloud to obtain a high-precision geometric model; The confidence level of semantic and geometric data is calculated based on a Bayesian reasoning framework. Through a dynamic weighting strategy, the semantic understanding of large models and the geometric accuracy of traditional algorithms are adaptively integrated in different regions. Use the large model Transformer module to reason about spatial relationships and correct geometric contradictions, optimize semantic boundaries with geometric dimensions, and iterate in both directions; The fusion model is smoothed using the moving least squares method, and a complete triangular mesh is generated using the Poisson surface reconstruction algorithm. Using the iterative closest point algorithm, the model is aligned with the preset standard reference object for point cloud registration, and finally a high-precision semantic 3D model is output.

7. A concrete working performance measurement system based on a multimodal visual large model, characterized in that: The invention comprises a memory, a processor and a program stored in the memory and executable on the processor, wherein the program can be loaded and executed by the processor to implement a method for measuring concrete working performance based on a multimodal visual large model as claimed in any one of claims 1 to 6.

Citation Information

Patent Citations

  • Multi-mode image fusion identification method and system based on multi-task collaborative learning

    CN119380000A