Concrete pouring process monitoring method and system based on multi-source data

By using multi-source data fusion technology, and combining visible light video and thermal infrared frames with vibration signals, a physical mapping model of vibration energy input and thermal field fluctuation response is established. This solves the problem that existing technologies cannot quantify the internal density of concrete, and enables accurate assessment and visualization of pouring quality.

CN121580084AActive Publication Date: 2026-02-27CHINA CONSTR EIGHT ENG DIV CORP LTD
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
CN202511891715.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-16
Publication Date
2026-02-27
Estimated Expiration
2045-12-16

AI Technical Summary

Technical Problem

Existing concrete pouring process monitoring solutions mainly rely on visible light computer vision technology, which makes it difficult to quantify internal density, lacks multi-source data fusion, and cannot accurately assess the uniformity of energy transfer and distribution within the concrete.

Method used

By acquiring visible light video frames, thermal infrared frames, and vibration signals of concrete pouring, the positioning and frequency analysis of the vibrator are performed. Combined with phase-locked-phase thermal imaging technology for vibration, a physical mapping model of vibration energy input, thermal field fluctuation response, and internal compaction state is established to achieve fusion evaluation of multi-source data.

Benefits of technology

It enables quantitative visualization of concrete pouring quality, effectively solving the problem that internal quality defects cannot be judged based on surface visual characteristics alone, and improving the accuracy and reliability of monitoring.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121580084A_ABST
    Figure CN121580084A_ABST
Patent Text Reader

Abstract

The invention discloses a concrete pouring process monitoring method and system based on multi-source data, relates to the field of intelligent monitoring, and particularly relates to the following steps: carrying out semantic segmentation and positioning on a pouring area and a vibrating rod by using a visible light video, and constructing an accurate space mask; meanwhile, the key vibration frequency is extracted from the vibration signal. On the basis, a vibration phase-locked thermal imaging technology is applied, the vibration frequency is used as a phase-locked reference, and harmonic energy transfer gain characteristics with the same frequency as vibration are extracted from a thermal infrared image sequence. And then temperature field uniformity evaluation is carried out by fusing the spatial position and the energy transfer characteristics, so that quantitative visualization of the concrete pouring quality is realized, and the technical problem that whether quality defects such as voids and pitted surfaces exist in the concrete or not cannot be judged only by means of surface visual characteristics is effectively solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of intelligent monitoring, and more specifically, to a method and system for monitoring the concrete pouring process based on multi-source data. Background Technology

[0002] Concrete pouring is a crucial step in building construction, and its quality directly determines the safety and durability of the building structure. During the pouring process, thorough vibration is the core method for removing air bubbles, eliminating voids, and ensuring structural density. With the development of intelligent construction technology, real-time evaluation of pouring quality using monitoring methods has become an industry consensus. Constructing a concrete pouring process monitoring solution based on multi-source data aims to overcome the limitations of single-sensor methods. By integrating multi-dimensional information such as vision, temperature, and vibration, it comprehensively perceives the complex conditions of the construction site, particularly addressing the problems of traditional manual inspections, such as strong subjectivity, narrow coverage, and delayed post-inspection. This enables precise quantification and real-time early warning of the internal density of the concrete.

[0003] However, existing concrete pouring process monitoring solutions mainly rely on visible light computer vision technology. This single-modal monitoring method has significant drawbacks in practical applications, making it difficult to truly quantify the internal density of concrete. Specifically, visible light cameras can only capture two-dimensional images or 2.5D morphology of the concrete surface, recording the apparent result of the "covering" action, rather than the state of the physical process of "compaction." Because there is no direct and stable physical mapping relationship between the pixel grayscale and texture of the concrete surface and the internal porosity and aggregate distribution, a smooth surface does not necessarily mean that internal air bubbles have been completely expelled. This easily leads to quality problems where the surface appears to be poured well, but the inside is actually honeycomb-like and pitted. In addition, key physical phenomena during the vibration process, such as the rise of internal air bubbles and thixotropy of slurry liquefaction, often occur inside the concrete or manifest as extremely weak features on the surface. In complex construction site environments, these signals are easily drowned out by noise such as changes in lighting and water reflection, making them difficult for visual algorithms to capture. More importantly, existing technologies lack a causal reasoning model that moves from surface phenomena to internal essence. Purely visual solutions ignore the crucial energy input for achieving "density"—vibration—and fail to establish a nonlinear coupling relationship between vibration energy and the compaction state. Although some solutions attempt to introduce vibration sensors or thermal imaging devices, these often operate independently, failing to effectively align and fuse the frequency characteristics of the vibration signal with the energy transfer process in the thermal infrared image. This makes it difficult to leverage the complementarity between multiple data sources to penetrate surface appearances and accurately assess the uniformity of energy transfer and distribution within the concrete.

[0004] Therefore, there is currently a lack of a monitoring scheme that can effectively integrate visible light positioning, thermal infrared energy analysis, and vibration frequency characteristics to solve the problem that visual semantics cannot quantify internal density. Summary of the Invention

[0005] To address the aforementioned problems in the existing technology, this application provides a method for monitoring the concrete pouring process based on multi-source data, comprising: Acquire visible light video frames, thermal infrared frames, and vibration signals of concrete pouring; The visible light video frames of concrete pouring are used to locate the pouring area and the vibrator to obtain the thermal infrared pouring area mask and the thermal infrared vibration influence area mask. The vibration signal is subjected to a fast Fourier transform to obtain the vibration frequency; Based on the vibration frequency, a phase-locked thermal imaging analysis of the sequence of thermal infrared frames of concrete pouring was performed to obtain a harmonic energy transfer gain map. A density score map was obtained by performing a density fusion evaluation based on temperature field uniformity and energy index on the harmonic energy transfer gain map, the thermal infrared casting area mask, and the thermal infrared vibration influence area mask. Based on the density fraction map and the thermal infrared casting zone mask, a visual heat map of casting quality is generated.

[0006] This application also provides a concrete pouring process monitoring system based on multi-source data, which includes: The data acquisition module is used to acquire visible light video frames, thermal infrared frames, and vibration signals of concrete pouring. The pouring area and vibrator positioning module is used to position the pouring area and vibrator in the visible light video frame of concrete pouring to obtain the thermal infrared pouring area mask and the thermal infrared vibration influence area mask. The vibration signal analysis module is used to perform a fast Fourier transform on the vibration signal to obtain the vibration frequency. The vibration phase-locked thermal imaging analysis module is used to perform vibration phase-locked thermal imaging analysis on the sequence of thermal infrared frames of concrete pouring based on the vibration frequency to obtain a harmonic energy transfer gain map. The density fusion evaluation module is used to perform a density fusion evaluation based on temperature field uniformity and energy index on the harmonic energy transfer gain map, the thermal infrared casting area mask and the thermal infrared vibration influence area mask to obtain a density score map. The casting quality visualization module is used to generate a casting quality visualization heat map based on the density fraction map and the thermal infrared casting area mask.

[0007] Compared with existing technologies, this application provides a method and system for monitoring the concrete pouring process based on multi-source data. Addressing the challenge that existing single-modal visual solutions cannot perceive the internal density of concrete, this application utilizes visible light video to perform semantic segmentation and localization of the pouring area and vibrator, constructing a precise spatial mask; simultaneously, it extracts key vibration frequencies from the vibration signals. Based on this, it applies phase-locked thermal imaging technology, using the vibration frequency as a phase-locked reference, to extract harmonic energy transfer gain features at the same frequency as the vibration from the thermal infrared image sequence. Since the internal density of concrete directly affects the transmission efficiency of vibration energy (manifested as thermal energy fluctuations), a physical mapping model of vibration energy input - thermal field fluctuation response - internal density state is successfully established. By fusing spatial location and energy transfer characteristics to assess temperature field uniformity, quantitative visualization of concrete pouring quality is achieved, effectively solving the technical problem that surface visual features alone cannot determine whether internal quality defects such as honeycomb or pitting exist. Attached Figure Description

[0008] The above and other objects, features and advantages of this application will become more apparent from the more detailed description of the embodiments of this application in conjunction with the accompanying drawings.

[0009] Figure 1 This is a flowchart of a concrete pouring process monitoring method based on multi-source data according to an embodiment of this application.

[0010] Figure 2 This is a schematic diagram of data flow in a concrete pouring process monitoring method based on multi-source data according to an embodiment of this application.

[0011] Figure 3 This is a flowchart of step 2 in the concrete pouring process monitoring method based on multi-source data according to an embodiment of this application.

[0012] Figure 4 This is a schematic diagram of the data flow in step 5 of the concrete pouring process monitoring method based on multi-source data according to an embodiment of this application.

[0013] Figure 5 This is a block diagram of a concrete pouring process monitoring system based on multi-source data according to an embodiment of this application. Detailed Implementation

[0014] The embodiments of this application will now be described in more detail with reference to the accompanying drawings. It should be understood that the drawings and embodiments of this application are for illustrative purposes only and are not intended to limit the scope of protection of this application.

[0015] In view of the shortcomings in the above-mentioned technical field, this application proposes a method for monitoring the concrete pouring process based on multi-source data. Figure 1This is a flowchart of a concrete pouring process monitoring method based on multi-source data according to an embodiment of this application. Figure 2 This is a schematic diagram of the data flow in the concrete pouring process monitoring method based on multi-source data according to an embodiment of this application. Figure 1 and Figure 2 As shown, the concrete pouring process monitoring method based on multi-source data according to an embodiment of this application includes: Step 1, acquiring visible light video frames, thermal infrared frames, and vibration signals of concrete pouring; Step 2, locating the pouring area and vibrator in the visible light video frames of concrete pouring to obtain a thermal infrared pouring area mask and a thermal infrared vibration influence area mask; Step 3, performing a fast Fourier transform on the vibration signal to obtain the vibration frequency; Step 4, performing phase-locked thermal imaging analysis on the sequence of thermal infrared frames of concrete pouring based on the vibration frequency to obtain a harmonic energy transfer gain map; Step 5, performing a density fusion evaluation based on temperature field uniformity and energy index on the harmonic energy transfer gain map, the thermal infrared pouring area mask, and the thermal infrared vibration influence area mask to obtain a density score map; Step 6, generating a visualization thermal map of pouring quality based on the density score map and the thermal infrared pouring area mask.

[0016] In step 1, visible light video frames, thermal infrared frames, and vibration signals of concrete pouring are acquired. It should be understood that the quality of concrete pouring is the cornerstone of ensuring the safety and durability of building structures, and internal density cannot be accurately determined solely by surface smoothness. Existing single visible light monitoring methods can only capture surface coverage, lacking the ability to perceive the internal vibration energy transfer and air bubble removal process, making it difficult to detect quality hazards such as "smooth surface, hollow interior" in a timely manner. Vibration operation is essentially a process of energy input and transfer; the mechanical energy of the vibrator is converted into heat and kinetic energy through the concrete medium, and its frequency characteristics are closely physically coupled with thermal field fluctuations. To overcome the bottleneck of visual semantics' inability to quantify internal density, a multi-dimensional information field including spatial location, heat distribution, and vibration frequency needs to be constructed. Therefore, acquiring spatiotemporally synchronized visible light video frames, thermal infrared frames, and vibration signals is the physical basis for establishing a causal model from vibration energy input to density response, and also a prerequisite for achieving cross-modal data fusion analysis.

[0017] In one embodiment of step 1, the specific processing is as follows: In the initial stage of implementing this monitoring method, a multi-source sensing hardware system needs to be deployed at the construction site and a rigorous system calibration needs to be completed. Then, a real-time data acquisition process is initiated to ensure accurate spatial alignment and strict temporal synchronization of data from different modalities. The entire implementation process is mainly divided into two stages: hardware calibration and synchronous acquisition. In the hardware calibration stage, because the imaging principles, resolutions, and field of view of standard industrial cameras and thermal infrared cameras are different, and their physical installation positions are offset, directly superimposing the two images will lead to severe parallax misalignment, making it impossible to accurately correspond to the physical state of the same spatial point. Therefore, a spatial alignment transformation matrix needs to be constructed. The spatial alignment transformation matrix is ​​a mathematical model that can map pixels in the visible light image coordinate system to their corresponding positions in the thermal infrared image coordinate system. It is represented as a 3×3 homography matrix or a projection matrix containing rotation and translation parameters. To obtain this matrix, after equipment deployment, a specially designed checkerboard calibration board, such as an 8×11 aluminum substrate with 30mm side length for each corner point, is required. This board is placed within the shared field of view of both the visible light and thermal infrared cameras, ensuring uniform heating and a significant temperature difference between the calibration board and the background to enable clear imaging by the thermal infrared camera. Multiple sets of synchronized image pairs are acquired by adjusting the angle and position of the calibration board. A corner detection algorithm is used to extract the checkerboard corner coordinates from the visible light image and the corresponding corner coordinates from the thermal infrared image. Then, based on a pinhole camera model and the least squares method, the intrinsic parameter matrices (including focal length and principal point coordinates) and distortion coefficients of each camera are calculated. Based on this, the rotation matrix and translation vector from the visible light camera coordinate system to the thermal infrared camera coordinate system are calculated, and finally combined to generate a spatial alignment transformation matrix. Once this matrix is ​​established, for any given visible light pixel coordinates, its precise position in the thermal infrared image can be calculated through matrix multiplication, thus providing a geometric reference for subsequent cross-modal mask transformations.

[0018] After calibration, the synchronous acquisition phase begins. A hardware trigger mode is used to ensure data timing consistency. A standard industrial camera, a high-sensitivity thermal infrared camera, and vibration sensors (such as high-frequency accelerometers or contact microphones) mounted on the vibratory tamping bar are connected to the same hardware trigger controller. The controller sends frequency-divided synchronization pulse signals according to preset sampling requirements. When the rising edge of the trigger signal arrives, each sensor simultaneously performs acquisition actions. Specifically, the standard industrial camera acquires and outputs visible light video frames of concrete pouring. These video frames are three-channel color digital images with a resolution set to 1920×1080 pixels. The data format is an uncompressed raw bitmap, which can clearly record the texture details of the construction site, the position of the vibratory tamping bar, and the worker's operating trajectory. The sampling frequency is set to 30 frames per second, serving as the basic input for subsequent semantic segmentation and target localization. Simultaneously, the thermal infrared camera acquires and outputs thermal infrared frames of concrete pouring. To satisfy the Nyquist sampling theorem to capture the high-frequency vibration thermal response and its harmonic components, the thermal infrared camera operates in high-speed acquisition mode, with a sampling frequency set to 600 frames per second (600Hz) or higher. The thermal infrared frame is a single-channel grayscale matrix or pseudo-color image generated based on the thermal radiation intensity of the object's surface, with a resolution set to 640×512 pixels. Each pixel value represents the absolute temperature or radiation energy value of the corresponding spatial location. A highly sensitive thermal infrared sensor (e.g., with a noise equivalent temperature difference (NETD) of less than 30 millikrvin) can capture minute temperature fluctuations caused by friction and energy dissipation during vibration, which contain crucial information about energy transfer within the concrete. The third data source acquired simultaneously is the vibration signal. A vibration sensor is tightly attached to the drive head or flexible shaft wall of the vibrator, continuously recording the time-domain waveform of vibration acceleration at a high-frequency sampling rate such as 4000 Hz or higher. This vibration signal is a one-dimensional time-series array, recording the changes in mechanical vibration intensity of the vibrator during operation. Due to hardware synchronization triggering, each visible light image and each thermal infrared image precisely correspond to a vibration signal segment within a specific time window.

[0019] In step 2, the concrete pouring visible light video frames are used to locate the pouring area and the vibrator to obtain thermal infrared masks for the pouring area and the vibration-affected area. Correspondingly, while thermal infrared imaging can capture minute changes in the temperature field and reveal the energy transfer process during concrete pouring, it suffers from inherent disadvantages such as low spatial resolution, blurred edges, and a lack of texture detail. Relying solely on thermal infrared images makes it difficult to accurately distinguish between objects of different materials, such as the concrete surface, reinforcing steel frame, wooden formwork, and construction workers. This makes subsequent energy analysis highly susceptible to interference from thermal radiation in non-working areas. To achieve accurate assessment of vibration quality, this application strictly limits the analysis scope to the effective fresh concrete area and the vibrator's effective range. Visible light video, with its high resolution and rich texture semantic information, is the best data source for scene analysis. Therefore, high-precision semantic segmentation and target localization of visible light video frames can accurately identify physical operation boundaries. These boundaries can then be mapped onto the thermal infrared field of view through cross-modal coordinate transformation, thereby generating an accurate spatial mask. This ensures that the compaction assessment algorithm is only executed on effective concrete entities and vibration operation surfaces, eliminating the interference of environmental background noise and improving the signal-to-noise ratio and reliability of monitoring results.

[0020] Figure 3 This is a flowchart of step 2 in the concrete pouring process monitoring method based on multi-source data according to an embodiment of this application. Figure 3 As shown, in one embodiment of step 2, the concrete pouring visible light video frame is used to locate the pouring area and the vibrator to obtain a thermal infrared pouring area mask and a thermal infrared vibration influence area mask, including: step 21, performing semantic segmentation and extraction of the concrete pouring area in the concrete pouring visible light video frame to obtain a pouring area mask; step 22, performing multi-scale target detection and localization of the vibrator in the concrete pouring visible light video frame to obtain a vibration influence area mask; step 23, performing cross-modal coordinate transformation on the vibration influence area mask and the pouring area mask based on the spatial alignment transformation matrix to obtain a thermal infrared pouring area mask and a thermal infrared vibration influence area mask.

[0021] In the above implementation, step 2 is specifically processed as follows: First, step 21 is executed. This step uses the visible light video frame of concrete pouring, which was synchronously acquired and output in step 1, as input data. Since the original visible light video frame acquired on-site usually has a high resolution, such as 1920×1080 pixels, directly inputting it into a deep neural network would lead to excessive computation and memory overflow. Therefore, image standardization processing is required first. The original red, green, and blue three-channel color image is adjusted to the standard input size required by the pre-trained model using a bilinear interpolation algorithm, for example, set to 512×512 pixels. Subsequently, the adjusted image data is normalized, mapping the pixel values ​​from the integer range of [0,255] to the floating-point range of [0,1], and the mean of the dataset is subtracted and divided by the standard deviation to accelerate model convergence. For example, if a certain channel of the input image is in the coordinate... The pixel value at that location is Normalized pixel values The calculation formula is ,in and These are the preset mean and standard deviation parameters, such as... =0.485, =0.229. The standardized image tensor is input into a semantic segmentation network using the deep learning-based U-Net architecture. The U-Net architecture consists of an encoder (downsampling path) and a decoder (upsampling path), forming a unique "U"-shaped structure. The encoder is responsible for extracting multi-scale features of the image. It contains multiple repeating convolutional blocks, each consisting of two 3×3 convolutional layers, a batch normalization layer, and a ReLU activation function, followed by a 2×2 max-pooling layer for downsampling. As the number of layers increases, the spatial size of the feature map is halved (e.g., from 512×512 to 256×256), while the number of channels doubles (e.g., from 64 to 128), thus effectively extracting high-level texture and semantic features that distinguish different materials such as concrete, steel bars, and formwork. The decoder is responsible for restoring the spatial resolution of the feature map. It performs upsampling through a 2×2 deconvolution operation and performs skip connections and channel concatenation between the upsampled feature map and the corresponding layer feature map in the encoder. This design allows the spatial location information lost in the encoder to be reintroduced into the decoding process, enabling the network to accurately locate the edge contours of concrete areas. Notably, the weight parameters (such as the weight matrix of the convolutional kernels) and bias terms in the model are obtained through supervised learning during offline training. A dataset containing a large number of labeled images of concrete pouring sites is constructed, where each pixel of each image is manually labeled as either "fresh concrete" or "background." During training, images are input into the network, and the cross-entropy loss function between the output and the true labels is calculated. The weights and biases are iteratively updated using backpropagation and a stochastic gradient descent optimizer until the model can accurately distinguish concrete areas. During the inference phase, the network outputs a 512×512×2 classification probability map, classifying only concrete and non-concrete, where each pixel contains the probability value for each category. Finally, a binarization mask is generated. The output classification probability map is parsed to extract the probability channel representing the fresh concrete category. A confidence threshold is set. For example, 0.5, for each pixel in the image. If it belongs to fresh concrete, the predicted probability is If a pixel is selected correctly, it is marked as foreground (value set to 1 or 255); otherwise, it is marked as background (value set to 0). This generates a binary black-and-white image, the pouring area mask. In this mask, the white area precisely covers the wet joints or concrete surface of the slab being poured, while the black background area removes the surrounding steel mesh, formwork supports, and hardened concrete structure. To maintain consistency with the original video frame, the 512×512 mask image can be restored back to the original 1920×1080 resolution using nearest neighbor interpolation, serving as the final pouring area mask output.

[0022] In one embodiment of step 22, multi-scale target detection and localization of the vibratory rod are performed on the visible light video frame of concrete pouring to obtain a mask of the vibration influence zone, including: step 221, inputting the visible light video frame of concrete pouring into the target detection model to obtain the bounding box coordinates and confidence of the vibratory rod head; step 222, using the center point of the vibratory rod in the bounding box coordinates of the vibratory rod head as the center, generating a circular region descriptor according to the preset effective vibration radius; step 223, filling the circular region corresponding to the circular region descriptor with white on a black background of the same size as the visible light video frame of concrete pouring to obtain a binary mask of the influence range of the vibratory rod as the mask of the vibration influence zone.

[0023] Step 221: The target detection model used here is preferably the YOLOv8 architecture, which is renowned for its excellent real-time performance and detection accuracy in industrial scenarios. The YOLOv8 network architecture mainly consists of three core parts: the backbone network, the neck network, and the head network. The backbone network adopts the CSPDarknet structure, utilizing cross-stage local connectivity and the SiLU activation function to extract features from the input visible light video frames at multiple scales, generating feature maps of different resolutions to capture the texture, edge, and shape features of the vibratory tamping rod. The neck network adopts the PANet structure, enhancing the semantic representation of multi-scale features through a bidirectional fusion path from top to bottom and bottom to top, ensuring that the model can detect both large vibratory tamping rods nearby and smaller targets at a distance. The head network adopts a decoupled head design, separating the classification task (determining whether it is a vibratory tamping rod) from the regression task (predicting bounding box coordinates) to improve localization accuracy. Before the model is put into use, it needs to undergo an offline training process. A dataset of concrete pouring site images with varying lighting, angles, and occlusion levels was constructed, and the vibratory tampers were manually labeled. During training, the images were input into the model, and the difference between the predicted bounding boxes and the ground truth bounding boxes was calculated using CIoU and DFL loss functions. Error was backpropagated using stochastic gradient descent (SGD) or the Adam optimizer, iteratively updating the weight parameters and biases in the network until the model converged. During online monitoring, real-time captured visible light video frames with a resolution of 1920×1080 pixels were input into the model. The model performed forward inference and output the detection results. For example, if the model identifies a vibratory tamper in the current frame with a confidence level of 0.95, it outputs its bounding box coordinates. .in, and Represents the x and y coordinates of the bounding box center point in the image coordinate system (e.g. =960, =540), and These represent the width and height of the bounding box, respectively (e.g., ...). =40, =300 pixels). These values ​​directly reflect the pixel-level position of the vibratory rod at the current moment.

[0024] Step 222: This step maps the physical laws of vibration to the image pixel space. The preset effective vibration radius... These parameters are not fixed but dynamically configured based on the physical properties of the concrete on site (especially its slump). A higher slump results in better fluidity, slower attenuation of the vibration waves, and a larger effective radius of action; conversely, a lower slump leads to a smaller effective radius of action for dry, stiff concrete. This parameter is typically stored in the system's configuration parameter file. For example, for pumped concrete with a slump of 180 mm, the effective vibration radius is set based on construction specifications and empirical data. =400 mm. To generate the corresponding region on the image, the physical radius needs to be converted to a pixel radius. This relies on the camera intrinsic parameters and the distance information between the camera and the casting plane obtained in the calibration phase of step 1, or the pixel equivalent ratio calculated based on on-site reference objects. (Unit: pixels / mm). For example, the pixel equivalent ratio of the current scene. =0.5, then the effective vibration radius on the image is calculated as follows: =400 × 0.5 = 200 pixels. Based on the determined center coordinates. and the calculated pixel radius This generates a mathematically circular region descriptor. This descriptor defines all regions that satisfy the inequalities. pixel set These points spatially constitute the current effective operating coverage area of ​​the vibratory rod.

[0025] Step 223: First, create a single-channel matrix in memory with the same size as the visible light video frame of the concrete pouring, i.e., 1920×1080, and set all elements to their initial values ​​of 0, representing a black background. This completely black background simulates the initial state without any vibration. Next, using the circular region descriptor generated in step 222, perform pixel filling operations on this matrix. Traverse each pixel coordinate in the matrix, or use the drawing functions of the graphics processing library, to select the pixels falling within the specified area. Center of the circle All pixel values ​​within a circular area with radius are set to 1, representing a white foreground. After this processing, the vibration-affected area mask is obtained. This is a binary image where white circular areas visually indicate the current effective working range of the vibrator, while black background areas are excluded from the evaluation range. This mask not only reflects the position of the vibrator but also quantifies its effective range by incorporating the rheological properties of the concrete. For example, in the masked image, a circular area with a radius of 200 pixels around coordinates (960, 540) is white, and the rest is black.

[0026] Finally, step 23 is implemented. First, perspective transformation is performed. This process uses a spatial alignment transformation matrix to perform point-by-point spatial mapping on each non-zero pixel in the input mask, representing the white pixels of the effective foreground area. The spatial alignment transformation matrix is ​​denoted as... It is a 3×3 homography matrix that describes the geometric mapping from the visible light image plane to the thermal infrared image plane through rotation, translation, scaling, and projection transformation parameters. For any foreground pixel in the visible light mask image, its coordinates are denoted as . Construct homogeneous coordinate vectors The corresponding position in the thermal infrared coordinate system is calculated using matrix multiplication. The specific calculation logic follows the formula: in, These are the transformed homogeneous coordinates. This is to obtain the actual pixel coordinates on the thermal infrared image plane. Perspective division is required, that is... For example, the coordinates of the center point of the vibratory rod detected in step 22 in the visible light image are (960, 540), which belongs to the white area of ​​the vibration influence zone mask. (e.g., spatial alignment transformation matrix) The parameters in the equation reflect that the field of view of the thermal infrared camera is approximately one-third that of the visible light camera, and that there is a certain degree of translation. Substituting these coordinates into the above formula, we can obtain the transformed coordinates. ≈(320.5, 256.2). This means that the vibratory rod area at the center of the visible light image corresponds to a specific location slightly to the left or right of the center in the thermal infrared image. The program performs this calculation on all white pixels in the mask, thus forming a set of floating-point target coordinates. Interpolation and resampling are then performed. The target coordinates are calculated after perspective transformation. Typically, the values ​​are non-integer floating-point values. Due to the significant difference in resolution between the two sensors (mapping from high resolution to low resolution), the directly mapped point cloud may be unevenly distributed on the thermal infrared pixel grid, resulting in holes or overlaps. To generate a continuous and regular binary mask image, nearest-neighbor interpolation must be used to remap the transformed coordinates onto the pixel grid of the concrete pouring thermal infrared frame. The computational logic is expressed as follows: In this formula, The thermal infrared casting area mask representing the output Or thermal infrared vibration affected area mask , For use as a cover for the pouring area or the area affected by vibration; These are the pixel coordinates in the visible light image coordinate system, specifically the pixel coordinates of the mask in the pouring area or the mask in the vibration-affected area. These are the pixel coordinates in the thermal infrared image coordinate system, i.e., the pixel coordinates of the thermal infrared casting area mask or the thermal infrared vibration influence area mask. This represents the interpolation and resampling function. Nearest neighbor interpolation is chosen instead of bilinear interpolation because the mask is a binary image (0 or 1), representing a yes or no logical state, and intermediate values ​​like 0.5 have no physical meaning. Nearest neighbor interpolation finds the closest source pixel value in the transformation map for each pixel in the thermal infrared grid, thus maintaining the sharpness of the mask edges and avoiding blurry grayscale transitions. Finally, the output is generated. After the above transformation and resampling, two new binary images are generated: a thermal infrared casting area mask and a thermal infrared vibration influence area mask. The resolution of these two masks is strictly aligned with the 640×512 specification of the thermal infrared camera. In the thermal infrared vibration influence area mask, a white circular area with a radius of 200 pixels in the original 1920×1080 image may become a white circular area with a radius of approximately 66 pixels after scaling and resampling, and its center position precisely corresponds to the insertion point of the vibrator rod exhibiting high-temperature characteristics in the thermal infrared frame. Similarly, the thermal infrared casting area mask accurately delineates the area covered by fresh concrete in the thermal infrared field of view.

[0027] In step 3, a Fast Fourier Transform (FFT) is performed on the vibration signal to obtain the vibration frequency. It is understandable that the vibration operation during concrete pouring is essentially a process of using mechanical vibration waves to force the coarse aggregates in the concrete mixture to rearrange and expel air bubbles. This process is accompanied by significant energy conversion and dissipation. Part of the mechanical energy of the vibrator is converted into the kinetic energy of the flowing concrete, and another part is converted into heat energy due to the internal friction between the aggregates, resulting in a slight local temperature rise. Due to the complex environment at the construction site, solar radiation, air convection, and the hydration heat of the concrete itself will create high-intensity background noise in the thermal infrared image, masking the subtle thermal features generated by vibration friction. To extract the effective thermal response caused solely by vibration from the noisy thermal image, signal processing technology based on the phase-locked loop (PLL) principle is required. The core of PLL analysis lies in having a precise reference frequency. Because the vibration frequency of the vibrator is affected by factors such as load changes, voltage fluctuations, or hydraulic instability during actual operation, it is not constant but fluctuates dynamically around the rated value. If the theoretical rated frequency is directly used as a reference, it will lead to mismatch in the PLL calculation and an inability to accurately reconstruct the energy distribution. Therefore, frequency domain analysis of the real-time acquired vibration signal is to accurately capture the actual working frequency of the vibrator at the current moment, and to provide a unique frequency reference for demodulating the same frequency energy characteristics in the thermal infrared sequence from the background noise, thereby establishing a strict physical correspondence between mechanical vibration and thermal field fluctuation.

[0028] In one embodiment of step 3, the vibration signal is subjected to a fast Fourier transform to obtain the vibration frequency, including: step 31, performing a fast Fourier transform on the vibration signal to obtain a vibration frequency domain signal; step 32, performing a peak search on the vibration frequency domain signal to obtain the frequency point with the largest amplitude as the vibration frequency.

[0029] In the above embodiment, step 3 is specifically processed as follows: First, step 31 is executed. The input to this step is the one-dimensional discrete-time series acquired in step 1 by an accelerometer or microphone mounted on the vibrating rod and quantized by an analog-to-digital converter (ADC). Let the sampling frequency be... The frequency is 4000Hz. To maintain consistency with the time window for subsequent thermal infrared frame sequence analysis, vibration data of a period of time, such as 1 second, corresponding to the current thermal infrared video frame is extracted as a processing unit. If the number of sampling points within this period is N, such as N=4000 points, then the discrete-time signal can be represented as follows: Where n is the time index, ranging from 0 to N-1. Before performing the transformation, to reduce spectral leakage caused by data truncation, the original signal needs to be... Apply a window function, such as a Hanning or Hamming window. The signal after applying the window function. The calculation formula is: Subsequently, the windowed time-domain signal was processed using the Fast Fourier Transform (FFT) algorithm. Convert to frequency domain signal The FFT is an efficient algorithm for the Discrete Fourier Transform (DFT). It maps a signal from the time domain to the frequency domain, revealing the intensity of each frequency component contained in the signal. It is a complex sequence containing amplitude and phase information for each frequency point. The output of this step is the vibration frequency domain signal, which essentially describes the distribution of vibration energy at different frequencies within the current 1-second sampling window.

[0030] Next, proceed to step 32. Due to the vibration frequency domain signal output in step 31... Since it's in complex form, to intuitively compare the energy levels of each frequency component, we first need to calculate its amplitude spectrum. For each frequency index... its magnitude It can be obtained through modulo operations on complex numbers: in, and Representing complex numbers respectively The real and imaginary parts of the integers are calculated. After the calculation, a real array reflecting the intensity of each frequency component is obtained. Based on this, a peak search is performed. Due to the presence of low-frequency interference (such as vibrations below 20 Hz caused by worker movement and wind) and high-frequency electronic noise in actual construction sites, an effective search frequency band is set to improve the robustness of the detection. For example, for commonly used high-frequency insertion vibrators, their rated frequency is typically between 100 Hz and 200 Hz, corresponding to 6000 to 12000 vibrations per minute. Therefore, the search algorithm only searches for the maximum value within the index range corresponding to this physical frequency band. Frequency resolution Determined by the sampling rate and the number of sampling points, the calculation formula is as follows: In this embodiment, =4000 / 4000=1Hz. The search logic iterates through all amplitude values ​​within the effective frequency band. Find the index corresponding to the point with the largest amplitude. That is, satisfying (For all valid ranges) After finding the index corresponding to the maximum amplitude, convert it into a physical frequency value, i.e., the tamping frequency. : Specifically, if in the spectral analysis of the aforementioned 4000 sampling points, the algorithm finds that in the index Amplitude value at =120 If the index is the largest and falls within the preset valid search interval, then the calculated vibration frequency is... =120×1=120Hz. This indicates that within the current monitoring time window, the vibrator is operating at a main frequency of 120 Hz, or 120 vibrations per second.

[0031] In step 4, based on the vibration frequency, a phase-locked thermal imaging analysis of the sequence of thermal infrared frames of concrete pouring is performed to obtain a harmonic energy transfer gain map. It should be understood that the formation of concrete compaction density is essentially a process of vibration energy propagation and dissipation in a multiphase medium. This process manifests at the microscopic level as aggregate rearrangement and air bubble overflow, and at the macroscopic physical field as thermal energy fluctuations at the same frequency as the vibration source. However, due to the hysteresis of thermal diffusion effects and the low-frequency thermal background drift caused by the construction site environment (such as direct sunlight and air convection), the transient weak thermal signals generated by internal vibration friction are easily submerged, making it impossible to accurately quantify the actual energy transfer efficiency of vibration by simply relying on temperature difference analysis of a single frame of thermal infrared image. To reveal the energy transfer mechanism in the causal chain of "vibration-compaction," a phase-locked thermal imaging technique based on time-series analysis is introduced. This technique, through frequency domain analysis of a continuous thermal infrared frame sequence, uses the vibration signal as a phase-locked reference to extract the tiny temperature fluctuations at the same frequency as the vibration hidden in random thermal noise. Therefore, trend de-simplification and full-spectrum feature analysis of thermal infrared frame sequences are performed to eliminate the interference of environmental DC components and construct a full-spectrum data space that can reflect the dynamic thermal response characteristics of each pixel.

[0032] In one embodiment of step 4, based on the vibration frequency, a phase-locked thermal imaging analysis of the sequence of thermal infrared frames of concrete pouring is performed to obtain a harmonic energy transfer gain map, including: step 41, performing data stacking and trend removal on the sequence of thermal infrared frames of concrete pouring to obtain a temperature fluctuation cube; step 42, performing a pixel-by-pixel time-series signal frequency domain transformation on the temperature fluctuation cube to obtain a full-spectrum feature map; step 43, based on the vibration frequency, performing thermal energy feature extraction and mapping based on vibration frequency locking on the full-spectrum feature map to obtain a harmonic energy transfer gain map.

[0033] In the above implementation, step 4 is specifically processed as follows: First, step 41 is executed. The processing object of this step is the sequence of thermal infrared frames of concrete pouring. This sequence originates from the raw thermal image data stream continuously acquired by a high-sensitivity thermal infrared camera at a specific frame rate, such as 600 frames / second, in step 1. Each frame image is a two-dimensional matrix with a resolution of 640×512 pixels, where the values ​​represent the radiation temperature value (unit: degrees Celsius or Kelvin) of the corresponding spatial location at that moment. In order to capture the dynamic changes during the vibration process, the algorithm does not process static images at a single moment, but extracts thermal infrared images containing the current time t and the preceding L consecutive frames as an analysis unit. The set time window length is determined based on the balance between the minimum effective frequency period multiple of the vibration signal and the real-time requirements of the system. For example, setting the time window length L=600 frames corresponds to 1 second of data at a sampling rate of 600Hz. This length can contain enough vibration periods to ensure the accuracy of spectrum analysis, while also taking into account the real-time requirements of computing resources. The data stacking operation is to sequentially stitch these consecutive L frames of images in the time dimension. In computer memory, construct a three-dimensional data cube with dimensions H×W×L, denoted as H×W×L. Here, H=512 represents the image height, W=640 represents the image width, and L=600 represents the temporal depth. Essentially, 600 thermal infrared images are stacked neatly like playing cards, forming a cuboid data block. Each spatial coordinate corresponds to a temperature change curve running along the time axis. However, this original data cube contains a significant amount of environmental background interference. For example, with changes in the angle of sunlight or fluctuations in overall air temperature, the base temperature of the concrete surface will slowly drift. The amplitude of this low-frequency change (DC component) is often much larger than the small high-frequency temperature rise caused by vibration friction. If not removed, this will cause energy leakage of the zero-frequency component in the frequency domain analysis, severely masking the effective signal. Therefore, trend removal processing is required. The algorithm traverses each spatial coordinate in the cube. A total of 512 × 640 = 327,680 pixels were collected, and their corresponding temporal vectors were extracted. For each vector, first calculate its arithmetic mean. This represents the average background temperature at that point within that one second. Next, subtracting this average value from each element of the vector yields the detrended fluctuation sequence. Its calculation logic follows the formula: ,in, This is a local time index, ranging from 0 to L-1. After this operation, the temperature sequence of all pixels is transformed into a pure fluctuation signal with 0 as the baseline, eliminating the DC component and highlighting the high-frequency dynamic change characteristics. These 327,680 processed fluctuation sequences are recombine in their original spatial positions to form the output temperature fluctuation cube. This is also a 512×640×600 dimension three-dimensional matrix, but it no longer stores absolute temperature values, but rather relative temperature fluctuation values.

[0034] Next, step 42 is executed. At this point, although the DC trend has been removed, the temperature fluctuation signal still contains multiple frequency components, including the vibration frequency signal of interest, as well as interference from other frequencies such as wind noise and electronic noise. To separate these components, the signal needs to be converted from the time domain to the frequency domain. Given the massive amount of data to be processed (over 320,000 time-series signals), this step relies on the powerful parallel computing capabilities of a graphics processing unit (GPU) to ensure real-time online monitoring. Parallel computing frameworks such as CUDA or OpenCL are used to process each spatial pixel in the temperature fluctuation cube. Allocate a dedicated computation thread. All threads simultaneously process their respective time-domain fluctuation signals. Perform a Fast Fourier Transform (FFT). The FFT operation decomposes a time-domain wave sequence into a superposition of sine and cosine waves, thus revealing the energy distribution of the signal at different frequencies. Specifically, for each pixel, the magnitude of its transform result is calculated to obtain the amplitude spectrum. The calculation process is rigorously described by the following formula: , in the formula, Representing pixels In frequency The spectral amplitude at a given point is an indicator of the intensity of temperature fluctuation at that frequency. This is the length of the sampling sequence; in this example, it is 600. For a moment Detrended temperature fluctuation values; The imaginary unit; The base of the natural logarithm is given. After the parallel operations described above, the original time-domain sequence of length L is transformed into a frequency-domain sequence of length L / 2, where, according to the Nyquist sampling theorem, the effective spectrum is only half the sampling rate. For the entire image, the original H×W×L time-domain cube is transformed into an H×W×(L / 2) frequency-domain data volume, i.e., a full-spectrum feature map. In this example, its dimensions are 512×640×300. This four-dimensional data structure (which is more complex if the real and imaginary parts are considered, but here we only refer to the amplitude spectrum) completely preserves the energy response characteristics of each pixel in the field of view at each resolvable frequency point. For example, at coordinates (100, 200), it indicates the amplitude of a temperature fluctuation at a frequency of 30Hz and the amplitude at a frequency of 120Hz.

[0035] Finally, proceed to step 43. First, a frequency locking operation is performed. This process involves locating key energy focal points on the frequency axis of the full-spectrum feature map. Since the full-spectrum feature map is a discrete frequency domain data volume, its frequency axis is divided into a series of discrete frequency points, spaced at frequency resolutions. 1Hz, determined by the sampling rate and the number of sampling points, is the vibration frequency calculated in step 3. 120Hz, the algorithm first calculates its corresponding index in the discrete spectrum. The calculation formula is: Meanwhile, considering the existence of nonlinear thermal response, it is also necessary to lock in its higher harmonic frequencies. Let the highest harmonic order be... If the value is 3, then the second harmonic needs to be calculated separately, i.e., 2× And the third harmonic, i.e., 3× Corresponding frequency index Where m takes values ​​of 2 and 3. For example, if the fundamental frequency is 120Hz, the frequency points to be locked are 120Hz, 240Hz, and 360Hz. It should be noted that for higher harmonics such as 360Hz, which are higher than the Nyquist frequency, the algorithm can be configured to use the aliasing features under undersampling for correction analysis, or rely on the dominant energy of the fundamental frequency and the second harmonic (240Hz) for analysis, because at a high sampling rate of 600Hz, they are all within the effective detection bandwidth. These specific frequency point indices are used to accurately extract energy slices directly related to vibration from the massive full-spectrum data. Then, weighted energy aggregation calculation is performed. For each pixel in the image... Extract the spectral amplitude at the locked frequency point from the full spectrum feature map. To comprehensively reflect the total energy contribution of the fundamental frequency and harmonics, a weighted summation method is used to calculate the harmonic energy transfer gain value at that point. The specific calculation formula is as follows: In this formula, Represents pixels The harmonic energy transfer gain value is a quantitative characterization of the vibration thermal response intensity at that point. Represents the fundamental frequency The energy at a given point is the square of the spectral amplitude. The summation symbol is used to accumulate the energy contribution of higher harmonics. For the harmonic order, traversing from 2 to . For the first Weighting coefficients for subharmonics. These weighting coefficients... These are preset empirical values ​​or parameters determined experimentally, and typically decrease with increasing harmonic order to reflect the physical fact that high-frequency energy decays more rapidly. For example, setting... =3, the weighting coefficient can be configured as follows: =0.5, =0.25. This means that the fundamental frequency energy dominates, the second harmonic energy weight is halved, and the third harmonic weight is further reduced. This weighted aggregation not only enhances the signal strength but also utilizes the complementarity of multi-band information, effectively suppressing random noise interference at a single frequency. Finally, image mapping is generated. Each calculated pixel... of The numerical values ​​were filled into a two-dimensional matrix of the same size as the original thermal infrared image, i.e., 512×640. This matrix was then normalized into a single-channel grayscale image or a floating-point data map, i.e., a harmonic energy transfer gain map. In this image, pixels with higher values, displayed as bright or warm tones, represent significant temperature rise at specific frequencies within the concrete at that location due to intense vibration friction, indicating that vibration energy was effectively transferred and converted; while areas with lower values, displayed as dark tones, indicate that energy transfer was hindered or severely attenuated.

[0036] In particular, during actual concrete vibration, concrete, as a complex viscoelastic medium, undergoes a dramatic evolution from a loose, under-vibrated state to a dense, compacted state. During this process, the damping characteristics and stiffness coefficient of the medium exhibit significant nonlinear changes. This abrupt change in physical properties inevitably triggers a strong nonlinear coupling effect, causing the vibration energy to no longer be limited to the actively input fundamental frequency band, but to be efficiently pumped or transferred to various harmonic frequency bands. That is, there is a dynamic energy transfer effect between the fundamental frequency energy and harmonic energy that is strongly correlated with the compaction degree. The aforementioned implementation method uses a fixed empirical weighting coefficient to statically linearly superimpose the fundamental frequency and harmonic energy. This approach not only lacks adaptability to different concrete mix proportions and field conditions, limiting the model's generalization ability, but also easily confuses the true phase-locked response with random background thermal noise. Especially in the early stages of vibration when the thermal signal is weak, the low signal-to-noise ratio easily leads to misjudgment of the compacted state. Therefore, simply observing the absolute energy value at a specific frequency point is insufficient to reveal the internal state. A dynamic analysis model that can perceive the intrinsic connections and energy flows between frequency points due to the evolution of physical states is needed. Therefore, this application adopts a harmonic energy transfer gain analysis method, which regards harmonic energy as the gain result excited by fundamental frequency energy in a nonlinear medium, and characterizes the density by calculating the ratio of harmonic energy excited by unit fundamental frequency energy.

[0037] Based on this, in a preferred embodiment of step 43, based on the vibration frequency, thermal energy feature extraction and mapping based on vibration frequency locking are performed on the full-spectrum feature map to obtain a harmonic energy transfer gain map, including: The background thermal noise baseline is estimated and subtracted from the full-spectrum feature map to obtain the noise baseline map. Since the original full-spectrum feature map inevitably contains a large amount of random thermal noise unrelated to the vibration machinery's movement, especially in the early stages of vibration or during deep vibration, the effective thermal signal is often weak. If this background noise is not removed, it will seriously interfere with the accuracy of subsequent energy ratio calculations. To separate the smoothly varying background noise from the mixed spectrum, a moving median filter algorithm is used to smooth the amplitude spectrum of each spatial pixel in the full-spectrum feature map. This algorithm utilizes the difference between the effective vibration signal (phase-locked response) and the morphological difference of the broadband, gently sloping base in the frequency domain. By filtering out sharp signal peaks through a sliding window, a clean noise baseline is fitted. The specific calculation formula is as follows: In the formula, Represents spatial coordinates Pixels at frequency Estimation of background noise amplitude at the location; It is the original spectral amplitude of the pixel in the full-spectrum feature map; Represents the median filtering function; It is the window width of the filter. The setting is determined based on the spectral resolution and the frequency domain broadening characteristics of the vibration signal. It is usually set to 3 to 5 times the effective signal peak width. For example, if the spectral resolution is 1Hz and the half-width at half-maximum (FWHM) of the vibration signal peak is approximately 3Hz, then the window width... It can be set to 15Hz (i.e., 15 sampling points). This step enables the active estimation and separation of random thermal noise, providing a high signal-to-noise ratio benchmark for subsequent calculations.

[0038] Based on the vibration frequency, the signal-to-noise enhancement energy of the fundamental frequency and harmonics is calculated from the full-spectrum feature map and noise baseline map to obtain the signal-to-noise enhancement fundamental frequency energy and the total signal-to-noise enhancement harmonic energy. This step aims to accurately quantify the pure thermal response energy directly caused by the vibration behavior, free from noise interference, from the mixed signal. Based on the vibration frequency obtained in step 3... The fundamental frequency and its higher harmonic frequencies are determined. For each pixel, signal purification is achieved by subtracting the corresponding noise energy in the noise baseline spectrum from the original energy of the full-spectrum feature map at these specific frequency points. To ensure the physical validity of the energy values, a non-negativity constraint is introduced in the calculation. The formula for calculating the fundamental frequency energy for signal-to-noise enhancement is as follows: In the formula, Coordinates Signal-to-noise enhancement fundamental frequency energy at the location; Is this point at the fundamental frequency of vibration? The original energy at the location (amplitude squared); This is the background noise energy at the same frequency. The formula for calculating the total harmonic energy of signal-to-noise enhancement is as follows: In the formula, It is the total harmonic energy of signal-to-noise enhancement; The highest harmonic order to be considered is usually 3 or 4, because higher harmonics have very weak energy and are easily affected by aliasing. The harmonic order; Let m be the frequency of the m-th harmonic. For example, if the vibration frequency... =120Hz. For a certain pixel, its original spectral amplitude at 120Hz is 50, and the noise baseline amplitude is 10; at 240Hz (second harmonic), its original amplitude is 20, and the noise amplitude is 5. Therefore, the signal-to-noise ratio (SNR) enhancement fundamental frequency energy at this point is 50^2 - 10^2 = 2400, and the SNR enhancement second harmonic energy is 20^2 - 5^2 = 375. This step no longer uses the contaminated original energy values, but instead uses the pure signal energy with a high SNR for subsequent physical modeling, significantly enhancing the robustness of the model.

[0039] The harmonic energy transfer gain (HETG) is calculated and mapped from the signal-to-noise enhancement (SNR) fundamental frequency energy and the total SNR harmonic energy to obtain the HETG graph. This is a key step in constructing a core index that can sensitively reflect changes in the coupling state within concrete. First, the ratio of total harmonic energy to fundamental frequency energy is calculated to obtain the HETG ratio. Then, a logarithmic transformation is performed on this ratio to obtain the final HETG index. The calculation formula is: In the formula, It is a very small positive real number, such as This ratio is used to prevent the denominator from being zero and to enhance numerical stability. It directly quantifies the nonlinear coupling strength. Harmonic energy transfer gain index. The calculation formula is: In the formula, This represents the logarithm with the natural logarithm as the base. Introducing the logarithmic transformation serves two purposes: firstly, it compresses the dynamic range, facilitating subsequent visualization; secondly, by utilizing the large slope of the logarithmic function in the small numerical range, it significantly improves the model's sensitivity in the crucial physical phase transition from undervibration to compaction, amplifying those small but vital state changes. For example, continuing with the above data, if... =2400, =375, if only the second harmonic is considered, then ≈0.156. (Calculated) =log(1.156)≈0.145. If the concrete becomes even denser, the nonlinear effect intensifies, and the harmonic energy increases to... =1200, then =0.5, HETG increased to 0.405. Finally, all pixels... The values ​​are mapped to a two-dimensional image matrix, resulting in a harmonic energy transfer gain map. In this map, dark areas represent under-vibrated, loose regions with low energy transfer efficiency, while bright areas intuitively indicate dense regions with significant nonlinear effects and efficient energy transfer, thus achieving precise visualization of the internal physical state of concrete. The harmonic energy transfer gain map generated by the above techniques has fundamental technical advantages. Unlike the fuzzy energy overlay map generated by traditional mechanisms, the pixel brightness of this map directly and quantitatively characterizes the nonlinear coupling strength and energy transfer efficiency within the concrete. Specifically, in the under-vibrated state of a loose medium, the energy transfer efficiency is low, the HETG value is low, and the image shows dark areas; as the medium becomes denser through vibration, the nonlinear effect intensifies, the HETG value increases, and the image turns into bright areas. This technique successfully transforms the difficult-to-observe "density" into an intuitive engineering indicator. By constructing a mathematical model that profoundly reflects the physical process, it overcomes the drawbacks of model simplification and noise interference, greatly improving the accuracy, sensitivity, and robustness of the system, thus providing solid technical support for achieving high-precision real-time quality control of the concrete pouring process.

[0040] In step 5, a density score map is obtained by fusing the harmonic energy transfer gain map, the thermal infrared pouring zone mask, and the thermal infrared vibration influence zone mask with density based on temperature field uniformity and energy indices. Correspondingly, during concrete vibration, density depends not only on the total input of vibration energy but, more importantly, on the uniformity and effective transfer efficiency of energy distribution in space. Excessive local energy may lead to segregation and bleeding, while insufficient energy will result in honeycomb-like pitting. Simply relying on the average energy of the entire map cannot reveal local quality defects and does not consider the interference from construction boundaries and non-working areas. Although the harmonic energy transfer gain map quantifies the thermal response intensity of each point to vibration excitation, these raw data are still pixel-level physical quantities and have not yet been converted into quality scores in an engineering sense. To achieve accurate quantification of vibration quality, the evaluation scope needs to be strictly limited to the physically effective working surface, and statistical indicators characterizing distribution characteristics need to be introduced. Therefore, combining spatial masking with dynamic meshing and local statistical analysis of harmonic energy maps aims to extract macroscopic energy intensity and distribution uniformity from microscopic pixel energy values, constructing a multidimensional evaluation system that reflects both vibration intensity and operational consistency, thereby achieving scientific scoring and grading of the internal compaction state of concrete.

[0041] Figure 4 This is a schematic diagram of the data flow in step 5 of the concrete pouring process monitoring method based on multi-source data according to an embodiment of this application. Figure 4As shown, in one embodiment of step 5, a density score map is obtained by performing a density fusion evaluation based on temperature field uniformity and energy index on the harmonic energy transfer gain map, the thermal infrared casting area mask, and the thermal infrared vibration influence area mask. This includes: step 51, performing dynamic grid division of the effective working surface of the harmonic energy transfer gain map based on the thermal infrared vibration influence area mask to obtain an effective evaluation grid set; step 52, performing statistical characteristic calculation of regional harmonic energy transfer gain for each sub-region in the effective evaluation grid set to obtain a statistical feature set; and step 53, performing a density score based on nonlinear normalization on the statistical feature set based on the uniformity penalty factor to obtain the density score map.

[0042] In the above implementation, step 5 is specifically processed as follows: First, step 51 is executed. ROI extraction is performed first. Using the thermal infrared vibration influence area mask as a logic filter, a bitwise AND operation is performed on the harmonic energy transfer gain map. Specifically, only the pixel value of 1 in the thermal infrared vibration influence area mask, i.e., the value in the harmonic energy transfer gain map corresponding to the white area, is retained, and the values ​​in other positions are set to zero or marked as invalid. This effectively extracts the effective gain map containing only the current vibration area. For example, if the effective radius of the vibrator in the thermal infrared image is approximately 66 pixels, the extracted image only has non-zero energy values ​​within this circular area, excluding the thermal radiation interference from the steel bars or formwork in the background. Next, meshing is performed. To refine the evaluation of local density, the rectangular bounding box area containing the extracted effective gain map is divided into K non-overlapping rectangular evaluation sub-region sets. Each sub-region The size can be set to a fixed dimension, such as a 16×16 pixel square, which corresponds to a micro-element area of ​​approximately 3.2cm×3.2cm in actual physical space. If the bounding box size of the vibration-affected zone is 132×132 pixels, it can be divided into approximately 64 such sub-grids. Validity filtering is then performed. Since the vibration-affected zone is circular, while the grid division is rectangular, some grids will inevitably lie at the edge of the circle, with only a small portion of pixels within that circle belonging to the effective vibration range. To avoid statistical bias due to insufficient samples, each sub-region needs to be traversed. Calculate the percentage of effective pixels within the mask area that falls under the influence of thermal infrared vibration. Set a preset threshold. This is determined by analyzing the relationship between the proportion of different pixels and the stability of statistical characteristics in historical monitoring data; for example, 0.6 represents 60%. If a certain grid... The effective pixel ratio within is less than If the proportion is less than or equal to the threshold, the grid is marked as invalid and discarded; otherwise, if the proportion is greater than or equal to the threshold, the grid is retained and added to the set of valid evaluation grids. This set ultimately contains a list of sub-regions located only in the core vibratory tamping area, ensuring the sufficiency and physical representativeness of the sample for subsequent statistical analysis.

[0043] Next, step 52 is executed. This step aims to condense the complex pixel-level energy data within each grid into statistical indicators with clear physical meaning. Specifically, this statistical feature set includes: the arithmetic mean of the harmonic energy transfer gain within the region and the standard deviation of the harmonic energy transfer gain within the region. The algorithm effectively evaluates each sub-region in the grid set. Perform a grid-by-grid traversal. For example, if the k-th sub-region contains... The effective pixels are those points that are simultaneously located within the grid and within the mask area. First, the first moment is calculated, which is the arithmetic mean of the harmonic energy transfer gain within the region. This index characterizes the average intensity of the nonlinear thermal response of the concrete medium to the vibration wave within this local micro-element region, directly reflecting the sufficiency of the vibration energy input. The calculation formula is: In the formula, is the average harmonic energy transfer gain of subregion k; This represents the total number of valid pixels within the region. For pixels The harmonic energy transfer gain value at a given location is calculated. For example, if the gain values ​​of all pixels within a certain grid are generally high, and the calculated mean reaches 0.8 (normalized units), it indicates that strong vibration energy was received at that location. Next, the second-order central moment is calculated, which is the standard deviation of the harmonic energy transfer gain within the region. This index characterizes the dispersion of energy distribution within the region, i.e., uniformity. Under ideal compaction conditions, the physical properties of all points inside the concrete tend to be consistent, and the response to vibration waves should also be uniform. If the standard deviation is too large, it means that there are energy hotspots and colds within the region, suggesting possible uneven aggregate distribution or incomplete air bubble removal. The calculation formula is as follows: In the formula, Let be the standard deviation of the harmonic energy transfer gain for sub-region k. A smaller value indicates a more consistent mechanical response and a more uniform compaction process within the region. For example, a calculated standard deviation of only 0.05 indicates a very uniform energy distribution within the grid. After the above calculations, a pair of statistical characteristic values ​​is generated for each valid evaluation grid k. These sets of feature pairs constitute the statistical feature set of the final output. This set quantifies not only whether the vibration is sufficient (mean) but also whether the vibration is uniform (standard deviation).

[0044] Final step 53 is implemented. This step is the final synthesis of the aforementioned physical feature extraction and statistical analysis. First, feature fusion calculation is performed. This step involves designing a scoring function that can scientifically reflect the physical essence of "density." According to the mechanism of concrete vibration compaction, a high average harmonic gain represents sufficient nonlinear interaction, meaning that the vibration energy is effectively transferred to the interior of the concrete, which is a positive indicator; while a high gain standard deviation represents uneven energy distribution, possibly indicating voids or aggregate accumulation, which is a negative indicator. To unify the two into a single score, the following nonlinear scoring model is constructed: In this formula, This represents the new overall density score for the k-th sub-region; a higher value indicates better pouring quality in that region. (Numerator) The average harmonic energy transfer gain in this region directly contributes to the increase in the fraction. (Denominator term) It served as a punishment. This is the standard deviation of harmonic energy transfer gain, used to measure the degree of non-uniformity. Known as the uniformity penalty factor, it is an adjustable hyperparameter used to control the severity of the penalty for non-uniformity. A constant of 1 is used for regularization, preventing the denominator from being zero and ensuring that, when the standard deviation is extremely small (ideally uniform), the fraction directly approaches the average gain value. Uniformity Penalty Factor This is typically determined by analyzing the statistical distribution differences between historical high-quality casting samples and defective samples (such as honeycomb-like pitting), for example, by setting... =5.0 means that for every 0.1 unit increase in non-uniformity, the denominator will increase significantly, leading to a sharp drop in the final score. This effectively eliminates inferior regions that appear to have sufficient energy but are extremely unevenly distributed. Specifically, assume there are two distinct sub-regions A and B. The average gain of region A... =0.8, standard deviation =0.05, indicating high and uniform energy; average gain of region B. =0.8, but standard deviation =0.3, high energy but extremely uneven, possibly with local voids. For region A, its score is: For region B, the score is: As can be seen, although the average energy received by the two regions is the same, region B is severely penalized due to its higher non-uniformity, receiving only half the score of region A. This scoring mechanism effectively identifies and distinguishes between two distinct quality states, meeting the requirements of uniformity and density in practical engineering. Finally, mapping generation is performed. A single-channel floating-point matrix matching the resolution of the input image grid is initialized, or an image matrix with a resolution of 512×640 is directly constructed. Each sub-region in the effective evaluation grid set is traversed. The calculated comprehensive score The backfill is applied to the coordinate range corresponding to the sub-region in the matrix. Background areas not included in the effective grid are filled with 0 or a specific invalid marker value. The resulting matrix is ​​the density score map. In this map, the grayscale value or numerical value of each pixel directly represents the overall pouring quality score of that micro-region. High-scoring areas, such as those with values ​​close to 0.8, indicate that the concrete is both well-vibrated and evenly distributed, belonging to high-quality pouring areas; while low-scoring areas, such as those with values ​​below 0.3, visually indicate potential honeycomb, pitting, or under-vibration risk areas. This map provides a direct data source for subsequently generating intuitive heatmaps.

[0045] In step 6, a visual heatmap of the pouring quality is generated based on the density score map and the thermal infrared pouring area mask. That is, although the density score map accurately quantifies the pouring quality of the concrete through a numerical matrix, it is often difficult for operators, supervising engineers, and managers on the construction site to intuitively judge the quality status immediately when faced with purely numerical or grayscale images. In a fast-paced, high-intensity construction environment, users need a visual feedback mechanism that can instantly and clearly convey where the work is up to standard and where there are problems. Transforming abstract quality scores into color codes that conform to human visual perception can greatly lower the threshold for information interpretation, helping workers quickly locate potential honeycomb, pitted, or under-vibration areas. Therefore, combining semantic masks to filter density scores and applying threshold-based segmented color mapping to generate a visual heatmap is to transform complex internal quality data into intuitive red, green, and yellow warning information. Through augmented reality overlay technology, this invisible information is directly projected onto the actual construction scene, thereby achieving real-time monitoring and efficient closed-loop management of pouring quality.

[0046] In one embodiment of step 6, the specific processing is as follows: First, a mask filtering process is performed. Since the density score map is calculated based on the entire map or a specific vibration grid, background areas (such as formwork, rebar mesh, or ground) may contain meaningless zero values ​​or noise values. To ensure that the final visualization result focuses only on the actual concrete entity and avoids false alarms for non-pouring areas, a thermal infrared pouring area mask is used as a mask to logically filter the density score map. Specifically, each pixel or sub-grid coordinate in the density score map is traversed. If the pixel value corresponding to this coordinate in the thermal infrared pouring area mask is 0 (background), the density score at that location is forcibly set to an invalid value or a marker value with completely zero transparency. This step cleans the data, ensuring that subsequent color rendering is only performed within the valid concrete area.

[0047] Next, color space mapping is performed. This is the crucial step in converting numerical values ​​into visual colors. The algorithm iterates through each sub-region k in the filtered density score map, reading its density score. And apply a segmented color mapping function according to the preset quality grading standard. This converts the scalar fraction into a standard RGB color vector. The mapping logic is strictly executed according to the following piecewise function: In this formula, The pixel color value corresponding to sub-region k on the heatmap is represented by a three-channel red, green, and blue (Red, Green, Blue) representation. This represents the density score of the region. The density threshold, for example, is set to 0.6. This value is determined based on the national building construction quality acceptance standard and on-site concrete mix design test data, and is the minimum energy uniformity score line that can guarantee the strength and impermeability of concrete. If the score is higher than this line, it will be displayed in green, indicating that the density is qualified. The defect alarm threshold is determined based on the lower bound (e.g., quantile value) of the statistical values ​​of low energy gain and high non-uniformity characteristics corresponding to quality defects such as honeycomb and pitting in historical projects. For example, it is set to 0.3. If the score is between these two values, it is displayed in yellow, representing critical / under-vibration, indicating that additional vibration is needed. If the score is below this line, it is displayed in red, representing serious defects / honeycomb risk, and re-vibration processing must be carried out immediately. For example, for region A calculated in step 5, its score is 0.64, which is greater than 0.6, so this region is rendered as bright green in the heat map; while region B has a score of 0.32, which is between 0.3 and 0.6, and is rendered as a warning yellow.

[0048] Finally, layer overlay is performed. This generated pseudo-color layer with clearly defined color coding is semi-transparent (e.g., setting the alpha channel opacity to 0.5), and then precisely overlaid on the original visible light video frame (after spatial transformation and alignment) or the projected view of the BIM model. The final output pouring quality visualization heatmap is not just a static image, but an augmented reality (AR) view. On the monitoring screen, managers can see the actual construction process, while a dynamically changing red, green, and yellow semi-transparent mask covers the concrete surface, visually displaying the internal density of every inch of concrete. This makes previously invisible quality hazards "visible," achieving true digital construction management.

[0049] In summary, the concrete pouring process monitoring method based on multi-source data, as described in this application, addresses the challenge of existing single-modal visual solutions being unable to perceive the internal density of concrete. This application utilizes visible light video to perform semantic segmentation and localization of the pouring area and vibrator, constructing a precise spatial mask; simultaneously, it extracts key vibration frequencies from the vibration signals. Based on this, it applies phase-locked thermal imaging technology, using the vibration frequency as a phase-locked reference, to extract harmonic energy transfer gain features at the same frequency as the vibration from the thermal infrared image sequence. Since the internal density of concrete directly affects the transfer efficiency of vibration energy (manifested as thermal energy fluctuations), a physical mapping model of vibration energy input - thermal field fluctuation response - internal density state is successfully established. By fusing spatial location and energy transfer characteristics to assess temperature field uniformity, quantitative visualization of concrete pouring quality is achieved, effectively solving the technical problem that surface visual features alone cannot determine whether internal quality defects such as honeycomb or pitting exist.

[0050] Figure 5 This is a block diagram of a concrete pouring process monitoring system based on multi-source data according to an embodiment of this application. Figure 5As shown, the concrete pouring process monitoring system 100 based on multi-source data according to an embodiment of this application includes: a data acquisition module 110, used to acquire visible light video frames, thermal infrared frames, and vibration signals of concrete pouring; a pouring area and vibrator positioning module 120, used to position the pouring area and vibrator of the visible light video frames of concrete pouring to obtain a thermal infrared pouring area mask and a thermal infrared vibration influence area mask; a vibration signal analysis module 130, used to perform a fast Fourier transform on the vibration signal to obtain the vibration frequency; and a vibration... The phase-locked thermal imaging analysis module 140 is used to perform phase-locked thermal imaging analysis on the sequence of thermal infrared frames of concrete pouring based on the vibration frequency to obtain a harmonic energy transfer gain map; the compaction fusion evaluation module 150 is used to perform compaction fusion evaluation based on temperature field uniformity and energy index on the harmonic energy transfer gain map, thermal infrared pouring area mask and thermal infrared vibration influence area mask to obtain a compaction score map; the pouring quality visualization module 160 is used to generate a pouring quality visualization heat map based on the compaction score map and thermal infrared pouring area mask.

[0051] Here, those skilled in the art will understand that the specific operations of each step in the above-mentioned multi-source data-based concrete pouring process monitoring system have been referenced above. Figures 1 to 4 The method for monitoring the concrete pouring process based on multi-source data has been described in detail, and therefore, its repeated description will be omitted.

Claims

1. A method for monitoring the concrete pouring process based on multi-source data, characterized in that, include: Acquire visible light video frames, thermal infrared frames, and vibration signals of concrete pouring; The visible light video frames of concrete pouring are used to locate the pouring area and the vibrator to obtain the thermal infrared pouring area mask and the thermal infrared vibration influence area mask. The vibration signal is subjected to a fast Fourier transform to obtain the vibration frequency; Based on the vibration frequency, a phase-locked thermal imaging analysis of the sequence of thermal infrared frames of concrete pouring was performed to obtain a harmonic energy transfer gain map. A density score map was obtained by performing a density fusion evaluation based on temperature field uniformity and energy index on the harmonic energy transfer gain map, the thermal infrared casting area mask, and the thermal infrared vibration influence area mask. Based on the density fraction map and the thermal infrared casting zone mask, a visual heat map of casting quality is generated.

2. The method for monitoring the concrete pouring process based on multi-source data according to claim 1, characterized in that, The visible light video frames of concrete pouring are used to locate the pouring area and vibrator to obtain a mask for the thermal infrared pouring area and a mask for the thermal infrared vibration-affected area, including: Semantic segmentation and extraction of the concrete pouring area are performed on visible light video frames of concrete pouring to obtain a mask of the pouring area. Multi-scale target detection and localization of vibrating rods were performed on visible light video frames of concrete pouring to obtain a mask of the vibration-affected zone. Based on the spatial alignment transformation matrix, cross-modal coordinate transformation is performed on the vibration-affected zone mask and the pouring zone mask to obtain the thermal infrared pouring zone mask and the thermal infrared vibration-affected zone mask.

3. The method for monitoring the concrete pouring process based on multi-source data according to claim 2, characterized in that, Multi-scale target detection and localization of the vibrator in visible light video frames of concrete pouring to obtain a mask of the vibration-affected zone, including: The visible light video frames of concrete pouring are input into the target detection model to obtain the bounding box coordinates and confidence level of the vibratory tamping rod head; Using the center point of the vibrator in the bounding box coordinates of the vibrator head as the center, a circular region descriptor is generated according to the preset effective vibration radius; On a black background of the same size as the visible light video frame of concrete pouring, the circular area corresponding to the circular area descriptor is filled with white to obtain a binary mask of the influence range of the vibrator, which serves as the vibration influence area mask.

4. The method for monitoring the concrete pouring process based on multi-source data according to claim 1, characterized in that, Performing a fast Fourier transform on the vibration signal to obtain the vibration frequency includes: The vibration signal is subjected to a fast Fourier transform to obtain the vibration frequency domain signal; Peak search is performed on the vibration frequency domain signal to obtain the frequency point with the largest amplitude as the vibration frequency.

5. The method for monitoring the concrete pouring process based on multi-source data according to claim 1, characterized in that, Based on the vibration frequency, a phase-locked thermal imaging analysis of the sequence of thermal infrared frames from concrete pouring was performed to obtain a harmonic energy transfer gain map, including: Data stacking and trend removal were performed on the sequence of thermal infrared frames of concrete pouring to obtain a temperature fluctuation cube. A pixel-by-pixel time-series signal frequency domain transformation is performed on the temperature fluctuation cube to obtain a full-spectrum feature map; Based on the vibration frequency, thermal energy features are extracted and mapped from the full-spectrum feature map using vibration frequency locking to obtain a harmonic energy transfer gain map.

6. The method for monitoring the concrete pouring process based on multi-source data according to claim 5, characterized in that, Based on the vibration frequency, thermal energy features are extracted and mapped from the full-spectrum feature map using vibration frequency locking to obtain a harmonic energy transfer gain map, including: The background thermal noise baseline is estimated and subtracted from the full-spectrum feature map to obtain the noise baseline map; Based on the vibration frequency, the signal-to-noise enhancement energy of the fundamental frequency and harmonics is calculated from the full spectrum characteristic map and the noise baseline map to obtain the signal-to-noise enhancement fundamental frequency energy and the signal-to-noise enhancement total harmonic energy. The harmonic energy transfer gain is calculated and mapped to the signal-to-noise enhancement fundamental frequency energy and the signal-to-noise enhancement total harmonic energy to obtain the harmonic energy transfer gain diagram.

7. The method for monitoring the concrete pouring process based on multi-source data according to claim 1, characterized in that, A density score map was obtained by fusing the harmonic energy transfer gain map, the thermal infrared casting zone mask, and the thermal infrared vibration influence zone mask with temperature field uniformity and energy index. The results included: Based on the thermal infrared vibration influence zone mask, the harmonic energy transfer gain map is dynamically divided into vibration effective working surfaces to obtain an effective evaluation grid set; Statistical characteristic calculations are performed on the regional harmonic energy transfer gain of each sub-region in the effective evaluation grid set to obtain a statistical characteristic set; Based on the uniformity penalty factor, a density score based on nonlinear normalization is performed on the statistical feature set to obtain the density score map.

8. The method for monitoring the concrete pouring process based on multi-source data according to claim 7, characterized in that, The statistical feature set includes: the arithmetic mean of harmonic energy transfer gain within the region and the standard deviation of harmonic energy transfer gain within the region.

9. A concrete pouring process monitoring system based on multi-source data, characterized in that, include: The data acquisition module is used to acquire visible light video frames, thermal infrared frames, and vibration signals of concrete pouring. The pouring area and vibrator positioning module is used to position the pouring area and vibrator in the visible light video frame of concrete pouring to obtain the thermal infrared pouring area mask and the thermal infrared vibration influence area mask. The vibration signal analysis module is used to perform a fast Fourier transform on the vibration signal to obtain the vibration frequency. The vibration phase-locked thermal imaging analysis module is used to perform vibration phase-locked thermal imaging analysis on the sequence of thermal infrared frames of concrete pouring based on the vibration frequency to obtain a harmonic energy transfer gain map. The density fusion evaluation module is used to perform a density fusion evaluation based on temperature field uniformity and energy index on the harmonic energy transfer gain map, the thermal infrared casting area mask and the thermal infrared vibration influence area mask to obtain a density score map. The casting quality visualization module is used to generate a casting quality visualization heat map based on the density fraction map and the thermal infrared casting area mask.

Citation Information

Patent Citations

  • Real-time monitoring system and method for concrete vibration

    CN110608769A

  • Concrete vibration construction quality monitoring method and system based on machine vision

    CN116912198A

  • Method for detecting concrete vibration compactness

    CN117890434A

  • A method, system, product and medium for intelligent monitoring of concrete pouring quality

    CN119758941A

  • Fresh concrete vibration apparent quality identification method based on convolutional neural network

    CN120579157A