A pipe piece liquid surface pouring detection method and device based on multi-modal perception

By fusing visual images and weight sensing data using multimodal perception technology, an adaptive attention mechanism is constructed, which solves the accuracy and stability problems of existing segment liquid level detection in complex environments, and realizes high-precision liquid level detection and automated control.

CN122365307APending Publication Date: 2026-07-10HUNAN XIJIA ROBOT CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
HUNAN XIJIA ROBOT CO LTD
Filing Date
2026-06-05
Publication Date
2026-07-10

AI Technical Summary

Technical Problem

Existing segment liquid level detection technologies suffer from limitations in complex industrial environments, including limited detection information, poor environmental adaptability, and insufficient ability to identify abnormal operating conditions. This results in insufficient detection accuracy and stability, making it difficult to meet the requirements of automated production.

Method used

A detection method based on multimodal perception is adopted. By fusing dual-view visual images and weight sensing data, multimodal features are constructed and adaptively controlled to achieve comprehensive analysis of liquid level height, distribution state and weight changes. Data is collected using a visual camera and a weight sensor, multimodal feature extraction and fusion are performed, and a spatiotemporal-modal collaborative adaptive attention mechanism is constructed for intelligent decision-making.

Benefits of technology

It significantly improves the accuracy and stability of liquid level detection, can identify the uniformity of liquid distribution and abnormal pouring conditions, reduces the false judgment rate, adapts to high viscosity and high pollution environments, reduces equipment maintenance frequency, and improves the level of automation and intelligence.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122365307A_ABST
    Figure CN122365307A_ABST
Patent Text Reader

Abstract

This invention relates to a method and apparatus for detecting the concrete slab surface during concrete slab pouring based on multimodal perception. The method includes: acquiring concrete slab surface image data from different perspectives and acquiring weight data of the slab forming mold; establishing a one-to-one correspondence between the concrete slab surface image data and the weight data, and constructing a multimodal input sample containing visual and weight features; extracting the visual and weight features from the multimodal input sample and fusing them to obtain multimodal initial features; obtaining global description vectors for different modalities based on the multimodal initial features; constructing a weight temporal dynamic embedding vector and constructing attention weights that adaptively adjust the contribution of visual and weight features; obtaining multimodal weighted features based on the attention weights and outputting the detection result indicating whether the concrete slab surface pouring is in place. This solution solves the problems of relying on manual experience for slab surface judgment and the high misjudgment rate of a single detection method during concrete slab pouring.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of industrial process detection and intelligent sensing technology, and in particular to a method and device for detecting the liquid level of pipe segments during casting based on multimodal sensing. Background Technology

[0002] In many stages of industrial production, the precise filling of liquid or semi-fluid materials into molds, containers, or carriers is a critical operation. This operation aims to meet the stringent requirements of subsequent molding, curing, or processing techniques regarding the consistency of material volume and quality. Whether the liquid level reaches the predetermined height, i.e., the liquid level condition, becomes a key control indicator affecting product quality and production stability. If the liquid level does not reach the predetermined height, it may lead to insufficient product dimensions and performance defects; while if the liquid level is too high, it may cause material waste and even affect the normal operation of production equipment.

[0003] However, in actual industrial settings, the liquid filling process is accompanied by a series of complex conditions, posing significant challenges to liquid level detection. For example, changes in liquid flow rate can exacerbate surface fluctuations, making it difficult to stably determine the liquid level position; foam formation not only alters the appearance of the liquid surface but can also interfere with detection signals; liquid splashing can cause instantaneous changes in the liquid level, increasing detection difficulty; liquid adhering to the container walls makes it difficult to accurately determine the true height of the liquid level; surface reflections can cause visual interference to the detection equipment; and fluctuations in ambient light further affect the accuracy of detection. These factors cause the liquid surface morphology to exhibit significant instability within a short period of time.

[0004] Currently, the main methods for detecting when the liquid level is in place are as follows: Weight detection method: This method determines the filling status based on a weight threshold. Its principle is to monitor the total weight change of the container and its contents; when the weight reaches a preset threshold, filling is considered complete. However, this method has a significant drawback: it cannot reflect abnormal liquid distribution within the container. For example, in cases of excessive foam or fluid adhering to the container walls, even if the actual liquid level has not reached the predetermined height, the added weight from the foam or adhering liquid may still lead the system to mistakenly classify it as fully filled, thus affecting product quality.

[0005] Monocular vision detection: This method determines the liquid surface position by analyzing the acquired images. Image processing technology is used to identify the liquid surface boundary, thereby determining the liquid level. However, in practical applications, this method is highly susceptible to factors such as lighting, reflections, and occlusion. Complex lighting conditions in industrial environments can reduce image contrast and blur features, making it difficult to accurately represent the true distribution of the liquid, thus affecting the accuracy of liquid level determination.

[0006] Contact-based liquid level sensing: This type of detection requires the sensor to be in direct contact with the liquid, determining the liquid level by sensing the interaction between the liquid and the sensor. However, in high-viscosity, easily agglomerated conditions, the sensor is prone to adhesion, blockage, or wear, which not only affects detection accuracy but also leads to a significant increase in maintenance costs. Furthermore, this contact-based detection method is unsuitable for continuous automated production, as frequent contact during production may disrupt the process, and sensor maintenance and replacement can impact production efficiency.

[0007] The aforementioned problems are particularly prominent in concrete segment casting scenarios. While existing segment liquid level detection systems have partially achieved automated control of the casting process using methods such as weight sensing, visual inspection, or liquid level sensing, they still exhibit several significant technical shortcomings under complex industrial casting conditions due to their limited detection methods and information dimensions. Relying on single-method detection is insufficient to accurately reflect the actual pouring state of concrete and lacks robustness: Existing systems often rely solely on one method—weight detection, monocular visual detection, or contact-based liquid level detection—as the basis for judgment. This single-method detection only reflects the change of a single physical quantity during the concrete pouring process and cannot comprehensively present the spatial distribution of concrete within the mold. For example, when concrete adheres to the mold walls, although the mold weight may reach the set value, the actual concrete may not have filled the mold evenly. In this case, the judgment based on weight detection will deviate from the actual pouring state. Similarly, monocular visual detection is also prone to misjudgment when faced with local accumulation, liquid level fluctuations, or foam interference, significantly reducing the accuracy of liquid level determination and resulting in poor system robustness.

[0008] Visual inspection suffers from limited dimensions and susceptibility to environmental interference, resulting in low stability in liquid surface recognition: Existing monocular vision-based detection schemes can only acquire two-dimensional image information, lacking effective perception capabilities for the spatial morphology and depth distribution of the liquid surface. In actual concrete pouring sites, factors such as dust, reflections, shadows, lighting variations, and occlusion by mold structures are prevalent. These factors can blur the liquid surface boundaries, even leading to false detections, severely impacting the stability and consistency of liquid surface height judgment and making it difficult to meet the needs of continuous detection under complex conditions. For example, under strong light, reflections from the liquid surface may make it difficult to accurately identify the liquid surface boundaries in the image; while occlusion by the mold structure may prevent the acquisition of some liquid surface information, thus affecting the detection results.

[0009] Weight-based detection methods cannot distinguish between pouring anomalies and are prone to misjudgment: These methods rely solely on changes in the overall weight of the mold to determine completion, failing to differentiate between abnormalities such as uneven concrete distribution, localized accumulation, wall residue, or foreign matter contamination. When concrete is not uniformly filled but the weight precisely reaches the set threshold, the system may misjudge it as complete, posing a potential risk to product quality. For example, if concrete accumulates in one corner of the mold, although the overall weight meets the standard, other areas may not be adequately filled, affecting the quality of the tunnel segments.

[0010] Contact-based liquid level detection methods suffer from poor adaptability and high maintenance costs: Some existing technologies use float-type or capacitive liquid level sensors for detection, which require direct contact with concrete. However, concrete has high viscosity and strong abrasiveness, making sensors prone to adhesion, clogging, or wear in such a pouring environment. Once these problems occur, detection accuracy decreases, failure rate increases, and maintenance costs rise significantly. Furthermore, because segment production lines require high-frequency, continuous operation, the frequent failures and maintenance requirements of contact-based liquid level sensors cannot meet the demands of automated production.

[0011] The detection logic is simple and lacks the ability to comprehensively judge multi-dimensional state information: Most existing liquid level detection systems use threshold comparison or simple rule judgment methods, lacking the ability to comprehensively analyze changes in liquid surface shape, weight change trends, and abnormal working conditions. In practical applications, facing complex pouring conditions, this simple detection logic is prone to problems such as false stop, missed stop, or false alarms. For example, when the liquid level fluctuates briefly but is not truly in place, the system may misjudge it as in place and stop pouring; or when abnormal working conditions occur, the system cannot identify and issue alarms in a timely and accurate manner, thus affecting production quality and efficiency.

[0012] It is evident that existing technologies for detecting the proper placement of liquid level in tunnel segments generally suffer from problems such as limited detection information, poor environmental adaptability, and insufficient ability to identify abnormal working conditions, making it difficult to simultaneously meet the requirements of detection accuracy, stability, and reliability in complex industrial environments. Summary of the Invention

[0013] The technical problem to be solved by the present invention is to provide a method and device for detecting the liquid level of pipe segments based on multimodal sensing.

[0014] To achieve the above-mentioned objectives, this invention provides a method for detecting the liquid level during segment casting based on multimodal sensing, comprising the following steps: S1. Collect concrete liquid surface image data inside the segment forming mold from different perspectives, and continuously collect and normalize the weight data of the segment forming mold. S2. Based on the time synchronization mechanism, the concrete liquid surface image data and weight data are mapped one-to-one, and a multimodal input sample containing visual features and weight features is constructed. S3. Extract visual and weight features from the multimodal input samples and fuse them to form the initial multimodal features; S4. Group the different modalities in the initial multimodal features and perform global information compression on each group to obtain global description vectors for different modalities; where the global description vectors for different modalities are the global description vectors for visual features and the global description vectors for weight features, respectively. S5. Collect continuous weight data to construct a weight temporal dynamic embedding vector, and fuse the weight temporal dynamic embedding vector with global description vectors of different modalities to construct attention weights that adaptively adjust the contribution of visual features and weight features. S6. Based on the attention weights, the initial multimodal features are recalibrated to obtain multimodal weighted features, and the detection result of whether the liquid level of the pipe segment is in place is output based on the multimodal weighted features.

[0015] According to one aspect of the present invention, in step S1, the step of acquiring image data of the concrete liquid surface inside the segment forming mold based on different perspectives involves acquiring images of the concrete liquid surface inside the segment forming mold using a dual-perspective method. The first perspective is a top-down view directly in front of the segment forming mold, used to obtain the height profile of the concrete liquid surface. The concrete liquid surface image data obtained from the first perspective is represented as follows: ; in, Represents the real number field. , These represent the height and width of the image, respectively. Indicates the number of channels; The second perspective is an angle tilted relative to the segment forming mold, used to obtain the uniformity of concrete surface distribution and local accumulation; the concrete surface image data obtained from the second perspective is represented as follows: .

[0016] According to one aspect of the present invention, in step S1, the step of continuously collecting and normalizing the weight data of the segment forming mold, the normalized weight data is expressed as follows: ; in, This indicates the maximum weight value during the pouring process. This indicates the minimum weight value during the pouring process. This indicates the weight data collected in real time, and .

[0017] According to one aspect of the present invention, step S3, which involves extracting visual features and weight features from the multimodal input samples and fusing them to obtain initial multimodal features, includes: S31. Based on multimodal input samples, image features from different perspectives are extracted and stitched together in the channel dimension to obtain fused multi-view visual features; S32. Extract weight features based on multimodal input samples; S33. Broadcast the weight features in the spatial dimension to make them consistent with the visual feature dimension, and fuse them with the visual features to obtain multimodal initial features.

[0018] According to one aspect of the present invention, in step S33, the obtained multimodal initial features are represented as follows: ; ; ; ; ; in, This indicates the fusion of visual features from multiple perspectives. Indicates the fusion mask features, This represents the weight characteristic after broadcasting expansion in the spatial dimension. Visual features representing first-person perspective images of the concrete liquid surface. Visual features representing second-person perspective images of the concrete liquid surface. This represents the weight features extracted based on multimodal input samples. ReLU represents the nonlinear qualitative activation function. , They represent the weights, , These represent the bias terms.

[0019] According to one aspect of the present invention, in step S4, in the step of grouping different modalities in the multimodal initial features and performing global information compression respectively, global average pooling is used to perform global information compression on the different modalities in the multimodal initial features.

[0020] According to one aspect of the present invention, in step S4, the global description vector of visual features is represented as: ; in, Represents a global description vector of visual features; The global description vector of weight features is represented as follows: ; in, This represents the global description vector of weight features. This indicates the number of weight sensors.

[0021] According to one aspect of the present invention, in step S5, in the step of collecting continuous-time weight data to construct a weight temporal dynamic embedding vector, and fusing the weight temporal dynamic embedding vector with global description vectors of different modalities to construct attention weights that adaptively adjust the contribution of visual features and weight features, the attention weights that adaptively adjust the contribution of visual features and weight features are expressed as follows: ; ; ; in, This represents the attention weights, which adaptively adjust the contribution of visual and weight features. The attention weights representing visual features, and , The attention weights represent the weight features, and , This represents the activation function Sigmoid. , This represents the learnable weight matrix for the visual feature fusion processing branch. , This represents the learnable weight matrix of the weight feature fusion processing branch.

[0022] According to one aspect of the present invention, in step S6, the step of recalibrating the initial multimodal features based on the attention weights to obtain multimodal weighted features is represented as follows: ; in, This represents multimodal weighted features.

[0023] To achieve the above-mentioned objective, the present invention provides an apparatus for detecting the liquid level during segment casting based on the aforementioned multimodal sensing method, comprising: A vision camera is used to acquire image data of the concrete liquid surface inside the segment forming mold from different perspectives; Weight sensor, used to continuously collect weight data of the segment forming mold; The data preprocessing module is used to normalize the weight data; The time synchronization module is used to establish a one-to-one correspondence between concrete liquid surface image data and weight data based on the time synchronization mechanism, and to construct multimodal input samples containing visual features and weight features. A multimodal feature extraction and fusion module is used to extract visual and weight features from multimodal input samples and fuse them to obtain multimodal initial features; and to group different modalities in the multimodal initial features and perform global information compression on each to obtain global description vectors for different modalities; wherein the global description vectors for different modalities are global description vectors for visual features and global description vectors for weight features, respectively; and to collect weight data over continuous time to construct a weight temporal dynamic embedding vector, and to fuse the weight temporal dynamic embedding vector with the global description vectors for different modalities to construct attention weights that adaptively adjust the contribution of visual and weight features; and to recalibrate the multimodal initial features based on the attention weights to obtain multimodal weighted features; The liquid level determination module is used to output the detection result of whether the liquid level of the pipe segment is in place based on the multimodal weighted features.

[0024] According to one aspect of the present invention, this approach simultaneously acquires multi-view image information and real-time weight data, and through multi-modal feature joint modeling, achieves comprehensive analysis of liquid level height, liquid distribution state, and weight changes. This approach accurately determines whether the pouring is in place, enabling high-precision judgment of liquid level height, distribution state, and pouring completion status, and significantly improving the intelligence and reliability of segment pouring quality inspection.

[0025] According to one aspect of the present invention, this approach overcomes the limitations of single visual or weight detection, significantly improves the accuracy and stability of liquid level detection in complex industrial scenarios, and provides an effective technical means for the automation and intelligent quality control of the liquid filling process.

[0026] According to one aspect of the present invention, this solution addresses the problems commonly found in existing concrete segment casting processes, such as reliance on manual experience for liquid level determination, high error rate of single detection methods, poor stability under complex working conditions, and insufficient automation.

[0027] According to one aspect of the present invention, this approach utilizes a multimodal fusion mechanism combining dual-view visual images and weight sensing data to comprehensively analyze liquid level height, liquid distribution state, and weight change trends. This effectively avoids the uncertainties caused by a single information source, significantly improves the accuracy and reliability of liquid level detection, and effectively eliminates misjudgments caused by complex working conditions such as foam, splashing, and wall adhesion.

[0028] According to one aspect of the present invention, this approach, based on a primary first-view perspective and a secondary second-view perspective, can not only accurately detect changes in liquid level height but also identify the uniformity of liquid distribution within the mold. Compared to single-view detection methods, this approach can effectively identify abnormal pouring conditions such as localized accumulation and liquid displacement, thereby improving the ability to perceive the actual pouring state.

[0029] According to one aspect of the present invention, this solution effectively enhances robustness under complex working conditions by constructing a dual adaptive attention weight for multimodal features; by introducing an attention fusion method based on SE-Block, the weights of visual and weight information are dynamically adjusted. When visual interference is severe (such as foam, reflection), the system automatically increases the weight of weight information; when weight changes are insensitive or there are residues adhering to the wall, the role of visual information is enhanced, thereby achieving dynamic trust in multi-source information and intelligent decision-making, significantly improving the detection stability in complex industrial environments.

[0030] According to one aspect of the present invention, this approach employs a non-contact detection method, which is adaptable to high-viscosity and highly polluted industrial environments. It effectively avoids direct contact between the sensor and the concrete slurry, reducing the risks of wear, contamination, and blockage, and lowering the frequency of equipment maintenance and operating costs. It is particularly suitable for long-term continuous production scenarios involving high-viscosity, easily caking industrial liquids.

[0031] According to one aspect of the present invention, this approach can automatically and in real time determine the liquid level status during the pouring process, providing a reliable basis for subsequent pouring control or shutdown decisions. It replaces manual visual judgment and experience-based operation, reduces human error and labor intensity, and improves the overall automation, intelligence, and operational consistency of the production line.

[0032] According to one aspect of the present invention, the proposed multimodal liquid level detection scheme does not rely on complex or high-cost dedicated hardware, and can be seamlessly integrated with existing pouring equipment, weighing systems and industrial cameras. It has strong versatility and scalability, and is suitable for automated filling and quality control scenarios of various industrial liquid materials.

[0033] According to one aspect of the present invention, this approach achieves high-precision, robust, and non-contact intelligent detection of the liquid surface pouring status of pipe segments by introducing multimodal information fusion and adaptive attention mechanism. This effectively solves the problems of unstable detection and high misjudgment rate in complex industrial environments of existing technologies, and provides a reliable and practical technical solution for automated quality control of the liquid pouring process.

[0034] According to one aspect of the present invention, unlike traditional methods that simply splice or weight visual features and weight data at the decision-making end, this approach maps one-dimensional weight sensing signals into high-dimensional features and fuses them with dual-view visual features in a unified channel space in the early stage. This allows weight information to deeply participate in subsequent feature expression and interaction, significantly enhancing the consistency and discriminative ability of multimodal representation.

[0035] According to one aspect of the present invention, the spatiotemporal-modal collaborative adaptive attention mechanism (ST-MAA) introduces modal grouping compression and weight temporal dynamic embedding into channel attention, which can adaptively adjust the weights of the vision and weight channels according to the casting conditions, and achieve highly robust intelligent decision-making in complex industrial scenarios without the need for manual rules. Attached Figure Description

[0036] Figure 1 This is a flowchart illustrating the steps of the segment liquid level detection method based on multimodal sensing of the present invention. Figure 2 The flowchart shows the segment liquid level detection method based on multimodal sensing of the present invention. Figure 3 This is a flowchart of the spatiotemporal-modal collaborative adaptive attention mechanism of the present invention. Detailed Implementation

[0037] To more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the embodiments of the present solution will be described in detail below.

[0038] In describing embodiments of the present invention, the terms "longitudinal," "lateral," "upper," "lower," "front," "rear," "left," "right," "vertical," "horizontal," "top," "bottom," "inner," and "outer" express orientations or positional relationships based on the orientations or positional relationships shown in the relevant drawings. They are only for the convenience of describing the present invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, the above terms should not be construed as limitations on the present invention.

[0039] The present invention will now be described in detail with reference to the accompanying drawings and specific embodiments. The embodiments cannot be described in detail here, but the embodiments of the present invention are not limited to the following embodiments.

[0040] Combination Figure 1 and Figure 2 As shown, according to one embodiment of the present invention, a method for detecting the liquid level during segment casting based on multimodal sensing includes the following steps: S1. Collect concrete liquid surface image data inside the segment forming mold from different perspectives, and continuously collect and normalize the weight data of the segment forming mold. S2. Based on the time synchronization mechanism, the concrete liquid surface image data and weight data are mapped one-to-one, and a multimodal input sample containing visual features and weight features is constructed. S3. Extract visual and weight features from the multimodal input samples and fuse them to form the initial multimodal features; S4. Group the different modalities in the initial multimodal features and perform global information compression on each group to obtain global description vectors for different modalities; where the global description vectors for different modalities are the global description vectors for visual features and the global description vectors for weight features, respectively. S5. Collect continuous weight data to construct a weight temporal dynamic embedding vector, and fuse the weight temporal dynamic embedding vector with global description vectors of different modalities to construct attention weights that adaptively adjust the contribution of visual features and weight features. S6. Recalibrate the initial multimodal features based on attention weights to obtain multimodal weighted features, and output the detection results of whether the liquid level of the pipe segment is in place based on the multimodal weighted features.

[0041] According to one embodiment of the present invention, in step S1, the step of acquiring image data of the concrete liquid surface inside the segment forming mold based on different perspectives involves acquiring images of the concrete liquid surface inside the segment forming mold using a dual-perspective method. The first perspective is a top-down view directly in front of the segment forming mold, used to obtain the height profile of the concrete liquid surface. The concrete liquid surface image data obtained from the first perspective is represented as follows: ; in, Represents the real number field, used to represent the size of the tensor of the data. , These represent the height and width of the image, respectively. Indicates the number of channels; for RGB images, it's 3 channels. The subscript... Indicates the collection time; Furthermore, the second perspective is an angle tilted relative to the segment forming mold, used to obtain the uniformity of concrete surface distribution and local accumulation. In this embodiment, the second perspective is an oblique upward angle relative to the segment forming mold or a diagonal angle (such as left diagonal or right diagonal) relative to the segment forming mold. Therefore, the concrete surface image data obtained from the second perspective is represented as follows: .

[0042] According to one embodiment of the present invention, in step S1, the step of continuously collecting and normalizing the weight data of the segment forming mold is based on a weight sensor installed on the segment forming mold, which continuously collects the weight data of the segment forming mold at a fixed frequency during the segment casting process. Therefore, for a single weight sensor... The weight data corresponding to the time is: ; in, Represents the real number field.

[0043] In this embodiment, the weight sensor can be expanded to include multiple sensors.

[0044] Furthermore, to avoid the impact of differences in the units of weight data from different batches on the feature extraction process, the weight data can be further normalized. The normalized weight data is then represented as follows: ; in, This indicates the maximum weight value during the pouring process. This indicates the minimum weight value during the pouring process. This indicates the weight data collected in real time, and .

[0045] According to one embodiment of the present invention, in step S2, in the step of establishing a one-to-one correspondence between the concrete liquid surface image data and weight data based on a time synchronization mechanism, and constructing a multimodal input sample containing visual features and weight features, the time synchronization mechanism ensures that the acquisition time of the concrete liquid surface image data and weight data meets the following requirements: ; in, This indicates the acquisition time for the concrete liquid surface image data. This indicates the time of data collection for weight.

[0046] Therefore, by concatenating the data corresponding to the acquisition time, a multimodal input sample containing visual and weight features can be formed, which is represented as follows: .

[0047] According to one embodiment of the present invention, step S3, which involves extracting visual and weight features from the multimodal input samples and fusing them to obtain initial multimodal features, includes: S31. Based on multimodal input samples, image features from different perspectives are extracted and stitched together along the channel dimension to obtain fused multi-view visual features; in this embodiment, the concrete liquid surface image data collected from the first and second perspectives can be input into a visual coding network with shared parameters for feature extraction, and then the extracted image features from the two perspectives can be mapped to: ; in, This represents the image features extracted from first-person perspective images of the concrete liquid surface. This represents the image features extracted from second-person perspective images of the concrete liquid surface. This represents the feature extraction function based on CSPDarknet.

[0048] To obtain fused visual features from multiple perspectives, the image features from two perspectives are concatenated along the channel dimension, which is represented as: ; in, It represents the integration of visual features from multiple perspectives.

[0049] In this embodiment, the liquid region in the concrete liquid surface image data is segmented, and a liquid mask is output, represented as follows: ; in, , representing the pixel mask of the liquid region, This represents an instance segmentation operation to fuse visual features from multiple perspectives. As input, a convolutional network learns the spatial distribution of the liquid region, and outputs a binary mask with the same spatial dimensions as the image. , used to identify the location of liquid pixels in an image of a concrete liquid surface.

[0050] Therefore, based on the obtained liquid mask as a pixel-level spatial weight map, it is possible to fuse visual features from multiple perspectives. Element-wise weighting is performed to ensure that subsequent feature extraction and fusion focus on the actual liquid surface area.

[0051] S32. Extract weight features based on multimodal input samples; In this embodiment, since the weight data is a one-dimensional continuous scalar, a multilayer perceptron (MLP) is used for feature mapping to obtain the corresponding weight features, and the weight features are represented as follows: ; in, This represents the weight features extracted based on multimodal input samples. ReLU represents the nonlinear qualitative activation function. , These represent the weights (i.e., the weights of the fully connected layer). , These represent the bias terms.

[0052] S33. Broadcast the weight features in the spatial dimension to align them with the visual feature dimension, and then fuse them with the visual features to obtain the initial multimodal features. In this embodiment, since the feature dimension of the weight features is different from that of the visual information, to achieve deep fusion of visual and weight information, the weight features are broadcast-expanded in the spatial dimension to align them with the visual feature dimension, and then fused with the visual features. Therefore, the weight features based on the broadcast-expanded features are: ; in, This indicates the weight characteristics after broadcast extension. This indicates a repeating broadcast operation along the spatial dimension, that is, repeating the input weight feature vector in the height direction. H Repeat in the width direction. W This expands to a three-dimensional tensor with the same spatial resolution as visual features. These represent the spatial height and width of the visual feature map before fusion, respectively.

[0053] Furthermore, the fused multimodal initial features are represented as follows: ; in, This represents the initial features of the fused multimodal data.

[0054] Therefore, the multimodal initial features obtained in step S33 are represented as follows: ; ; ; ; ; in, This indicates the fusion of visual features from multiple perspectives. Represents the fusion mask features, representing the fusion of visual features. Element-wise weighting is performed to ensure that subsequent feature extraction and fusion focus on the actual liquid surface area. This represents the weight characteristic after broadcasting expansion in the spatial dimension. Visual features representing first-person perspective images of the concrete liquid surface. Visual features representing second-person perspective images of the concrete liquid surface. This represents the weight features extracted based on multimodal input samples. ReLU represents the nonlinear qualitative activation function. , They represent the weights, , These represent the bias terms.

[0055] like Figure 3 As shown, according to one embodiment of the present invention, to avoid the inability of simple splicing or fixed weighting to reflect the differences in importance of different modalities in different scenarios, the initial features of the multimodal structures are further optimized based on a spatio-temporal and modality-aware adaptive attention mechanism (ST-MAA). By explicitly introducing modality grouping compression and temporal dynamics embedding into the channel attention framework based on the spatio-temporal and modality-aware adaptive attention mechanism, dual adaptive control of the contribution of visual features and weight features is achieved.

[0056] Furthermore, in step S4, the step of grouping different modalities in the multimodal initial features and performing global information compression on each to obtain global description vectors for different modalities is achieved by modality grouping compression. In this embodiment, the global description vectors for different modalities are the global description vectors for visual features and the global description vectors for weight features, respectively; wherein, global average pooling is used to perform global information compression on different modalities in the multimodal initial features; specifically, for the multimodal initial features... Grouping by modality source and performing global average pooling on each: For the visual feature group, initial features of multimodal features are... Global average pooling is performed on the visual channels to obtain channel-level global description vectors, which in turn yield the corresponding global description vectors of visual features. ; in, Represents a global description vector of visual features; For the weight feature set, for the multimodal initial features Global average pooling is performed on the medium weight channels to obtain channel-level global description vectors, which in turn yield the corresponding global description vectors for the weight features. ; in, This represents the global description vector of weight features. This indicates the number of weight sensors.

[0057] According to one embodiment of the present invention, in step S5, the step of collecting continuous weight data to construct a weight temporal dynamic embedding vector, and fusing the weight temporal dynamic embedding vector with global description vectors of different modalities to construct attention weights that adaptively adjust the contribution of visual features and weight features, specifically, the weight temporal dynamic embedding vector is constructed by collecting weight data sequences of the current time and n historical frames. Taking 8 frames as an example, the obtained weight data sequence is as follows: It generates heavy temporal dynamic embedding vectors through a lightweight 1D convolutional encoder. It is used to characterize the stability state of the pouring process (such as continuous feeding, short pauses or complete stability).

[0058] Furthermore, the global description vectors of different modalities are fused with the weighted temporal dynamic embedding vectors respectively. Specifically, attention weights for each modality can be generated through a two-way parallel fully connected network, which can be expressed as: ; ; in, The attention weights representing visual features, and , The attention weights represent the weight features, and , This represents the activation function Sigmoid. , The learnable weight matrix represents the visual feature fusion processing branch (i.e., one fully connected network path). , This represents the learnable weight matrix of the weight feature fusion processing branch (i.e., another fully connected network).

[0059] Therefore, the attention weight, which adaptively modulates the contribution of visual features and weight features, is represented as follows: ; ; ; in, This represents the attention weights, which adaptively adjust the contribution of visual and weight features. The attention weights representing visual features, and , The attention weights represent the weight features, and , This represents the activation function Sigmoid. , The learnable weight matrix represents the visual feature fusion processing branch (i.e., one fully connected network path). , This represents the learnable weight matrix of the weight feature fusion processing branch (i.e., another fully connected network).

[0060] According to one embodiment of the present invention, in step S6, the multimodal weighted features are recalibrated based on attention weights to obtain multimodal weighted features, wherein the multimodal weighted features are represented as follows: ; in, This represents multimodal weighted features.

[0061] Furthermore, in step S6, the step of outputting the detection result of whether the segment liquid level is in place based on the multimodal weighted features, the multimodal initial features and the weight time-series dynamic embedding vector are jointly input into the end-to-end trained liquid level placement discrimination model. The liquid level placement discrimination model outputs the detection result of whether the segment liquid level is in place. The liquid level placement discrimination model is trained to: suppress the output of "placement in place" confidence when the weight time-series dynamic embedding vector indicates a recent feeding trend; and only output the "placement in place" detection result when the weight time-series dynamic embedding vector indicates that the weight has entered a sustained stable plateau period and the weight has reached a preset threshold. In this embodiment, the liquid level placement discrimination model adopts a lightweight structure of Global Average Pooling (GAP) followed by a single-layer fully connected layer (FC) to fuse the multimodal initial features and the weight time-series dynamic embedding vector and output the final detection result. Specifically, the fused multimodal weighted features are first obtained. Then, the multimodal weighted features are passed through a global average pooling layer to reduce their spatial dimension. Compression into channel statistics yields a global description vector. Then the global description vector The input is fed into a fully connected layer (FC), followed by a Sigmoid activation function, which directly outputs the probability confidence p∈[0,1] of whether the pouring is in place.

[0062] In this embodiment, dual adaptive regulation can be achieved based on the constructed multimodal weighted features: (1) Anti-interference mode switching: When the visual image is interfered with by foam, wall hangings or strong reflections, the attention weight model In the middle, attention weight can be automatically reduced. Increase attention weight This ensures that decision-making relies on reliable physical sensing. (2) Anti-pause misjudgment mechanism: When a short pause occurs during the pouring process (such as during the material feeding gap), even if the current weight data is close to the target value, if the weight time sequence is dynamically embedded in the vector If a recent feeding trend is observed, the "in place" confidence level is suppressed; only if the weight remains consistently stable (i.e., the weight time-series dynamic embedding vector) is the confidence level suppressed. Only when the plateau period is reached and a preset threshold is reached will a high-confidence "in place" signal be output.

[0063] Therefore, based on the constructed multimodal weighted features, under the action of the spatiotemporal-modal collaborative adaptive attention mechanism, there is no need to manually set modal weights or post-processing rules. The trust allocation of the scene is fully adapted through end-to-end training, which significantly improves the robustness and judgment accuracy of the system in complex industrial environments.

[0064] According to one embodiment of the present invention, the present invention provides an apparatus for detecting the liquid level during segment casting based on the aforementioned multimodal sensing method, comprising: A vision camera is used to acquire image data of the concrete liquid surface inside the segment forming mold from different perspectives. In this embodiment, two vision cameras are provided. The first vision camera is arranged from the front top view position of the segment forming mold to provide a first perspective. The second vision camera is set at an angle relative to the segment forming mold, and its oblique upward perspective relative to the segment forming mold or its diagonal perspective (such as left diagonal or right diagonal) relative to the segment forming mold.

[0065] Weight sensor, used to continuously collect weight data of the segment forming mold; The data preprocessing module is used to normalize the weight data; The time synchronization module is used to establish a one-to-one correspondence between concrete liquid surface image data and weight data based on the time synchronization mechanism, and to construct multimodal input samples containing visual features and weight features. The multimodal feature extraction and fusion module is used to extract visual and weight features from multimodal input samples and fuse them to obtain multimodal initial features; and to group different modalities in the multimodal initial features and perform global information compression on each to obtain global description vectors for different modalities; wherein the global description vectors for different modalities are the global description vectors for visual features and the global description vectors for weight features, respectively; and to collect continuous-time weight data to construct a weight temporal dynamic embedding vector, and to fuse the weight temporal dynamic embedding vector with the global description vectors for different modalities to construct attention weights that adaptively adjust the contribution of visual and weight features; and to recalibrate the multimodal initial features based on the attention weights to obtain multimodal weighted features; The liquid level determination module is used to output the detection result of whether the liquid level of the pipe segment is in place based on multimodal weighted features.

[0066] In this embodiment, the multimodal feature extraction and fusion module includes a multimodal liquid surface detection model, which is constructed based on the YOLOv8-Seg network. The multimodal liquid surface detection model comprises a visual feature encoder, a weight feature encoder, and a multimodal feature fusion module. The visual feature encoder is constructed based on a shared-parameter visual encoding network, used to extract viewpoint information from concrete liquid surface image data from different perspectives, and to stitch together the viewpoint information from different perspectives to output fused multi-view visual features. The weight feature encoder is constructed using a multilayer perceptron, used to map corresponding weight features from weight data, and to broadcast and expand the weight features in the spatial dimension to maintain consistency with the visual feature dimension. The multimodal feature fusion module is used to fuse the broadcast-expanded weight features and visual features to obtain multimodal initial features, and to convert the multimodal initial features into multimodal weighted features based on a spatio-temporal and modality-aware adaptive attention mechanism (ST-MAA).

[0067] Specific limitations regarding the device for detecting the liquid level in tunnel lining segments based on multimodal sensing can be found in the limitations of the method described above, and will not be repeated here. Each module in the device for detecting the liquid level in tunnel lining segments based on multimodal sensing can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device, or stored in the memory of a computer device as software, so that the processor can call and execute the operations corresponding to each module.

[0068] In this embodiment, the memory may be, but is not limited to, Random Access Memory (RAM), Read Only Memory (ROM), Programmable Read-Only Memory (PROM), Erasable Programmable Read-Only Memory (EPROM), Electrically Erasable Programmable Read-Only Memory (EEPROM), etc.

[0069] In this embodiment, the processor can be an integrated circuit chip with signal processing capabilities. The processor can be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it can also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components.

[0070] The above description is merely an example of a specific solution of the present invention. For any devices and structures not described in detail herein, it should be understood that they are implemented using common devices and methods already available in the art.

[0071] The above description is merely one embodiment of the present invention and is not intended to limit the invention. Various modifications and variations can be made to the invention by those skilled in the art. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.

Claims

1. A method for detecting the liquid level during segment casting based on multimodal sensing, characterized in that, Includes the following steps: S1. Collect concrete liquid surface image data inside the segment forming mold from different perspectives, and continuously collect and normalize the weight data of the segment forming mold. S2. Based on the time synchronization mechanism, the concrete liquid surface image data and weight data are mapped one-to-one, and a multimodal input sample containing visual features and weight features is constructed. S3. Extract visual and weight features from the multimodal input samples and fuse them to form the initial multimodal features; S4. Group the different modalities in the initial multimodal features and perform global information compression on each group to obtain global description vectors for different modalities; where the global description vectors for different modalities are the global description vectors for visual features and the global description vectors for weight features, respectively. S5. Collect continuous weight data to construct a weight temporal dynamic embedding vector, and fuse the weight temporal dynamic embedding vector with global description vectors of different modalities to construct attention weights that adaptively adjust the contribution of visual features and weight features. S6. Based on the attention weights, the initial multimodal features are recalibrated to obtain multimodal weighted features, and the detection result of whether the liquid level of the pipe segment is in place is output based on the multimodal weighted features.

2. The method for detecting the liquid level during segment casting based on multimodal sensing according to claim 1, characterized in that, In step S1, the step of acquiring image data of the concrete liquid surface inside the segment forming mold based on different perspectives involves acquiring images of the concrete liquid surface inside the segment forming mold using a dual-view approach. The first perspective is a top-down view directly in front of the segment forming mold, used to obtain the height profile of the concrete liquid surface. The concrete liquid surface image data obtained from the first perspective is represented as follows: ; in, Represents the real number field. , These represent the height and width of the image, respectively. Indicates the number of channels; The second perspective is an angle tilted relative to the segment forming mold, used to obtain the uniformity of concrete surface distribution and local accumulation; the concrete surface image data obtained from the second perspective is represented as follows: 。 3. The method for detecting the liquid level during segment casting based on multimodal sensing according to claim 2, characterized in that, In step S1, the step of continuously collecting and normalizing the weight data of the segment forming mold, the normalized weight data is expressed as follows: ; in, This indicates the maximum weight value during the pouring process. This indicates the minimum weight value during the pouring process. This indicates the weight data collected in real time, and .

4. The method for detecting the liquid level during segment casting based on multimodal sensing according to claim 3, characterized in that, Step S3, which involves extracting visual and weight features from the multimodal input samples and fusing them to obtain the initial multimodal features, includes: S31. Based on multimodal input samples, image features from different perspectives are extracted and stitched together in the channel dimension to obtain fused multi-view visual features; S32. Extract weight features based on multimodal input samples; S33. Broadcast the weight features in the spatial dimension to make them consistent with the visual feature dimension, and fuse them with the visual features to obtain multimodal initial features.

5. The method for detecting the liquid level during segment casting based on multimodal sensing according to claim 4, characterized in that, In step S33, the obtained multimodal initial features are represented as follows: ; ; ; ; ; in, This indicates the fusion of visual features from multiple perspectives. Indicates the fusion mask features, This represents the weight characteristic after broadcasting expansion in the spatial dimension. Visual features representing first-person perspective images of the concrete liquid surface. Visual features representing second-person perspective images of the concrete liquid surface. This represents the weight features extracted based on multimodal input samples. ReLU represents the nonlinear qualitative activation function. , They represent the weights, , These represent the bias terms.

6. The method for detecting the liquid level during segment casting based on multimodal sensing according to claim 5, characterized in that, In step S4, the step of grouping different modalities in the multimodal initial features and performing global information compression on each modality is to use global average pooling to compress global information on the different modalities in the multimodal initial features.

7. The method for detecting the liquid level during segment casting based on multimodal sensing according to claim 6, characterized in that, In step S4, the global description vector of visual features is represented as: ; in, Represents a global description vector of visual features; The global description vector of weight features is represented as follows: ; in, This represents the global description vector of weight features. This indicates the number of weight sensors.

8. The method for detecting the liquid level during segment casting based on multimodal sensing according to claim 7, characterized in that, In step S5, the step of collecting continuous weight data to construct a weight temporal dynamic embedding vector, and fusing the weight temporal dynamic embedding vector with global description vectors of different modalities to construct attention weights that adaptively adjust the contribution of visual features and weight features, is expressed as follows: ; ; ; in, This represents the attention weights, which adaptively adjust the contribution of visual and weight features. The attention weights representing visual features, and , The attention weights represent the weight features, and , This represents the activation function Sigmoid. , This represents the learnable weight matrix for the visual feature fusion processing branch. , This represents the learnable weight matrix of the weight feature fusion processing branch.

9. The method for detecting the liquid level during segment casting based on multimodal sensing according to claim 8, characterized in that, In step S6, where the initial multimodal features are recalibrated based on the attention weights to obtain multimodal weighted features, the multimodal weighted features are represented as follows: ; in, This represents multimodal weighted features.

10. An apparatus for detecting the liquid level during segment casting based on multimodal sensing as described in any one of claims 1 to 9, characterized in that, include: A vision camera is used to acquire image data of the concrete liquid surface inside the segment forming mold from different perspectives; Weight sensor, used to continuously collect weight data of the segment forming mold; The data preprocessing module is used to normalize the weight data; The time synchronization module is used to establish a one-to-one correspondence between concrete liquid surface image data and weight data based on the time synchronization mechanism, and to construct multimodal input samples containing visual features and weight features. The multimodal feature extraction and fusion module is used to extract visual and weight features from multimodal input samples and fuse them to form multimodal initial features. Furthermore, the initial multimodal features are grouped into different modalities and global information compression is performed on each group to obtain global description vectors for different modalities; wherein, the global description vectors for different modalities are global description vectors for visual features and global description vectors for weight features, respectively; and, weight data over continuous time is collected to construct a weight temporal dynamic embedding vector, and the weight temporal dynamic embedding vector is fused with the global description vectors for different modalities to construct an attention weight that adaptively adjusts the contribution of visual features and weight features; and, the initial multimodal features are recalibrated based on the attention weight to obtain multimodal weighted features. The liquid level determination module is used to output the detection result of whether the liquid level of the pipe segment is in place based on the multimodal weighted features.