MFMPosePredict model for mobile formwork attitude prediction and prediction method
The MFMPosePredict model improves construction safety and precision by integrating instance and joint interaction modeling for mobile formwork systems, addressing pose abnormalities and enhancing real-time monitoring.
Patent Information
- Application Number
- CN202510372447.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-27
- Publication Date
- 2025-07-15
AI Technical Summary
During the construction process, mobile mold frames are susceptible to external factors to cause abnormal posture or deformation, which affects construction accuracy and may cause safety accidents. It is difficult for the existing technology to predict and adjust the posture in real time.
The MFMPosePredict model is adopted to detect the mobile model frame in the image through multi-objective pose estimation, and use the interaction information between target instances and inter-joint joints, and combine the cross-instance interaction modeling module, the cross-joint interaction modeling module and the adaptive feature fusion module to generate reliable pose representations to achieve real-time pose prediction and early warning.
It significantly improves the accuracy of mobile mold frame posture recognition, ensures construction safety and efficiency, monitors the mold frame status in real time through the AI system and promptly warns of potential safety hazards.
Smart Images

Figure CN120318752A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of intelligent construction safety management, and particularly to an MFMPosePredict model and a prediction method for predicting the attitude of a moving formwork. Background Art
[0002] Moving Formwork Method (MFM) construction is an efficient construction method used in bridge construction, especially suitable for the on-site casting of continuous girder bridges and prestressed concrete box girders. This method can significantly reduce the installation and disassembly time of formwork, improve construction efficiency, and ensure the quality and safety of the structure.
[0003] With the rapid development of transportation infrastructure construction, road and bridge construction has become an important part of it. As an advanced construction technology, the moving formwork has been widely used in road and bridge construction due to its unique advantages. The moving formwork construction technology can significantly improve construction efficiency, contribute to improving construction safety and project quality, can reduce construction costs, and has good adaptability and flexibility, and can be applied to different types of bridge structures and complex construction conditions, such as viaducts, high-pier highway bridges, and cross-sea bridges, etc.
[0004] During the construction process, the moving formwork is easily affected by external factors (such as wind force, temperature, load changes, etc.), resulting in abnormal attitude or deformation, which in turn affects the construction accuracy and may even lead to safety accidents. Introducing an AI model can predict the attitude change of the formwork in real time. Summary of the Invention
[0005] Aiming at the problems existing in the prior art, in order to help construction workers timely master and predict the attitude change of the moving formwork and make corresponding adjustments to ensure more accurate construction results; help construction workers identify abnormal formwork attitudes in advance, issue early warnings, and prevent safety accidents caused by out-of-control attitudes, the present invention proposes an MFMPosePredict model for predicting the attitude of a moving formwork. In the model, the input image first undergoes feature extraction through a backbone network. The feature data extracted from the backbone network is processed in two paths: one path passes through a key point decoder, and the decoded result is F joint ; the other path passes through an instance decoder, and its output is fused with the feature data extracted from the backbone network to obtain a result of F inst ; then, F inst passes through an instance-joint relationship branch to obtain a feature F I→J ; F joint passes through a joint-instance relationship branch to obtain a feature F J→I ; finally, F I→J and F J→IAll are input into the pose decoder to generate a reliable pose representation.
[0006] Based on the above solution, in the instance-joint relationship branch, the instance representation F inst will be used as the input of the cross-instance interaction modeling module together with the position embedding F pos After being processed by the cross-instance interaction modeling module, Then, it is fused with the joint information F joint through an adaptive feature fusion module based on channel attention; subsequently, a cross-joint interaction modeling module is used to model the interaction between joints, and the learned interaction feature F I→J is obtained.
[0007] Based on the above solution, in the joint-instance relationship branch, the joint representation F joint and the position embedding F pos are used as the input of the cross-joint interaction modeling module. After being processed by the cross-joint interaction modeling module, After that, it is fused with the instance representation F inst through an adaptive feature fusion module based on channel attention. Subsequently, through the cross-instance interaction modeling module, the feature F J→I is obtained.
[0008] This application also provides a method for predicting the pose of a mobile formwork, which uses the above model.
[0009] A server includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps of the above method for predicting the pose of a mobile formwork are implemented.
[0010] A computer-readable storage medium stores a computer program. When the computer program is executed by a processor, the steps of the above method for predicting the pose of a mobile formwork are implemented.
[0011] An MFMPosePredict model for mobile formwork attitude prediction in the present invention uses multi-object pose estimation to detect the mobile formwork in an image and locate key points for each object. The MFMPosePredict model proposed in the present invention simultaneously considers the interaction information between target instances or joints, and utilizes the complementarity of the interaction between target instances and the interaction between joints to improve the performance of multi-object pose estimation. Relationship information is crucial for accurately positioning joints because the correlation between joints helps to learn pose structure information. The cross-instance interaction modeling module in the model incorporates representations and position embeddings containing instance position information into the calculation, which can enhance the perceptual representation, enhance the distinguishability of individual information, clearly distinguish the postures of multiple mobile formwork targets, and improve the accuracy of the model. The cross-joint interaction modeling module in the model uses calculations such as feature reshaping and transposition to integrate the joint information of the mobile formwork, generate joint features, and incorporate the original joint features into the aggregated features, which can enhance the information discrimination ability of the features and ultimately make the posture recognition more accurate. At the same time, the MFMPosePredict model of the present invention uses an adaptive feature fusion module, which is implemented based on the channel attention mechanism, can activate features related to the pose estimation task, activate and represent the pose features of the mobile formwork, and help improve the recognition accuracy of the model. The present invention uses AI technology to perform safety detection on the attitude of the mobile formwork during mobile formwork construction, which can significantly improve the safety and efficiency of construction. By using computer vision algorithms, the AI system can monitor the state of the formwork in real time and give early warnings of potential safety hazards. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] Figure 1 It is the overall architecture diagram of the MFMPosePredict model of this application;
[0013] Figure 2 It is the structural diagram of the cross-instance interaction modeling module in the model of this application;
[0014] Figure 3 It is the structural diagram of the cross-joint interaction modeling module in the model of this application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0015] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0016] Embodiment 1
[0017] The pose of the moving formwork refers to its position and angular state in space, including levelness, verticality, torsion degree, etc. In actual construction, the formwork may tilt horizontally or shift vertically during movement or construction, and even twist as a whole or in part, resulting in structural instability.
[0018] Given the abnormal pose and easy deformation defects of the moving formwork in construction, it is particularly important to introduce an AI pose prediction model. By predicting the pose changes of the formwork through the AI model, adjustments can be made in advance to ensure construction accuracy. Monitor the formwork pose in real time, warn of potential safety risks, and prevent accidents.
[0019] The present invention provides an MFMPosePredict model for predicting the pose of a moving formwork. The essence of the model is multi-object pose estimation. Multi-object pose estimation (MOPE) can detect all relevant objects (the moving formwork in this invention) in an image and locate key points for each object.
[0020] Given an image J containing multiple objects, the goal of multi-object pose estimation is to estimate the position of the pose key points of each object, which can be expressed as:
[0021]
[0022] where m represents the number of objects in the image, n represents the number of key points for each object, and K j (i) represents the j-th pose key point of the i-th target object in image J. m and n represent the number of target objects in image J and the number of key nodes for each target object, respectively.
[0023] Given an image where H and W represent the height and width of the image respectively. The image contains N target object instances (the moving formwork in this invention). The proposed model aims to detect the joint points of the n-th instance. In this invention, the joints are the main beam arms, legs, lifting legs, guide beams, etc. of the target object (i.e., the moving formwork), and each part is similar to the key parts such as the arms, thighs, and head of the human body. The joint points of the n-th target object can be represented by the joint heatmap where represents the k-th joint heatmap with size h×w, and K is the number of joint points.
[0024] In a multi-object scenario, rich interaction information, including interaction between target object instances and interaction between joints, plays a crucial role. Interaction between target object instances helps to locate the current target object by means of information from other target objects, while interaction between joints helps to accurately locate the current joint by using information from other joints.
[0025] The MFMPosePredict model proposed by the present invention simultaneously considers the interaction information between target instances or joints, and utilizes the complementarity of the interaction between target instances and the interaction between joints to improve the performance of multi-object pose estimation. Relationship information is crucial for accurately positioning joints because the correlation between joints helps to learn pose structure information.
[0026] The overall framework of the MFMPosePredict model of the present invention takes an image as input and uses a pre-trained feature encoder to extract the visual features of the image, denoted as (where c represents the number of channels, h represents the image height, and w represents the image width). The visual feature F fuses the information of the background and the target object.
[0027] To effectively model the interaction between target instances and joints, we use an instance decoder and a joint decoder to extract the instance-aware representation (where represents the feature with dimension d of the nth target object, h represents the height, and w represents the width) and the joint-aware representation (where represents that the nth target object has K joint features).
[0028] Specifically, as Figure 1 shown, in the MFMPosPredict model of the present invention, the input image first undergoes feature extraction through the backbone network. The feature data extracted from the backbone network is processed in two paths. One path is through the keypoint decoder, and the decoded result is F joint ; the other path passes through the instance decoder, and its output is fused with the feature data extracted from the backbone network to obtain the result F inst . F inst and F joint respectively contain information about a single target object and independent joints. The subsequent modules of the model take F inst and F joint as inputs to model the interaction between target instances and joints.
[0029] First, for the instance-joint relationship branch (IJR), the instance representation F inst will be combined with the position embedding F pos , as the input of the cross-instance interaction modeling module (CIM), and after being processed by the cross-instance interaction modeling module, we get Then, since lacks information about joints, is adaptively fused with the joint information F jointFusion is performed to enhance joint information. Subsequently, a cross-joint interaction modeling module (CJM) is used to model the interaction between joints, and the learned interaction feature F is obtained. I→J The formula for this process is denoted as:
[0030]
[0031] where represents the correlation between instances. The channel attention block and the convolutional layer Conv(·) constitute the adaptive feature fusion module, and [·,·] represents the concatenation operation.
[0032] Since F I→J is the feature obtained first by the target instance-interaction modeling module (CIM) and then by the cross-joint interaction modeling module (CJM), it contains the interaction information at the target instance level and the joint level. Therefore, F I→J can be regarded as the interaction information from coarse to fine.
[0033] The principle of the joint-instance relationship branch (JIR) is similar to that of the instance-joint relationship branch (IJR). Specifically, the joint representation F joint and the position embedding F pos , as the input of the cross-joint interaction modeling module (CJM), are processed by the cross-joint interaction modeling module (CJM) to obtain After that, through the adaptive feature fusion module (ADFM) based on channel attention, it is fused with the instance representation F inst , and then, through the cross-instance interaction modeling module (CIM), the feature F J→I is obtained. The feature F J→I extracted by the JIR branch can be regarded as the interaction information from fine to coarse granularity.
[0034] To make full use of the complementary interaction information, the features of both the instance-joint relationship branch (IJR) and the joint-instance relationship branch (JIR) are input into the pose decoder to generate a reliable pose representation. The formula is as follows:
[0035]
[0036] As a specific implementation, the implementation details of the cross-instance interaction modeling module (CIM), the cross-joint interaction modeling module (CJM), and the adaptive feature fusion module are as follows:
[0037] 1. Cross-instance interaction modeling module (CIM)
[0038] As Figure 2As shown, the input of the Cross-Instance Interaction Modeling module (CIM) includes the instance representation F inst and the position embedding F containing instance position information pos . We incorporate F pos into CIM to enhance the position information of each object. The instance is represented as a series of center maps, and the coordinates of the maximum value are used as the position information
[0039] To facilitate the modeling of appearance and position correlations, first, a reshape operation is performed on F inst to generate three sets of instance representations ( and ), where is generated by transposing the vector in addition to reshaping. Then, the matrix multiplication between and is calculated to obtain the appearance correlation
[0040] Meanwhile, the position embedding F pos is transposed to obtain For F ops , another branch is denoted as and has the same value as F pos . Then, the matrix multiplication between and is calculated to obtain the position correlation
[0041] Then, the matrices of S inst and S pos are added element-wise, and the result is subjected to the Softmax operation to aggregate and obtain the relevant attention map The calculation formula for the above process is:
[0042]
[0043] Here, i and j represent the elements in the i-th row and j-th column respectively
[0044] Each element Att inst in
[0045] measures the interaction degree between any two instances inst This enables the model to localize the joints of the current instance based on the information of other instances. The matrix multiplication of Att and is performed to generate the relationship-based feature Since the information about the instances ininst , reshape , and the result is the same as F inst perform matrix addition, and the resulting value is denoted as
[0046]
[0047] These calculations are enhancements of instance-aware representations, which can enhance the distinguishability of individual information.
[0048] The cross-instance interaction modeling module adopts the representation F inst and the position embedding F containing instance position information pos integrated into the calculation, which can improve the enhancement of perceptual representation, enhance the distinguishability of individual information, enable the postures of multiple mobile formwork targets to be clearly distinguished, and improve the accuracy of the model.
[0049] 2. Cross-Joint Interaction Modeling Module (CJM)
[0050] As Figure 3 shown, the cross-joint interaction modeling module (CJM) takes joint features as input. Subsequently, 3 groups of 1×1 convolutional layers are used as joint feature extractors and applied to F joint to generate three groups of discriminative joint features respectively. To facilitate the modeling of the correlation between joints, the first group of features is reshaped and transposed to obtain The second group of features is reshaped to obtain The third group of features is reshaped to obtain The complete formula is as follows:
[0051]
[0052] Each row in , or each row in , represents a complete joint. Similar to the modeling of the correlation between instances, the spatial affinity between joints is calculated through multiplication and Softmax to generate the relevant attention map The formula is as follows:
[0053]
[0054] Att joint Each element in represents the interaction degree between joint k and i. This design enables the proposed model to locate the current joint based on the information of other joints.
[0055] Finally, first reshape Att joint , and then the reshaped result and Perform matrix multiplication to integrate joint-related information and generate relationship-based joint features, denoted as RF joint In order to enhance the discriminative ability of single joint information, we also transform the original joint feature F joint Integrate it into the aggregation feature. Specifically, it is to integrate RF joint and the original joint feature F joint Perform matrix addition calculation to obtain output features
[0056]
[0057] The cross-joint interaction modeling module uses calculations such as feature reshaping and transposition to integrate the joint information of the mobile mold frame, generate joint features, and integrate the original joint features into the aggregated features, which can enhance the information recognition ability of the features and ultimately make posture recognition more accurate.
[0058] 3. Adaptive Feature Fusion Module (ADFM)
[0059] The adaptive feature fusion module (ADFM) is implemented based on the channel attention mechanism, which aims to activate features related to the pose estimation task. Taking the instance-joint relationship (IJR) branch as an example, we first use global average pooling (GAP) to transform and F joint The spatial information of is compressed into a channel indicator. Then, through linear projection, the sigmoid nonlinear function is used to generate channel attention specifically used to activate important instance or joint features. Finally, the activated features are modulated and fused through 1×1 convolution. The formula for the whole process is as follows:
[0060]
[0061] Among them, Att c represents channel attention. [·,·] and ⊙ denote concatenation and channel multiplication, respectively. σ(·), MLP(·), and GAP(·) refer to sigmoid function, linear projection, and global average pooling, respectively.
[0062] The model of the present invention adopts an adaptive feature fusion module (ADFM), which is implemented based on a channel attention mechanism and can activate features related to the posture estimation task, so that the posture features of the mobile mold frame are activated and represented, thereby helping to improve the recognition accuracy of the model.
[0063] 4. Posture decoder
[0064] Pose decoder: To obtain robust pose features, channel attention and spatial attention are introduced into the pose decoder to adaptively highlight task-related features. Then, two convolutional layers are used to generate joint heatmaps as follows:
[0065]
[0066] where denotes the enhanced pose features by integrating interaction information from different objects. SA(·) and CA(·) refer to the spatial attention mechanism and channel attention mechanism in CBAM respectively. F I→J and F J→I are relationship-based features extracted by the instance-joint relationship (IJR) branch and joint-instance relationship (JIR) branch respectively.
[0067] This model adopts a pose decoder, which introduces channel attention and spatial attention into the pose decoder, helping to obtain more robust pose features of the mobile formwork and highlighting task-related features. The accuracy of this model is improved.
[0068] Embodiment 2
[0069] Based on the model in Embodiment 1, this application provides a method for constructing a prediction model (MFMPosePredict) for the pose of the mobile formwork, including the following steps:
[0070] (1) Data acquisition
[0071] Install high-definition cameras or industrial cameras at key positions of the mobile formwork, such as the four corners and support structures of the mobile formwork, to capture images and videos of the formwork.
[0072] (2) Data preprocessing and feature extraction
[0073] Image processing: Use computer vision techniques such as edge detection and shape recognition to extract features of the formwork from the images, such as the contour, auxiliary legs, bearers, connection parts, etc.
[0074] (3) Training of the MFMPosePredict model
[0075] Use the processed pictures in step (2) as the input of the MFMPosePredict model for model training. During training, the loss function includes an instance segmentation loss and a joint heatmap loss. The former is used to extract robust instance representations, and the latter prompts the model to learn task-related information.
[0076]
[0077] Among them, α is a balance factor used to adjust the importance of different loss terms. is the focal loss between the instance center map output by the model and the ground truth. is the mean square error between the joint heat map output by the model and the ground truth.
[0078] To train this model, we collected a total of 2000 pictures of the moving formwork construction of road and bridge construction in Yantai, Shandong. We divided all the data into a training set, a test set and a validation set according to the ratio of 2:1:1. Use the training set data to train the model and update the model parameters through backpropagation. Regularly use the validation set data to evaluate the model performance and adjust the parameters. When the performance metrics (such as accuracy, recall, mAP, etc.) of the model on the validation set tend to be stable and no longer improve significantly, training can be considered stopped. In particular, set the maximum number of training epochs to 500, and automatically stop when the training reaches the maximum number of epochs. After stopping, save the optimal model for prediction.
[0079] We evaluated the model constructed in Example 2 on the Yantai dataset, adopted the average precision (AP) as the evaluation metric, and demonstrated the joint localization accuracy at different thresholds by reporting AP, AP50, and AP75. The evaluation results show that we achieved excellent results in this dataset, with the average precision (AP) reaching 69.0, AP50 being 89.8, and AP75 being 76.4.
[0080] The above are only the preferred embodiments of the present application and are not used to limit the present application. For those skilled in the art, the present application can have various changes and modifications. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.
Claims
1. An MFMPosePredict model for mobile formwork attitude prediction, characterized in that, In the described model, the input image first undergoes feature extraction through the backbone network; the feature data extracted from the backbone network is processed in two paths: one path passes through the keypoint decoder, and the decoded result is F joint ; the other path passes through the instance decoder, and its output is fused with the feature data extracted from the backbone network to obtain the result F inst ; then, F inst passes through the instance-joint relationship branch to obtain the feature F I→J , F joint passes through the joint-instance relationship branch to obtain the feature F J→I ; finally, F I→J and F J→I are both input into the pose decoder to generate a reliable pose representation.
2. The MFMPosePredict model for mobile formwork attitude prediction according to claim 1, wherein In the instance-joint relationship branch, the instance represents F inst will be combined with the position embedding F pos As the input of the cross-instance interaction modeling module, after being processed by the cross-instance interaction modeling module, we get Then Through the adaptive feature fusion module based on channel attention and the joint information F joint are fused; Subsequently, a cross-joint interaction modeling module is used to model the interaction between joints, and the learned interaction feature F I→J is obtained.
3. The MFMPosePredict model for mobile formwork attitude prediction according to claim 1, wherein In the joint-instance relationship branch, the joint represents F joint and the positional embedding F pos are used as the input of the cross-joint interaction modeling module. After being processed by the cross-joint interaction modeling module, subsequently, it is fused with the instance representation F inst through the adaptive feature fusion module based on channel attention. Subsequently, through the cross-instance interaction modeling module, the feature F J→I is obtained.
4. The MFMPosePredict model for mobile formwork attitude prediction according to claim 2 or 3, characterized in that The input of the cross-instance interaction modeling module includes an instance representation F inst and a positional embedding F containing instance location information pos ; First, reshape F inst to generate three sets of instance representations ( and ), where is generated by transposing the vector in addition to reshaping; then calculate the matrix multiplication between and to obtain the appearance correlation Meanwhile, transpose the positional embedding F pos , to obtain For F pos , another branch is denoted as and F pos have the same value; then calculate with in the matrix multiplication to obtain the positional correlation Then, for S inst and S pos perform matrix addition, i.e., element-wise addition, and perform the Softmax operation on the result to aggregate and obtain the relevant attention map Meanwhile, for Att inst and perform matrix multiplication to generate relationship-based features Finally, is reshaped, and the reshaped result is added to F inst by matrix addition, and the resulting value is denoted as 5. The MFMPosePredict model for mobile formwork attitude prediction according to claim 2 or 3, characterized in that, The cross-joint interaction modeling module takes joint features as input; subsequently, three groups of 1×1 convolutional layers are used as joint feature extractors and applied to F joint to respectively generate three groups of discriminative joint features; The first set of features is reshaped and transposed to obtain The second set of features is reshaped to obtain The third set of features is reshaped to obtain Each line in , or each line in , represents a complete joint; similar to the correlation modeling between instances, Each line in calculates the spatial affinity between joints by multiplication and Softmax to generate a correlation attention map After that, Att joint is reshaped, and then the reshaped result and are subjected to matrix multiplication to integrate joint-related information and generate relationship-based joint features, denoted as RF joint ; Finally, perform matrix addition on RF joint and the original joint feature F joint to obtain the output feature 6. The MFMPosePredict model for mobile formwork attitude prediction according to claim 2 or 3, characterized in that In the adaptive feature fusion module of the instance-joint relationship branch, first, global average pooling (GAP) is used to and F joint 's spatial information is compressed into a channel indicator; then, through linear projection, followed by using the sigmoid non-linear function, channel attention dedicated to activating important instance or joint features is generated; finally, the activated features are modulated and fused through 1×1 convolution.
7. The MFMPosePredict model for mobile formwork attitude prediction according to claim 6, wherein The calculation formula of the adaptive feature fusion module is as follows: Among them, Att c represents channel attention; [·,·] and ⊙ represent concatenation operation and channel multiplication respectively; σ(·), MLP(·) and GAP(·) refer to sigmoid function, linear projection and global average pooling respectively.
8. A prediction method for the attitude of a moving formwork, characterized in that, Use the model according to any one of claims 1-7.
9. A server, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, the steps of the prediction method for the mobile formwork attitude as described in claim 8 are implemented.
10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, the steps of the prediction method for the mobile formwork attitude as described in claim 8 are implemented.