Ship final assembly construction safety quality real-time monitoring method based on AI image recognition
Through AI image recognition technology of multi-source heterogeneous data, a multi-modal deep learning model is built, which solves the problem of low efficiency and accuracy in traditional monitoring methods, and realizes intelligent and real-time safety and quality monitoring of the ship assembly and construction process.
Patent Information
- Application Number
- CN202510726932.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-03
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2045-06-03
AI Technical Summary
Traditional safety and quality monitoring methods for ship assembly construction rely on manual inspection, which has low monitoring efficiency and accuracy, making it difficult to discover potential safety hazards and quality problems in real time, and fail to fully integrate and utilize a variety of data sources.
Using an AI image recognition method, multi-modal deep learning model is constructed by collecting multi-source heterogeneous data (high-definition visible light images, ultrasonic scanning signals, infrared temperature field and three-dimensional point cloud data), using attention mechanisms to perform feature fusion, identify quality hazards and safety risks, and generate real-time monitoring reports.
It has achieved full coverage monitoring of the ship assembly and construction process, improved the accuracy and real-time monitoring, provided scientific decision-making support, optimized resource allocation and emergency response.
Smart Images

Figure CN120234769A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of safety and quality monitoring, and particularly to a real-time monitoring method and storage medium for the safety and quality of ship general assembly construction based on AI image recognition. Background Art
[0002] Ship general assembly construction is a highly complex and comprehensive industrial process, involving multiple links and various processes, such as steel structure welding, equipment installation, painting, etc. In this process, ensuring safety and quality is crucial, and any minor defect or oversight may lead to serious consequences, such as structural failure, equipment malfunction, or even casualties. However, traditional safety and quality monitoring methods mainly rely on manual inspections and post-event checks, which have limitations. For example, manual inspections are easily affected by the experience and attention of inspectors, and it is difficult to detect potential safety hazards and quality problems in real time. Traditional methods are difficult to achieve full coverage monitoring of the entire construction process, and key areas and links are easily missed. A large amount of data, such as images, signals, temperatures, etc., will be generated during the ship general assembly construction process, but traditional methods have not fully integrated and utilized these data, resulting in low monitoring efficiency and accuracy. Summary of the Invention
[0003] The present invention aims to at least solve the technical problem of low monitoring efficiency and accuracy in the prior art, and particularly innovatively proposes a real-time monitoring method and storage medium for the safety and quality of ship general assembly construction based on AI image recognition.
[0004] To achieve the above object of the present invention, the present invention provides a real-time monitoring method for the safety and quality of ship general assembly construction based on AI image recognition, and the method includes: S1. Collect multi-source heterogeneous data, and preprocess the multi-source heterogeneous data to generate a spatio-temporally aligned multi-modal data set. The multi-source heterogeneous data includes high-definition visible light image data, ultrasonic scanning time-series signals, infrared temperature field data, and three-dimensional point cloud data; S2. Construct a multi-modal deep learning model based on the multi-modal data set, and use the multi-modal deep learning model to extract texture features, high-dimensional time-series vectors, and infrared temperature gradient features from the multi-source heterogeneous data; S3. Use the attention mechanism to perform weighted fusion on the texture features, high-dimensional time-series vectors, and infrared temperature gradient features to obtain a weighted fusion comprehensive feature vector. Based on the comprehensive feature vector, use the multi-modal deep learning model to identify quality hazards in the ship general assembly construction process, and output a quality monitoring result; S4. Based on the texture features and three-dimensional point cloud data, use the multi-modal deep learning model to identify whether personnel are wearing protective equipment devices, and output a safety monitoring result; S5. Obtain the movement trajectories of personnel and the movement trajectories of devices based on the high-definition visible light image data and the three-dimensional point cloud data, predict the collision risk probability using a multi-modal deep learning model based on the movement trajectories of personnel and the movement trajectories of devices, and perform safety warnings based on the collision risk probability; S6. Generate a real-time safety and quality monitoring report based on the quality monitoring results, safety monitoring results, and collision risk probability.
[0005] As an alternative embodiment of the present invention, optionally, the output of the quality monitoring results in step S3 includes: S301. Perform self-attention calculations on the texture features, high-dimensional time series vectors, and infrared temperature gradient features respectively; S302. Use the texture features as queries, the high-dimensional time series vectors and infrared temperature gradient features as keys and values respectively, and calculate the cross-modal attention weights accordingly; S303. Perform weighted summation on the texture features, high-dimensional time series vectors, and infrared temperature gradient features based on the cross-modal attention weights to obtain a fused comprehensive feature vector; S304. Input the fused comprehensive feature vector into a classifier, and the classifier classifies the fused comprehensive feature vector according to the preset quality hazard classification criteria to obtain the quality monitoring results.
[0006] As an alternative embodiment of the present invention, optionally, the output of the safety monitoring results in step S4 includes: S401. Extract the geometric feature vectors from the three-dimensional point cloud data, enhance the attention of the texture features and the geometric feature vectors, and perform weighted fusion to obtain an enhanced fused feature vector; S402. Input the enhanced fused feature vector into the two-layer fully connected network in the multi-modal deep learning model, and use the two-layer fully connected network to classify the enhanced fused feature vector to determine whether the personnel are wearing protective equipment devices to obtain the safety monitoring results.
[0007] As an alternative embodiment of the present invention, optionally, obtaining the movement trajectories of personnel and the movement trajectories of devices based on the high-definition visible light image data and the three-dimensional point cloud data in step S5 includes: S501. Align the preprocessed high-definition visible light image data with the three-dimensional point cloud data; S502. Identify the personnel bounding box and the device bounding box using image processing algorithms based on the image frames in the high-definition visible light image data; S503. Calculate the center point of the personnel bounding box based on the personnel bounding box, and calculate the center point of the device bounding box based on the device bounding box; S504. Perform ground segmentation on the three-dimensional point cloud data, and obtain the three-dimensional positions of the personnel and the equipment based on the segmented three-dimensional point cloud data; S505. Map the center point of the personnel bounding box and the center point of the equipment bounding box into the three-dimensional space to obtain the predicted positions of the personnel and the equipment; S506. Based on the predicted positions of the personnel and the equipment, combined with the timestamp information, use the Kalman filter algorithm to predict the trajectories of the personnel and the equipment for the predicted positions of the personnel and the equipment, and obtain the movement trajectories of the personnel and the equipment.
[0008] As an optional embodiment of the present invention, optionally, predicting the collision risk probability using a multimodal deep learning model based on the movement trajectories of the personnel and the equipment in step S5 includes: S507. Perform time alignment operations on the movement trajectories of the personnel and the equipment, and extract the corresponding time series features; S508. Based on the time series features, use the recurrent neural network module in the multimodal deep learning model to predict the changing trend of the relative distance between the personnel and the equipment; S509. Set a safety distance threshold. When the predicted relative distance is less than the safety distance threshold, it is determined that there is a collision risk, and the collision risk probability is calculated.
[0009] On the other hand, the present invention also provides a computer-readable storage medium, including: A memory, on which a computer program is stored; A processor, configured to execute the program in the memory to implement a real-time monitoring method for the safety and quality of ship general assembly construction based on AI image recognition.
[0010] Advantages of the present invention: The present invention comprehensively utilizes various heterogeneous data such as high-definition visible light images, ultrasonic scanning timing signals, infrared temperature fields, and three-dimensional point clouds, making up for the monitoring blind spots caused by traditional monitoring methods relying on a single data source. For example, high-definition visible light images can intuitively reflect the appearance quality of ships, ultrasonic scanning timing signals can detect internal structure defects, infrared temperature field data can monitor temperature changes in real time to prevent fires, and three-dimensional point cloud data can accurately measure spatial dimensions and positional relationships. The multi-modal deep learning model can automatically extract the features of different data and fuse them, avoiding the subjectivity and omissions of manual inspections and improving the accuracy of monitoring. The present invention also realizes the automatic extraction and fusion of the features of different modal data by constructing a multi-modal deep learning model specifically for processing multi-source heterogeneous data. The attention mechanism in the model can dynamically adjust the weights of the features of each modality, enabling the model to focus on the most critical information for the monitoring task in complex scenarios, further improving the accuracy of monitoring. The present invention realizes the intelligence of the monitoring process, can automatically analyze data, identify potential hazards, predict risks, and generate monitoring reports. This level of intelligence not only improves the accuracy and real-time performance of monitoring, but also provides scientific decision-making support for management personnel, optimizing resource allocation and emergency response.
[0011] Additional aspects and advantages of the present invention will be given in part in the following description, become apparent in part from the following description, or be learned through the practice of the present invention. Description of the Drawings
[0012] The above and / or additional aspects and advantages of the present invention will become apparent and be readily understood from the description of the embodiments in conjunction with the following drawings, in which: Figure 1 is a flowchart of the real-time monitoring method for the safety and quality of ship general assembly construction based on AI image recognition of the present invention. Detailed Embodiments
[0013] The embodiments of the present invention will be described in detail below. The examples of the embodiments are shown in the drawings, where the same or similar reference numerals represent the same or similar elements or elements with the same or similar functions throughout. The embodiments described below by referring to the drawings are exemplary and are only used to explain the present invention and should not be construed as a limitation of the present invention.
[0014] Embodiment 1
[0015] As Figure 1 shown, a real-time monitoring method for the safety and quality of ship general assembly construction based on AI image recognition, the method includes: S1. Collect multi-source heterogeneous data, and preprocess the multi-source heterogeneous data to generate a spatio-temporally aligned multi-modal dataset. The multi-source heterogeneous data includes high-definition visible light image data, ultrasonic scanning time-series signals, infrared temperature field data, and three-dimensional point cloud data; It should be noted that in step S1, the process of collecting multi-source heterogeneous data is completed by various sensors and devices deployed at the ship assembly construction site. The high-definition visible light image data is captured by a high-definition camera, the ultrasonic scanning time-series signals are generated by ultrasonic detection equipment, the infrared temperature field data is collected by an infrared thermal imager, and the three-dimensional point cloud data is obtained through a three-dimensional laser scanner. These data may have problems of time asynchrony and spatial misalignment during collection. Therefore, in the preprocessing stage, spatio-temporal alignment operations need to be performed to ensure the consistency and accuracy of the data. The spatio-temporal alignment operations include steps such as timestamp correction and spatial coordinate transformation to generate a spatio-temporally aligned multi-modal dataset.
[0016] S2. Build a multi-modal deep learning model based on the multi-modal dataset, and use the multi-modal deep learning model to extract texture features, high-dimensional time-series vectors, and infrared temperature gradient features from the multi-source heterogeneous data; It should be noted that step S2 needs to build a deep learning model based on a spatio-temporally aligned multi-modal dataset. This model processes different modal data through a parallel encoding network: uses a convolutional neural network (CNN) to extract texture features in high-definition visible light images (such as welding crack edges, coating peeling areas), uses a long short-term memory network (LSTM) or a Transformer module to capture high-dimensional time-series vectors in ultrasonic scanning signals (reflecting the dynamic changes of internal structure defects), and extracts gradient features of infrared temperature field data through a heat map analysis network (identifying abnormal temperature rise areas); the model training adopts a multi-task learning framework, and combines quality hazard annotation data (such as defect type, location, severity) for end-to-end optimization. The loss function fuses classification loss (cross-entropy) and regression loss (defect localization error) to ensure that the model has both accurate defect recognition and spatial localization capabilities.
[0017] S3. Use the attention mechanism to perform weighted fusion on the texture features, high-dimensional time-series vectors, and infrared temperature gradient features to obtain a weighted fusion comprehensive feature vector. Based on the comprehensive feature vector, use the multi-modal deep learning model to identify quality hazards in the ship assembly construction process and output quality monitoring results; It should be noted that step S3 realizes the intelligent identification of quality hazards through a multi-modal attention mechanism: First, self-attention calculations are respectively performed on the texture features, high-dimensional time-series vectors, and infrared temperature gradient features to mine the key information within each modality; Subsequently, using the texture features as queries, and the high-dimensional time-series vectors and infrared temperature gradient features as keys and values, cross-modal attention weights are calculated to represent the importance of the associations between different modalities; Based on these weights, the three types of features are weighted and summed to generate a comprehensive feature vector that fuses multi-source information. This vector can dynamically adjust the contribution degrees of each modality to adapt to different defect types (such as cracks, soldering joints, coating defects); Finally, the comprehensive vector is input into a classifier, and the quality monitoring results are output according to preset criteria, realizing the complementarity of multi-modal information and the precise positioning of defects, and significantly improving the recognition robustness of quality hazards in complex scenarios.
[0018] S4. Based on the texture features and the three-dimensional point cloud data, use the multi-modal deep learning model to identify whether a person is wearing a protective equipment device and output the safety monitoring results; It should be noted that step S4 realizes the identification of protective equipment wearing by fusing texture features and three-dimensional point cloud geometric features. First, geometric feature vectors (such as local geometric attributes like normal vectors, curvatures, etc.) are extracted from the three-dimensional point cloud data, and attention enhancement processing is performed on the texture features and geometric features: The attention weights of the two sets of features are calculated through the self-attention mechanism to dynamically adjust the saliency of different feature channels and enhance the perception ability of key areas of the protective equipment (such as the edges of safety helmets, reflective strips on protective clothing). Subsequently, the enhanced features are weighted and fused to generate a comprehensive feature vector that contains visual and spatial information. This vector is input into a two-layer fully connected network. The first layer uses the rectified linear unit (ReLU) activation function for non-linear feature mapping, and the second layer combines the Softmax function to output a safety monitoring probability vector. Finally, a safety monitoring result is generated through threshold judgment (such as a probability > 0.8 is determined as correct wearing). This process compensates for the limitations of texture recognition under occlusion and lighting changes through geometric constraints, and at the same time uses texture details to improve the recognition accuracy of geometric features for small target devices, forming a complementary multi-modal detection mechanism.
[0019] S5. Based on the high-definition visible light image data and the three-dimensional point cloud data, obtain the movement trajectories of the person and the device, predict the collision risk probability using the multi-modal deep learning model based on the movement trajectories of the person and the device, and conduct safety warnings based on the collision risk probability; It should be noted that step S5 realizes dynamic collision risk prediction through multi-modal data fusion. First, the preprocessed high-definition visible light image and the 3D point cloud are aligned in space and time. The image object detection algorithm (such as YOLOv5) is used to identify the bounding boxes of personnel and equipment, and the coordinates of their center points are calculated. At the same time, the ground segmentation of the point cloud is performed to extract the 3D spatial coordinates of personnel and equipment, and the image center point is mapped to the 3D space to form the estimated position. Combining the timestamp information, the Kalman filter algorithm is used to predict the trajectory of the estimated position to generate a continuous motion trajectory. Subsequently, the time alignment of the trajectories of personnel and equipment is performed, and the time series features including speed, acceleration, and direction angle are extracted and input into a recurrent neural network (such as LSTM) to predict the relative distance change within the future time window. When the predicted relative distance is less than the safety threshold (such as 1.5 meters), the collision risk probability is calculated by combining the Sigmoid function. This probability is jointly determined by the current actual distance, the distance change amount, and the sensitivity coefficient, and finally a collision risk value between 0 and 1 is output to achieve millisecond-level safety warning. The entire process overcomes the limitations of a single sensor through the complementarity of images and point clouds, and uses a time series model to capture the motion pattern to form an end-to-end trajectory prediction and risk assessment framework.
[0020] S6. Generate a real-time safety and quality monitoring report based on the quality monitoring results, safety monitoring results, and collision risk probability.
[0021] It should be noted that step S6, as the system output terminal, realizes the panoramic perception of the safety and quality situation through the decision-level fusion of multi-source heterogeneous data. This step receives in real time the quality hidden danger identification results output by S3 (including the probability vector and spatial positioning of 6 types of quality problems such as welding defects and structural deformations), the safety monitoring results output by S4 (including the confidence levels of 5 types of violations such as personnel identity binding information and not wearing safety helmets and reflective vests), and the collision risk probability output by S5 (fusion of the personnel-equipment motion trajectory prediction results and the 3D space distance threshold determination). The spatio-temporal alignment mechanism is used to synchronously process the three major data sources, and the confidence intervals of each monitoring index are dynamically updated through the Bayesian network to construct a four-dimensional monitoring matrix (quality, personnel safety, equipment safety, collision risk). Based on the preset hierarchical warning rule library, a monitoring report including heat map distribution, risk ranking list, and 3D scene reproduction is automatically generated. Among them, the quality hidden danger module uses the AR overlay technology to mark the defect positions, the safety monitoring module integrates the historical violation frequency statistics, and the collision warning module embeds the dynamic trajectory prediction visualization component. The report generation cycle is less than 200 ms, supporting multi-terminal adaptive rendering, and forming a closed-loop management link of "monitoring - warning - disposal - feedback".
[0022] In summary, this embodiment provides an efficient and accurate real-time monitoring method for the safety and quality of ship general assembly construction. By collecting multi-source heterogeneous data and using a deep learning model for feature extraction and fusion, comprehensive monitoring of quality hazards, safety protection, and collision risks during the ship general assembly construction process is achieved. This method not only improves the accuracy and real-time performance of monitoring but also provides scientific decision-making support for managers, helping to optimize resource allocation and emergency response. In addition, the embodiment of the present invention further improves the intelligent level of monitoring by constructing a multi-modal deep learning model specifically for processing multi-source heterogeneous data, making the monitoring process more automated and efficient.
[0023] As an alternative embodiment of the present invention, optionally, the output of the quality monitoring result in step S3 includes: S301. Perform self-attention calculations on the texture feature, high-dimensional time series vector, and infrared temperature gradient feature respectively; It should be noted that in step S301, independent relationship modeling and feature enhancement of the three heterogeneous features are performed through the self-attention mechanism. For the texture feature, this step divides the multi-scale texture feature map extracted by the convolutional neural network into several spatial regions, and by calculating the attention weights of each region with all regions, local significant features are enhanced and irrelevant backgrounds are suppressed; for the high-dimensional time series vector, a self-attention mechanism in the time dimension is used to capture periodic defect features in the ultrasonic scan signal, and query-key-value triples are generated by sliding the time window to highlight waveform patterns strongly related to quality hazards in the time series data; the self-attention calculation of the infrared temperature gradient feature focuses on the spatial distribution of the temperature field, and an abnormal temperature gradient region is strengthened by constructing a pixel-level attention matrix while maintaining the structural information of the global temperature distribution. The self-attention calculations of the three features all adopt the scaled dot-product attention formula, generate query, key, and value vectors through learnable linear transformations, obtain the attention weight matrix after normalization by the softmax function, and finally output enhanced texture features, time series feature vectors, and temperature gradient features after feature recalibration, providing feature representations with intra-modal context awareness capabilities for subsequent cross-modal fusion.
[0024] S302. Use the texture feature as the query, the high-dimensional time series vector and the infrared temperature gradient feature as the key and value respectively, and calculate the cross-modal attention weights accordingly; It should be noted that step S302 realizes the deep interaction and fusion of multi-source features through a cross-modal attention mechanism. In this process, the texture feature is used as the query vector to guide the model to actively explore the potential associations with the high-dimensional temporal vector (key) and the infrared temperature gradient feature (value). Specifically, when implementing, first, the three-modal features are respectively passed through a learnable linear projection layer to generate a query matrix Q, a key matrix K, and a value matrix V. Among them, Q is generated from the texture feature, K is generated from the temporal feature, and V is generated from the temperature gradient feature. Subsequently, the cross-modal attention weights are calculated using the scaled dot-product attention formula, that is, the transpose of Q is multiplied by K and then divided by the square root of the feature dimension, and then normalized by the softmax function to obtain the attention distribution matrix. This matrix reflects the dynamic association strength between the texture feature and the other two-modal features. Among them, a high attention weight indicates that the corresponding temporal signal or temperature gradient in the region has a significant impact on the judgment of quality hazards. Finally, the attention weight matrix is applied to the value matrix V through a weighted sum operation to generate an enhanced texture feature representation that fuses temporal and temperature information. This feature not only retains the spatial details of the original texture but also incorporates the dynamic characteristics of the temporal data and the abnormal patterns of the temperature field, providing a multi-modal joint representation for the accurate identification of quality hazards.
[0025] S303. Perform a weighted sum on the texture feature, the high-dimensional temporal vector, and the infrared temperature gradient feature based on the cross-modal attention weights to obtain a fused comprehensive feature vector. It should be noted that step S303 realizes the deep fusion of multi-modal features through a weighted sum operation. In this process, the cross-modal attention weight matrix is respectively applied to the texture feature vector, the high-dimensional temporal vector, and the infrared temperature gradient feature vector, and the original features are dynamically adjusted through element-wise multiplication. Specifically, the attention weight coefficient reflects the importance distribution of different modal features in the quality hazard identification task. The larger the weight, the more critical the feature region is for the current decision. The weighted texture feature strengthens its association with the temporal signal and the abnormal patterns of the temperature field. The temporal vector selects the dynamic patterns highly correlated with the spatial texture change through the attention mechanism, while the temperature gradient feature is recalibrated to highlight the temperature abnormal regions strongly related to quality hazards such as welding quality and structural deformation. Finally, the three weighted feature vectors are concatenated along the channel dimension to form a comprehensive feature vector that fuses multi-modal information. This vector not only retains the spatial resolution and semantic information of the original features but also realizes information complementarity and redundancy suppression between modalities through the attention mechanism, providing a quality hazard feature representation with strong discriminative power for the classifier.
[0026] S304. Input the fused comprehensive feature vector into a classifier, and the classifier classifies the fused comprehensive feature vector according to a preset quality hazard classification standard to obtain the quality monitoring result.
[0027] It should be noted that in step S304, the fused comprehensive feature vector is input into a classifier, which performs non - linear mapping and decision boundary division on the feature vector based on a preset quality hazard classification standard. The classifier adopts a deep neural network structure, which learns typical feature patterns of quality hazards from historical data through supervised learning. The multi - layer non - linear activation functions inside the network can capture complex interaction relationships between features. In the inference stage, the comprehensive feature vector undergoes dimensional transformation through a fully - connected layer and then outputs the probability distribution of various quality hazards through the Softmax function. The preset classification standard defines the judgment thresholds for quality hazard categories such as welding defects, structural deformation, and assembly misalignment. The classifier compares the probability values with the thresholds and outputs the final quality monitoring result vector, which not only contains the hazard category label with the highest probability but also retains the confidence scores of each potential category, providing an interpretable diagnostic basis for quality control personnel.
[0028] As an alternative embodiment of the present invention, optionally, the expression for outputting the quality monitoring result in step S304 is:
[0029] Wherein, α i represents the attention weight coefficient of the i th modal feature, exp ( ) represents the exponential function, W i T represents the transpose of the learnable attention weight vector corresponding to the i th modal feature, Concat ( ) represents the concatenation operation, F1 represents the texture feature vector, F2 represents the high - dimensional time - series vector, F3 represents the infrared temperature gradient feature, W j T represents the transpose of the learnable attention weight vector corresponding to the j th modal feature, Q result represents the finally output quality monitoring result vector, Classifier ( ) represents the classifier function.
[0030] As an alternative embodiment of the present invention, optionally, the output of the safety monitoring result in step S4 includes: S401. Extract the geometric feature vector from the three - dimensional point cloud data, perform attention enhancement on the texture feature and the geometric feature vector, and perform weighted fusion to obtain an enhanced fused feature vector; S402. Input the enhanced fused feature vector into the two-layer fully connected network in the multimodal deep learning model, use the two-layer fully connected network to classify the enhanced fused feature vector, determine whether the personnel are wearing protective equipment, and obtain the safety monitoring result.
[0031] It should be noted that step S402 inputs the enhanced fusion feature vector into a two-layer fully connected network. The first layer of the network performs a linear transformation on the input features through a learnable weight matrix, and introduces nonlinearity by applying the ReLU activation function to mine complex patterns related to the identification of protective equipment. The second layer of the network further performs dimensional compression and feature reorganization on the first layer output, and finally outputs the probability value of personnel wearing protective equipment in a standardized manner through the Sigmoid function. The two-layer structure effectively integrates texture details and three-dimensional geometric information through step-by-step feature abstraction, and the network parameters are optimized through historical data supervised training, so that the output probability can accurately reflect the personnel protection status and provide real-time quantitative basis for safety supervision.
[0032] As an optional embodiment of the present invention, optionally, the expression for obtaining the safety monitoring result in step S402 is:
[0033] Among them, F fused represents the fused feature vector, α img represents the weight of the texture feature vector, F img represents the texture feature vector, α pcl represents the weight of the geometric eigenvector, F pcl represents the geometric eigenvector, P safety represents the safety monitoring probability vector, Softmax ( ) represents the normalized exponential function, W3 represents the output layer weight matrix, ReLU ( ) represents the rectified linear unit activation function, W2 represents the weight matrix of the second-layer fully connected network, W1 represents the weight matrix of the first-layer fully connected network, b1 represents the bias vector of the first-layer fully connected network, b2 represents the bias vector of the second-layer fully connected network, and b3 represents the bias vector of the output layer.
[0034] As an optional embodiment of the present invention, optionally, in step S5, obtaining the movement trajectory of a person and the movement trajectory of a device based on the high-definition visible light image data and the three-dimensional point cloud data includes: S501, aligning the preprocessed high-definition visible light image data with the three-dimensional point cloud data; S502, identifying a person boundary box and a device boundary box using an image processing algorithm based on an image frame in the high-definition visible light image data; It should be noted that in step S502, a deep learning-based object detection algorithm is used to analyze each frame of the preprocessed high-definition visible light image sequence. This algorithm extracts multi-scale features in the image through a convolutional neural network, generates candidate regions using the anchor box mechanism, and filters out the optimal detection boxes through non-maximum suppression. Considering the characteristics of the shipbuilding scenario, the model is specially trained with personnel features such as safety helmets and work clothes, as well as device features such as cranes and welding equipment, and can accurately output the bounding box coordinates, confidence levels, and class labels of personnel and devices. The detection process integrates image context information, effectively differentiates the overlapping of personnel and the occlusion of devices in dense operation areas, ensures high robustness even under complex lighting conditions, and provides an accurate spatial positioning basis for subsequent trajectory tracking.
[0035] S503. Calculate the center point of the personnel bounding box based on the personnel bounding box, and calculate the center point of the device bounding box based on the device bounding box; S504. Perform ground segmentation on the three-dimensional point cloud data, and obtain the three-dimensional positions of personnel and devices based on the segmented three-dimensional point cloud data; It should be noted that in step S504, a three-dimensional point cloud segmentation algorithm is used to separate the ground and non-ground of the original point cloud. This algorithm first optimizes the data density through voxel grid downsampling, then uses cloth simulation filtering to remove outlier noise points, and then extracts a continuous ground point set based on normal vector estimation and region growing algorithms. After constructing a digital elevation model using the ground point set, through height threshold segmentation and Euclidean clustering analysis, the personnel activity area and the space occupied by devices are accurately distinguished. For personnel position extraction, the algorithm combines prior knowledge of human height, filters out point clusters within a specific height range in the segmented non-ground point cloud, and uses the principal component analysis method to calculate the centroid of the point cluster as the three-dimensional coordinates of the personnel; device positioning is determined by matching the geometric features of the device and screening with volume thresholds, and combining the point cloud density distribution characteristics to determine the center position of the device. The entire process effectively addresses the complex terrain and dynamic occlusion problems in the shipbuilding scenario through a multi-scale feature fusion strategy, achieving centimeter-level accuracy in spatial positioning.
[0036] S505. Map the center point of the personnel bounding box and the center point of the device bounding box into three-dimensional space to obtain the estimated positions of personnel and devices; It should be noted that step S505 realizes the mapping conversion from two-dimensional image coordinates to three-dimensional space positions through multi-modal data fusion. First, a conversion matrix from the image coordinate system to the world coordinate system is constructed using camera calibration parameters. Combining the pixel coordinates of the center point of the detected person bounding box and the center point of the device bounding box in the image frame, the initial three-dimensional coordinates are calculated through the pinhole camera model. Subsequently, the depth information in the three-dimensional point cloud data is introduced for optimization, and a weighted average is performed on the image mapping result and the ground elevation data obtained from the point cloud segmentation, where the weight coefficient is jointly determined by the point cloud density and the image detection confidence. For the complex spatial structure of the shipbuilding scenario, a nearest neighbor search algorithm based on a spatial grid is adopted, and the nearest neighbor points in the image mapping point cloud are used as spatial anchor points to construct a non-linear optimization model to iteratively correct the initial coordinates. Finally, through Kalman filtering processing in the time series, instantaneous noise interference is eliminated, and the smooth and continuous estimated positions of the person and the device are output. This process effectively solves the scale ambiguity problem of the monocular camera and enables the three-dimensional positioning accuracy to reach the centimeter level, providing a reliable data basis for subsequent trajectory prediction.
[0037] S506. Based on the estimated positions of the person and the device, combined with the timestamp information, use the Kalman filter algorithm to predict the trajectories of the estimated positions of the person and the device, and obtain the movement trajectories of the person and the device.
[0038] It should be noted that step S506 uses the extended Kalman filter algorithm to achieve spatio-temporal correlation prediction of the movement trajectory. First, the estimated positions of the person and the device obtained from the three-dimensional space mapping are used as the observation input, and a six-dimensional state vector including position coordinates, velocity vectors, and acceleration components is constructed, and a time series correlation is established through the timestamp information. In the prediction stage, the state vector is recursively estimated based on the constant velocity model, and a process noise covariance matrix is introduced to characterize the uncertainty of the movement. In the update stage, the predicted value is corrected using the observed value at the current moment, and the trust weights of the observed values in different dimensions are dynamically adjusted through the observation noise covariance matrix. For the common non-linear movement trajectories in the shipbuilding scenario, an iterative Kalman filter method is adopted to locally linearize the state transition function and the observation function through the Jacobian matrix, improving the adaptability to complex movement patterns. Finally, a smooth trajectory sequence including timestamp information is output to achieve sub-second real-time update, providing a high-precision spatio-temporal reference for subsequent collision risk prediction.
[0039] As an alternative embodiment of the present invention, optionally, in step S5, predicting the collision risk probability using a multi-modal deep learning model based on the movement trajectories of the person and the device includes: S507. Align the movement trajectories of the person and the device in time, and extract the corresponding time series features; It should be noted that in step S507, the dynamic time warping algorithm is used to perform time alignment on the movement trajectories of personnel and equipment. By constructing a time warping function, the differences in sampling frequencies of different sensors are compensated. The locally weighted scatterplot smoothing method is used to process the noise mutation points in the trajectories. Based on the aligned trajectory data, multi-dimensional time series features including three-dimensional coordinates, movement speed, and direction angle change rate are extracted. The sliding time window method is used to generate fixed-length feature sequences, where the window length is dynamically adjusted according to the typical movement speeds of personnel and equipment in the shipbuilding scenario to ensure that the feature sequences can completely cover the key movement patterns before potential collisions. The feature extraction process incorporates physical constraints such as speed threshold filtering and trajectory smoothing to improve the robustness of the features to outliers, and finally generates a standardized feature vector sequence with spatio-temporal consistency.
[0040] S508. Based on the time series features, use the recurrent neural network module in the multi-modal deep learning model to predict the changing trend of the relative distance between personnel and equipment; It should be noted that in step S508, the gated recurrent unit (GRU) network is used to deeply encode the time series features. This network adaptively captures the long-term spatio-temporal dependencies in the trajectory features through a gating mechanism. Specifically, time series features such as the three-dimensional coordinates, velocity vectors, and acceleration magnitudes of personnel and equipment are input into the GRU layer, and the movement patterns in the time dimension are modeled using the hidden state transfer mechanism. The network performs a linear transformation on the input features through a learnable parameter matrix, combines the hidden state of the previous moment, and dynamically adjusts the information flow using the update gate and the reset gate to generate a hidden state vector containing the temporal evolution law. On this basis, the hidden state of the final moment is mapped to the relative distance change prediction value through a fully connected layer. The prediction time window length is set according to the typical reaction times of personnel and equipment in the shipbuilding scenario, and the mean squared error loss function and the Adam optimizer are used to optimize the network parameters to ensure that the prediction results can accurately reflect the dynamic change trend of the relative distance in the future time period.
[0041] S509. Set a safety distance threshold. When the predicted relative distance is less than the safety distance threshold, it is determined that there is a collision risk, and the collision risk probability is calculated.
[0042] It should be noted that step S509 realizes safety warning through a collision risk probability calculation formula. The formula uses the Sigmoid function to map the linear output to the interval (0, 1) to represent the collision probability. Among them, the safety distance threshold is used as a key parameter to define the risk boundary. The current actual distance and the predicted relative distance change amount jointly determine the change trend of the collision probability. The distance change sensitivity coefficient is used to adjust the influence weight of the distance change on the probability, so that when the predicted relative distance is less than the safety threshold, the system can dynamically calculate the collision risk probability and trigger a hierarchical warning mechanism. The larger the probability value, the higher the collision risk, and the warning level will increase accordingly, thus realizing refined safety control based on spatio-temporal prediction.
[0043] As an alternative embodiment of the present invention, optionally, the expression for calculating the relative distance change trend in step S508 is:
[0044] where Δ d t+Δt represents the relative distance change amount from the current moment t to the future moment t +Δ t . If Δ d t+Δt <0, it means that the distance between the person and the device is decreasing; if Δ d t+Δt >0, it means that the distance between the person and the device is increasing. f RNN ( ) represents the function representation of the multi-modal deep learning model, and h t person represents the trajectory feature vector of the person at the moment t , and h t device represents the trajectory feature vector of the device at the moment t , and Δ t represents the predicted time window length.
[0045] As an alternative embodiment of the present invention, optionally, the expression for calculating the collision risk probability in step S509 is:
[0046] where p collision represents the collision risk probability, σ ( ) represents the Sigmoid activation function, d safe represents the safety distance threshold, d t the actual distance between the person and the device at the current moment t Δd t+Δt Indicates from the current moment t to a future moment t +Δ t of the relative distance change amount, τ Indicates the distance change sensitivity coefficient, P t person Indicates the three-dimensional position coordinates of the person, P t device Indicates the three-dimensional position coordinates of the device.
[0047] Embodiment 2
[0048] A computer-readable storage medium, comprising: A memory, on which a computer program is stored; A processor, configured to execute the program in the memory to implement the real-time monitoring method for the safety and quality of ship general assembly construction based on AI image recognition in Embodiment 1.
[0049] It should be noted that the electronic device in the embodiments of the present disclosure includes a processor and a memory for storing processor-executable instructions. Among them, the processor is configured to implement the real-time monitoring method for the safety and quality of ship general assembly construction based on AI image recognition described in any one of the foregoing when executing the executable instructions.
[0050] Here, it should be pointed out that the number of processors can be one or more. At the same time, in the electronic device of the embodiments of the present disclosure, an input device and an output device may also be included. Among them, the processor, the memory, the input device, and the output device may be connected through a bus or in other ways, which are not specifically limited here.
[0051] As a computer-readable storage medium, the memory can be used to store software programs, computer-executable programs, and various modules, such as: the programs or modules corresponding to the real-time monitoring method for the safety and quality of ship general assembly construction based on AI image recognition in the embodiments of the present disclosure. The processor executes various functional applications and data processing of the electronic device by running the software programs or modules stored in the memory.
[0052] The input device can be used to receive input numbers or signals. Among them, the signal can be a key signal related to the user settings and function control of the device / terminal / server. The output device may include a display device such as a display screen.
[0053] Although the embodiments of the present invention have been shown and described, those of ordinary skill in the art can understand that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the claims and their equivalents.
Claims
1. A real-time monitoring method for the safety and quality of ship general assembly construction based on AI image recognition, characterized in that, The method includes: S1. Collect multi-source heterogeneous data, and preprocess the multi-source heterogeneous data to generate a spatio-temporally aligned multi-modal dataset. The multi-source heterogeneous data includes high-definition visible light image data, ultrasonic scanning time-series signals, infrared temperature field data, and three-dimensional point cloud data; S2. Build a multi-modal deep learning model based on the multi-modal dataset, and use the multi-modal deep learning model to extract texture features, high-dimensional time-series vectors, and infrared temperature gradient features from the multi-source heterogeneous data; S3. Use an attention mechanism to perform weighted fusion on the texture features, high-dimensional time-series vectors, and infrared temperature gradient features to obtain a comprehensively weighted fusion feature vector. Based on the comprehensively weighted fusion feature vector, use the multi-modal deep learning model to identify quality hazards in the ship assembly construction process and output a quality monitoring result; S4. Based on the texture features and three-dimensional point cloud data, use the multi-modal deep learning model to identify whether personnel are wearing protective equipment and output a safety monitoring result; S5. Obtain the movement trajectories of personnel and equipment based on the high-definition visible light image data and three-dimensional point cloud data. Based on the movement trajectories of personnel and equipment, use the multi-modal deep learning model to predict the probability of collision risk, and conduct safety warnings based on the collision risk probability; S6. Generate a real-time safety and quality monitoring report based on the quality monitoring result, safety monitoring result, and collision risk probability.
2. The real-time monitoring method for the safety and quality of ship general assembly construction based on AI image recognition according to claim 1, characterized in that, The output of the quality monitoring result in step S3 includes: S301. Perform self-attention calculations on the texture features, high-dimensional time-series vectors, and infrared temperature gradient features respectively; S302. Use the texture features as queries, the high-dimensional time-series vectors and infrared temperature gradient features as keys and values respectively to calculate cross-modal attention weights; S303. Perform weighted summation on the texture features, high-dimensional time-series vectors, and infrared temperature gradient features based on the cross-modal attention weights to obtain a comprehensively fused feature vector; S304. Input the comprehensively fused feature vector into a classifier, and the classifier classifies the comprehensively fused feature vector according to a preset quality hazard classification standard to obtain the quality monitoring result.
3. The real-time monitoring method for the safety and quality of ship general assembly construction based on AI image recognition according to claim 2, wherein The expression for outputting the quality monitoring result in step S304 is: ; Among them, α i represents the attention weight coefficient of the i th modal feature, exp ( ) represents the exponential function, and W i T represents the transpose of the learnable attention weight vector corresponding to the i th modal feature, Concat ( ) represents the concatenation operation, F1 represents the texture feature vector, F2 represents the high-dimensional time series vector, F3 represents the infrared temperature gradient feature, and W j T represents the transpose of the learnable attention weight vector corresponding to the j th modal feature, and Q result represents the vector of the final output quality monitoring result, Classifier ( ) represents the classifier function.
4. The real-time monitoring method for the safety and quality of ship general assembly construction based on AI image recognition according to claim 1, characterized in that The output of the safety monitoring result in step S4 includes: S401. Extract geometric feature vectors from the three-dimensional point cloud data, perform attention enhancement on the texture features and geometric feature vectors, and perform weighted fusion to obtain an enhanced fused feature vector; S402. Input the enhanced fused feature vector into a two-layer fully connected network in the multi-modal deep learning model, and use the two-layer fully connected network to classify the enhanced fused feature vector to determine whether personnel are wearing protective equipment to obtain the safety monitoring result.
5. The real-time monitoring method for the safety and quality of ship general assembly construction based on AI image recognition according to claim 4, wherein The expression for obtaining the safety monitoring result in step S402 is: ; Among them, F fused represents the fused feature vector, α img represents the weight of the texture feature vector, F img represents the texture feature vector, α pcl represents the weight of the geometric feature vector, F pcl represents the geometric feature vector, P safety represents the security monitoring probability vector, Softmax ( ) represents the normalization exponential function, W3 represents the output layer weight matrix, ReLU ( ) represents the rectified linear unit activation function, W2 represents the weight matrix of the second fully connected network, W1 represents the weight matrix of the first fully connected network, b1 represents the bias vector of the first fully connected network, b2 represents the bias vector of the second fully connected network, and b3 represents the bias vector of the output layer.
6. The real-time monitoring method for the safety and quality of ship general assembly construction based on AI image recognition according to claim 1, characterized in that, Obtaining the movement trajectories of personnel and the movement trajectories of equipment based on the high-definition visible light image data and the three-dimensional point cloud data in step S5 includes: S501. Performing an alignment operation on the preprocessed high-definition visible light image data and the three-dimensional point cloud data; S502. Identifying the personnel bounding box and the equipment bounding box by using an image processing algorithm based on the image frames in the high-definition visible light image data; S503. Calculating the center point of the personnel bounding box based on the personnel bounding box, and calculating the center point of the equipment bounding box based on the equipment bounding box; S504. Performing ground segmentation on the three-dimensional point cloud data, and obtaining the three-dimensional positions of personnel and the three-dimensional positions of equipment based on the segmented three-dimensional point cloud data; S505. Mapping the center point of the personnel bounding box and the center point of the equipment bounding box into three-dimensional space to obtain the predicted positions of personnel and the predicted positions of equipment; S506. Based on the predicted positions of personnel and the predicted positions of equipment, combining the timestamp information, and using the Kalman filtering algorithm to perform trajectory prediction on the predicted positions of personnel and the predicted positions of equipment to obtain the movement trajectories of personnel and the movement trajectories of equipment.
7. The real-time monitoring method for the safety and quality of ship general assembly construction based on AI image recognition according to claim 6, characterized in that, Predicting the collision risk probability by using a multi-modal deep learning model based on the movement trajectories of personnel and the movement trajectories of equipment in step S5 includes: S507. Performing time alignment operation on the movement trajectories of personnel and the movement trajectories of equipment, and extracting corresponding time series features; S508. Based on the time series features, using the recurrent neural network module in the multi-modal deep learning model to predict the changing trend of the relative distance between personnel and equipment; S509. Setting a safety distance threshold, and when the predicted relative distance is less than the safety distance threshold, determining that there is a collision risk and calculating the collision risk probability.
8. The real-time monitoring method for the safety and quality of ship general assembly construction based on AI image recognition according to claim 7, characterized in that, The expression for calculating the changing trend of the relative distance in step S508 is: ; where, Δ d t+Δt represents the relative distance change from the current moment t to the future moment t +Δ t . If Δ d t+Δt < 0, it means the distance between the person and the device is decreasing; if Δ d t+Δt > 0, it means the distance between the person and the device is increasing. f RNN ( ) represents the function representation of the multimodal deep learning model, h t person represents the trajectory feature vector of the person at the moment t . h t device represents the trajectory feature vector of the device at the moment t . Δ t represents the predicted time window length.
9. The real-time monitoring method for the safety and quality of ship general assembly construction based on AI image recognition according to claim 7, characterized in that, The expression for calculating the collision risk probability in step S509 is: ; Among them, p collision represents the collision risk probability, σ ( ) represents the Sigmoid activation function, d safe represents the safety distance threshold, d t the current moment t the actual distance between the person and the device, Δ d t+Δt represents from the current moment t to the future moment t +Δ t the relative distance change amount, τ represents the distance change sensitivity coefficient, P t person represents the three-dimensional position coordinates of the person, P t device represents the three-dimensional position coordinates of the device.
10. A computer-readable storage medium, characterized in that, Including: A memory on which a computer program is stored; A processor for executing the program in the memory to implement the real-time monitoring method for the safety and quality of ship general assembly construction based on AI image recognition according to any one of claims 1 to 9.
Citation Information
Patent Citations
Intelligent inspection system based on digital power plant
CN114373245A
Ship target fine granularity identification method and device based on heterogeneous data sharing semantics
CN118154962A
Industrial defect detection method, system and device and storage medium
CN118967672A
Road disease automatic identification method based on sparse self-encoding network
CN119963927A
Methods for target detection based on visible cameras, infrared cameras, and lidars
US20240355105A1
Cited By
Radio frequency access door based on multi-band cooperative control and article passing identification method
CN120449912A
Channel environment safety monitoring and early warning method and system based on multi-sensor fusion
CN120538606A
Multi-stage blasting rock mass structural surface intelligent identification method based on deep learning
CN120580524A
A multi-stage blasting rock mass structural plane intelligent identification method based on deep learning
CN120580524B
Pest and disease detection model construction method and pest and disease detection method based on space-air-ground multi-source data fusion
CN120932126A