A visual confirmation method for leak detection during the filling process of a liquid tank truck

Through deep learning and posture estimation technology, the leakage measurement and operation during the filling process of liquid tankers is monitored in real time, and the safety hazards caused by the reliance on manual operation in the existing technology are solved, and intelligent safety monitoring and early warning are realized.

CN119380282BActive Publication Date: 2025-08-19NANJING BOILER & PRESSURE VESSEL SUPERVISION & INSPECTION INST +4
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411548328.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-01
Publication Date
2025-08-19
Estimated Expiration
2044-11-01

AI Technical Summary

Technical Problem

During the filling process of existing liquid tankers, leakage measurement depends on personnel operational norms, and there is a risk of missing or incomplete actions, which makes it difficult to effectively monitor safety hazards.

Method used

Deep learning and attitude estimation technology are used to monitor the tank truck filling process through the camera, and the completion status of the leak detection action is identified and confirmed, including object detection and attitude recognition, and the completion status of the crane tube connection, spraying soapy water and observation sequence are determined.

Benefits of technology

It realizes intelligent monitoring of leak-testing actions, automatically detects missing actions, improves the safety of the filling process, builds a safety warning mechanism, and avoids safety accidents.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119380282B_ABST
    Figure CN119380282B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for visually confirming the leak detection action during the filling process of a liquid tank truck. During the filling process of the liquid tank truck, the method identifies and confirms the operator's leak detection action on the crane pipe connection area based on the video captured by the camera, and is used to monitor and confirm whether the personnel's leak detection operation before the tank truck is filled is completed in accordance with the required specifications. The method includes: using a trained deep learning model to perform target detection, identify the vehicle, operating box, crane pipe, spray bottle and on-site personnel, and confirm the crane pipe connection scene; using a deep learning model to perform posture estimation, and identify the body movements of the person holding the spray bottle and spraying soapy water near the crane pipe; using multi-target association detection, according to the relative position of the person and the crane pipe connection area, identify the continuous process of the person observing the bubble situation near the crane pipe.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer vision recognition technology, and in particular to a visual intelligent confirmation method for leak detection during the filling process of a liquid tank truck. Background Art

[0002] When transporting hazardous liquid chemicals by tanker, filling the tanker is a critical operation with unique risks. Due to the physical and chemical properties of hazardous liquids, a leak can lead to serious workplace accidents, resulting in significant injuries and losses. Therefore, ensuring the tightness of the connection between the tanker and filling equipment, as well as the tightness of the pipeline, is crucial.

[0003] For hazardous chemicals such as Class A and Class B liquids, the use of crane pipes is an important means to ensure safe production, protect the environment, and comply with regulations. The crane pipe is equipped with valves and control devices that can accurately control the flow and stop of the liquid. Operators can control it from a safe position to avoid direct contact with hazardous liquids, which improves the safety and accuracy of operations. During general operations, after connecting the crane pipe, the operator will perform a leak test on the connection area by spraying soapy water. The specific method is to spray soapy water on the connection and observe whether bubbles are generated to determine whether the connection is tight and whether the airtightness meets the standards. This detection process relies on the operator's operational standardization, and there may be missing or incomplete actions. Summary of the Invention

[0004] In order to overcome the shortcomings of the prior art, the purpose of the present invention is to propose a visual intelligent confirmation method for leak detection during the filling process of a liquid tanker. By processing the monitoring video data of the filling area, detecting and confirming the correct completion of the leak detection process, it is possible to more intelligently and safely monitor personnel to complete the tank leak detection operation and effectively avoid the occurrence of safety accidents. The method monitors the operation area through a camera, collects the operating actions and continuous processes between the operator and the tanker and crane in real time, and uses image recognition technology to confirm whether the key targets appear in the picture. Through this automated visual monitoring method, the completion status of the leak detection action can be effectively identified and determined, the absence of the leak detection action can be automatically detected, and the safety of the operation can be improved.

[0005] In order to achieve the above object, the technical solution adopted by the present invention is as follows:

[0006] A visual confirmation method for leak detection during the filling process of a liquid tanker. During the period when personnel are performing leak detection operations on the tanker during the filling process, the monitoring video frame rate collected by a fixed-position camera is 30 frames per second, and one frame is taken every five frames. For each frame of the image, image processing, target and posture detection, and scene status determination are performed based on a deep learning method. The first step is to identify the position of the vehicle, locate the operating box and the crane pipe target, determine that the crane pipe is in a state of connection with the interface in the tanker operating box, and determine the position of the crane pipe connection area in the picture; the second step is to detect the posture of the person and the spray bottle target in the picture, and identify the action posture of the person holding the spray bottle; the third step is to identify the relative position of the person and the operating box and the crane pipe connection, and determine that the person is near the crane pipe connection; the fourth step is to identify the spraying action of the person holding the spray bottle; the fifth step is to identify the continuous process of the person staying and observing at the crane pipe connection. Specifically including the following steps:

[0007] S1: Capture on-site monitoring videos of container filling and leak detection operations under various lighting conditions, capturing scene images including the vehicle, operating box, crane pipe, spray bottle, and on-site personnel, as well as environmental images at a 7:1 ratio.

[0008] S2: Capture on-site monitoring videos of container filling and leak detection operations under various lighting conditions, capturing scene images including the vehicle, operating box, crane hose, spray bottle, and on-site personnel, as well as environmental images at a 7:1 ratio.

[0009] S3: Uses "transfer learning" and an adaptive learning rate algorithm to train the target detection model, enabling target detection and positioning of vehicles, operating boxes, crane hoses, spray bottles, and on-site personnel.

[0010] S4: According to the relative position and angle between the crane pipe and the operation box, it is determined that the crane pipe has been connected to the injection port of the tank truck;

[0011] S5: Run the posture estimation model to extract key points of the person's body, calculate the angle between adjacent key point vectors, and identify related actions such as holding a watering can and spraying soapy water;

[0012] S6: Based on the continuous detection results of the relative position of the personnel and the crane pipe and the personnel's related actions, confirm that the personnel conducted continuous observation after spraying soapy water and confirmed that the leak detection operation was completed correctly.

[0013] S7: If the three states and action processes of crane pipe connection, soapy water spraying, and continuous observation are all completed in sequence, it is determined that the personnel have correctly performed the standardized leak detection action.

[0014] The present invention provides a visually intelligent method for confirming leak detection during the filling process of liquid tank trucks, enabling real-time monitoring of personnel performing leak detection operations on the tanks, such as spraying soapy water from a watering can. By leveraging deep learning and posture estimation, the method captures personnel and related facilities and equipment, detects the absence of necessary actions during the leak detection process, verifies the process's compliance, establishes a safety early warning mechanism, effectively eliminates potential safety hazards, and provides more intelligent and safe monitoring of the petrochemical product filling process in tank trucks. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] Figure 1 It is a flow chart of the method execution in the present invention.

[0016] Figure 2 Schematic diagram of detection scene examples and targets of the present invention. DETAILED DESCRIPTION

[0017] The technical solution of the present invention will be described in detail below with reference to the accompanying drawings:

[0018] Figure 1 This is a flow chart of the method of the present invention. Figure 1 As shown, the tank leak detection operation posture detection method described in the present invention includes the following steps:

[0019] S1: Under various lighting conditions, 360° video of on-site tanker safety inspections, filling and leak detection operations, and environmental images are captured from four angles. The images include the vehicle, operating box, crane, spray bottle, and on-site personnel, along with the surroundings, at a 7:1 ratio.

[0020] S2: Divide the dataset into a ratio of 8:2:1 to obtain the training set, validation set, and test set for machine learning. Then, perform target annotation to generate labels and use generative adversarial network technology to expand the dataset to cover small sample targets and achieve generalization effect. By introducing a dynamic partitioning strategy based on data distribution, the training set, validation set, and test set are made more representative, ensuring that each set contains sufficient sample diversity.

[0021] S3: Uses "transfer learning" and an adaptive learning rate algorithm to train the target detection model, enabling target detection and positioning of vehicles, operating boxes, crane hoses, spray bottles, and on-site personnel.

[0022] S4: According to the relative position and angle between the crane pipe and the operation box, it is determined that the crane pipe has been connected to the injection port of the tank truck;

[0023] S5: Run the posture estimation model to extract key points of the person's body, calculate the angle between adjacent key point vectors, and identify related actions such as holding a watering can and spraying soapy water;

[0024] S6: Based on the continuous detection results of the relative position of the personnel and the crane pipe and the personnel's related actions, confirm that the personnel have conducted continuous observation after spraying soapy water and confirm that the leak detection operation has been completed correctly;

[0025] S7: If the three states and action processes of crane pipe connection, soapy water spraying, and continuous observation are all completed in sequence, it is determined that the personnel have correctly performed the standardized leak detection action.

[0026] The acquisition of on-site monitoring video of leak detection operations and expansion of the data set described in S1 includes the following sub-steps:

[0027] S11. Under varying lighting conditions, capture 360° video of a tanker safety inspection at a chemical plant from four azimuths. Intercept scene images containing the vehicle, operating box, crane hose, spray can, and on-site personnel, as well as general scene images, from the video stream at a 7:1 ratio. Collect 3,500 images containing specific targets and 500 general scene images, and crop each image into several sub-images with a long side of 1,000 pixels and a short side of 600 pixels.

[0028] S12. Using incremental learning methods, the extracted initial image is geometrically transformed and combined with generative adversarial networks (GANs) to expand the dataset.

[0029] The specific method for geometric transformation of the image described in S12:

[0030] Horizontal Flip: Flip the image on the horizontal axis to generate samples with different perspectives.

[0031] Translation: Moves the image along the x-axis or y-axis in the image plane, simulating capturing the image at a different location.

[0032] Random scaling: randomly enlarge or reduce the image to increase the robustness of the model to images of different sizes;

[0033] The steps for expanding the dataset using a generative adversarial network described in S12 are as follows:

[0034] S121. Use the initial image dataset to train a GANs model. GANs consists of two neural networks: a generator and a discriminator. The generator's task is to generate new images, while the discriminator's task is to distinguish between generated images and real images.

[0035] S122. Use the generator to generate new image samples that are visually similar to the original dataset but have different details and variations;

[0036] S123, merging the new images generated by GANs with the original dataset (including images expanded by geometric transformation) to form a larger dataset;

[0037] As described in S2, the factory image dataset is divided into a ratio of 8:2:1 to obtain a training set, a validation set, and a test set, and then annotated and labeled, including the following sub-steps:

[0038] S21. Classify the dataset according to the monitoring location, ambient light, and the number of key targets included;

[0039] S22. Divide the classified data set into a ratio of 8:2:1 to obtain a training set, a validation set, and a test set.

[0040] S23. Use semi-supervised learning for labeling. First, manually label a part of the data, then use these labeled data to train an initial model, and then use this model to automatically label the remaining data, thereby ensuring the labeling quality to a certain extent.

[0041] The training target detection model described in S3 uses the "transfer learning" method to train the YOLOv8 target detection model as follows:

[0042] S31. Pose Estimation Model Network Structure Design. We use the "Real-Time Multi-person One-stage" RTMO model, with the CSPDarknet backbone network. This reduces the propagation of repeated gradient information to improve the network's learning capabilities and model inference speed.

[0043] The main components and structure of CSPDarknet:

[0044] Conv Layer: A standard convolutional layer that uses a convolution operation with a kernel size of 3x3, with BatchNormalization and Leaky ReLU as the activation function.

[0045] Residual Blocks: This retains the Residual Block from ResNet, which addresses the vanishing gradient problem in deep networks. Each residual block consists of two to three convolutional layers, passing input information through skip connections.

[0046] Cross-Stage Partial (CSP) Connection: The CSP structure splits the input feature map into two parts. One part passes directly through the residual block, while the other part skips the residual block and is later concatenated with the output feature map that passed through the residual block. Through this splitting and fusion, CSP can effectively reduce the propagation of repeated gradient information, reduce computational complexity, and improve the model's generalization ability.

[0047] Downsampling: A downsampling operation is performed at the end of each stage, using a convolution or pooling operation with a stride of 2 to halve the resolution of the feature map while increasing the number of channels to extract higher-level features.

[0048] Stage Structure: CSPDarknet consists of multiple stages, each of which contains several CSP modules. As the stage progresses, the number of channels in the feature map gradually increases and the resolution gradually decreases.

[0049] SPP Module (Spatial Pyramid Pooling Module): An SPP module is connected at the end of the backbone network to further increase the receptive field and capture multi-scale features.

[0050] The last three feature maps are processed using a hybrid encoder to produce two spatial features with downsampling rates of 16 and 32, respectively. Each pixel in these features is mapped to a grid cell evenly distributed on the original image plane. The network head uses a double convolution block at each spatial level to generate a score and corresponding pose feature for each grid cell. These pose features are used to predict bounding boxes, keypoint coordinates, and visibility, and a dynamic coordinate classifier is used to generate detailed information for 1-D heatmap prediction.

[0051] S32. Divide the dataset. Use the k-fold cross-validation method to divide the dataset marked in S2. That is, divide the dataset into k subsets. Each subset is used as a test set, and the rest is used as a training set. This fully utilizes the data for model training and evaluation, reducing the randomness of model performance.

[0052] S33. Model training and evaluation. Start with a learning rate of 0.001 and adjust it based on the change in loss during training. The batch size determines the number of samples used in each update of the model parameters. Smaller batch sizes may lead to more unstable training but help escape local minima. Larger batch sizes may make training more stable, but may sacrifice some generalization ability. For a dataset of 6,000 images, try batch sizes of 32, 64, or 128 and observe their effects on training speed and stability.

[0053] The steps for estimating the person's posture described in S4 are as follows:

[0054] S41. Deploy the posture estimation model RTMO in the project and extract 17 key points of the human body in the video, including 0-nose, 1-left eye, 2-right eye, 3-left ear, 4-right ear, 5-left shoulder, 6-right shoulder, 7-left elbow, 8-right elbow, 9-left wrist, 10-right wrist, 11-left hip, 12-right hip, 13-left knee, 14-right knee, 15-left ankle and 16-right ankle.

[0055] S42. Based on the coordinates of the key points, the limb vector is obtained. The inner product formula of the vector is used to derive the limb angle using the inverse tangent function. Based on the different limb angles and their degrees, judgment conditions are set to estimate the human body movement. The human body movement is combined with the position of the person and the leak detection operation of the tester on the tank. The specific calculation method is as follows:

[0056] Bending: Take the center point c1 of the left and right shoulders, the center point c2 of the left and right hips, and the center point c3 of the left and right knees, and calculate the vector and vector The limb angle θ is obtained according to the vector inner product formula.

[0057] Inner product formula:

[0058] Derived θ:

[0059] in, and are the upper body vector and the lower body vector, and θ is the angle between the upper and lower body.

[0060] Judgment condition: θ<140°

[0061] Extend your arm: Take the hand key point q1, the arm center point q2, and the shoulder key point q3, and calculate the vector and vector The limb angle α is obtained according to the vector inner product formula.

[0062] Inner product formula:

[0063] Derived α:

[0064] in, and are the forearm vector and the upper arm vector respectively, and α is the angle between the forearm and the upper arm.

[0065] Judgment condition: α>90°

[0066] Spraying: The key points of the hand coincide with the target of the spray bottle. The person is located in the crane pipe connection area. Hold the spray bottle with the arm extended forward for 3 to 5 seconds to complete the spraying action. Stay and observe for a while to confirm.

Claims

1. A method for visually confirming leak detection during the filling process of a liquid tank truck, characterized in that: During the leak detection operation of the tank truck during the filling process, a fixed-position camera is used to collect monitoring video. For each frame of the image, image processing, target and posture detection, and scene status determination are performed based on deep learning methods. The specific steps include: S1: Capture on-site monitoring videos of container filling and leak detection operations under various lighting conditions, capturing scene images including the vehicle, operating box, crane pipe, spray bottle, and on-site personnel, as well as environmental images at a 7:1 ratio. S2: Divide the dataset into a ratio of 8:2:1 to obtain training, validation, and test sets for machine learning. Then, perform target annotation to generate labels and use generative adversarial network technology to expand the dataset to cover small sample targets. S3: Uses "transfer learning" and an adaptive learning rate algorithm to train the target detection model, enabling target detection and positioning of vehicles, operating boxes, crane hoses, spray bottles, and on-site personnel. S4: According to the relative position and angle between the crane pipe and the operation box, it is determined that the crane pipe has been connected to the injection port of the tank truck; S5: Run the posture estimation model to extract key points of the person's body, calculate the angle between adjacent key point vectors, and identify a series of related actions including but not limited to holding a watering can and spraying soapy water; S6: Based on the continuous detection results of the relative position of the personnel and the crane pipe and the personnel's related actions, confirm that the personnel have conducted continuous observation after spraying soapy water and confirm that the leak detection operation has been completed correctly; S7: If the three states and action processes of crane pipe connection, soapy water spraying, and continuous observation are all completed in sequence, it is determined that the personnel have correctly performed the standardized leak detection action.

2. A method for visually confirming leak detection during the filling process of a liquid tank truck according to claim 1, characterized in that: The video images of the production site area described in step S1 are collected in an environment with different lighting conditions, and the video of the on-site tank truck safety inspection operation is collected from 360° without blind spots from the monitoring in four directions. The images including the vehicle, operation box, crane pipe, spray bottle and on-site personnel as well as the general scene images are intercepted from the video stream at a ratio of 7:

1.

3. The method for visually confirming leak detection during the filling process of a liquid tank truck according to claim 1, characterized in that: In step S2, "expanding the dataset using generative adversarial network technology" uses an incremental learning method to perform geometric transformation on the extracted initial image and combine it with generative adversarial networks (GANs) to expand the dataset. The specific methods for geometric transformation of the image include but are not limited to: (1) Horizontal flip: flip the image on the horizontal axis to generate samples with different perspectives; (2) Translation: moving the image along the x-axis or y-axis on the image plane to simulate the capture of the image at different positions; (3) Random scaling: randomly enlarging or reducing the image to increase the robustness of the model to images of different sizes; The steps of expanding the dataset using a generative adversarial network are as follows: S121. Use the initial image dataset to train a GANs model. GANs consists of two neural networks: a generator and a discriminator. The generator's task is to generate new images, while the discriminator's task is to distinguish between generated images and real images. S122. Use the generator to generate new image samples that are visually similar to the original dataset but have different details and variations; S123. Merge the new images generated by GANs with the original dataset, including the images expanded by geometric transformation, to form a larger dataset.

4. The method for visually confirming leak detection during the filling process of a liquid tank truck according to claim 1, characterized in that: The "target label generation" described in step S2 is annotated using semi-supervised learning. First, a portion of the data is manually annotated, and then an initial model is trained using this annotated data. This model is then used to automatically annotate the remaining data to ensure the quality of the annotation.

5. The method for visually confirming leak detection during the filling process of a liquid tank truck according to claim 1, characterized in that: The training target detection model in step S3 is based on the YOLOv8 pre-training model, and uses an adaptive learning rate algorithm to dynamically adjust the learning rate according to the gradient and loss during the model training process. The loss function selects a combination of DFL Loss and CIoU Loss to improve the detection box positioning accuracy. The transfer learning method is used to speed up the training model, and an early stopping mechanism is set to avoid overfitting and improve the model generalization ability.

6. The method for visually confirming leak detection during the filling process of a liquid tank truck according to claim 1, characterized in that: The training of the target detection model in step S3 is carried out in the following steps: S31. Pose Estimation Model Network Structure Design: The "Real-Time Multi-person One-stage" (RTMO) model is used, and the backbone network uses CSPDarknet. This reduces the propagation of repeated gradient information to improve the network's learning ability and the model's inference speed. S32. Divide the dataset: Use the k-fold cross-validation method to divide the dataset marked in S2, that is, divide the dataset into k subsets, each subset is used as a test set, and the rest is used as a training set, making full use of the data for model training and evaluation, reducing the randomness of model performance; S33. Model training and evaluation: The learning rate is tried starting from 0.001 and adjusted according to the change in loss during training. The batch size determines the number of samples used each time the model parameters are updated. The impact of the batch size on the training speed and stability is observed during training.

7. The method for visually confirming leak detection during the filling process of a liquid tank truck according to claim 1, characterized in that: The determination of the crane pipe connection status described in step S4 is based on the target detection and positioning of the tank truck operating box, the target detection and positioning of the crane pipe equipment, and the calculation of the relative distance and overlap degree of the positions of the two targets in the picture, thereby determining the connection status of the crane pipe at the injection port of the operating box.

8. The method for visually confirming leak detection during the filling process of a liquid tank truck according to claim 1, characterized in that: The posture estimation described in step S5 uses a single-stage posture estimation framework (Real-Time Multi-person One-stage) RTMO. In the YOLO architecture, dual one-dimensional heat maps are used to represent key points to seamlessly integrate coordinate classification to extract the coordinates of key points of the human body. These coordinates are used to calculate the angles and relative positions of various parts of the human body to determine the posture action: when a person picks up a spray bottle, the spray bottle target overlaps with the hand key point; when a person sprays soapy water to detect leaks, the person is near the crane pipe connection area, holding the spray bottle toward the connection, and the upper arm and forearm are extended forward at an obtuse angle; The continuous observation described in step S6 is obtained by detecting the duration of time the person stands near the crane connection of the operating box holding a spray bottle, and calculating: the position of the person, the movement of holding the spray bottle, and the position of the operating box and the crane, which are obtained by the aforementioned target detection and posture detection methods.

9. A method for visually confirming leak detection during the filling process of a liquid tank truck according to claim 8, characterized in that: The specific steps for estimating the posture of a person described in step S5 are as follows: S41. Deploy the posture estimation model RTMO in the project and extract 17 key points of the human body in the video, including 0-nose, 1-left eye, 2-right eye, 3-left ear, 4-right ear, 5-left shoulder, 6-right shoulder, 7-left elbow, 8-right elbow, 9-left wrist, 10-right wrist, 11-left hip, 12-right hip, 13-left knee, 14-right knee, 15-left ankle and 16-right ankle; S42. Based on the coordinates of the key points, the limb vector is obtained. The limb angle is calculated using the inverse tangent function by deriving the inner product formula of the vector. Based on different limb angles and their degrees, judgment conditions are set to estimate human body movements. The human body movements are combined with the position of the person and the leak detection operation of the tester on the tank body. The specific calculation method is as follows: Bending: Take the center point c1 of the left and right shoulders, the center point c2 of the left and right hips, and the center point c3 of the left and right knees, and calculate the vector and vector The limb angle θ is obtained according to the vector inner product formula: Inner product formula: Derived θ: in, and are the upper body vector and the lower body vector, respectively, and θ is the angle between the upper and lower body; Judgment condition: θ<140° Extend your arm: Take the hand key point q1, the arm center point q2, and the shoulder key point q3, and calculate the vector and vector The limb angle α is obtained according to the vector inner product formula: Inner product formula: Derived α: in, and are the forearm vector and the upper arm vector, respectively, and α is the angle between the forearm and the upper arm; Judgment condition: α>90° Spraying: The key points of the hand coincide with the target of the spray bottle. The person is located in the crane pipe connection area. Hold the spray bottle with the arm extended forward for 3 to 5 seconds to complete the spraying action. Stay and observe for a while to confirm.

Citation Information

Patent Citations

  • Security detection system and security detection method

    CN117876949A

  • Intelligent visual monitoring and early warning method for safety of tank entering operation personnel of tank car

    CN118587648A