Lifting safety detection method and system based on artificial intelligence

By using multimodal structured target extraction, multi-class target trajectory modeling, and dual risk factor assessment, the problems of limited field of view and insufficient dynamic risk identification in lifting safety inspection are solved, and efficient safety risk assessment and prevention are achieved.

CN120876833AInactive Publication Date: 2025-10-31SHANDONG ZHINUO SAFETY TECHNOLOGY CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202511006300.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-22
Publication Date
2025-10-31
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing crane safety inspection methods suffer from limited monitoring field of view, inaccurate target identification, difficulty in capturing target spatial location in dynamic construction environments, and traditional risk assessment methods rely on static spatial location judgment, resulting in insufficient dynamic risk perception and difficulty in identifying approach conflict behavior between targets and unauthorized entry into high-risk work areas.

Method used

A multimodal structured target extraction method is adopted, which combines RGB images and depth information to extract targets and obtain their 3D positions; a multi-class target trajectory modeling method is used for spatiotemporal trajectory prediction; and a risk assessment method that integrates two risk factors is used to assess conflict and violation risks.

Benefits of technology

It improves the ability to perceive dynamic risks in lifting operation scenarios, accurately predicts the trend of target position changes, and enhances the level of accident prevention.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120876833A_ABST
    Figure CN120876833A_ABST
Patent Text Reader

Abstract

The invention discloses a hoisting safety detection method and system based on artificial intelligence, and belongs to the technical field of hoisting safety management, and the method comprises the steps of panoramic view perception, target extraction, spatial-temporal trajectory modeling and safety risk assessment. According to the invention, a multi-modal structured target extraction method is adopted for target extraction, RGB images and depth information are integrated, the defect that the recognition rate of a single camera is reduced under shielding and illumination changes is overcome, the three-dimensional position and the motion state of the target can be obtained, and the interaction relationship between the target and a restricted area can be judged; according to the method, spatial-temporal trajectory modeling is carried out by adopting a multi-class target trajectory modeling method, the position change trend of various types of targets is accurately predicted, and the future possible moving region range of the targets is depicted through trajectory distribution modeling, so that prospective support is provided for risk identification; safety risk assessment is carried out by adopting a risk scoring method fusing double risk factors, the comprehensive safety risk is comprehensively assessed, and the accident prevention level in the lifting operation environment is remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of crane safety management technology, specifically referring to a crane safety detection method and system based on artificial intelligence. Background Technology

[0002] AI-based crane safety detection integrates AI technologies such as image recognition, trajectory modeling, and risk analysis to perform real-time perception and risk assessment of key targets such as loads, personnel, hooks, and vehicles at crane operation sites. It aims to provide proactive and intelligent safety assurance for crane operations and help achieve safety management in high-risk scenarios.

[0003] However, existing crane safety inspection processes suffer from several technical problems: limited monitoring field of view, inaccurate target identification, and difficulty in capturing the spatial position of targets in dynamic construction environments; significant differences in the movement patterns of different types of targets make it difficult to accurately predict their future movement trends, leading to lagging dynamic risk identification and a high misjudgment rate; and traditional safety risk assessment methods largely rely on static spatial position judgments and have a single risk assessment dimension, resulting in insufficient perception of dynamic risks at the work site, difficulty in effectively identifying approach conflict behaviors between targets, and targets illegally entering high-risk work areas, leading to weak risk prediction capabilities at crane operation sites. Summary of the Invention

[0004] To address the aforementioned issues and overcome the shortcomings of existing technologies, this invention provides an artificial intelligence-based method and system for crane safety detection. Addressing the technical problems of limited monitoring field of view, inaccurate target identification, and difficulty in capturing target spatial positions in dynamic construction environments during existing crane safety detection processes, this solution creatively employs a multimodal structured target extraction method, integrating RGB images and depth information. This not only overcomes the shortcomings of single cameras in terms of reduced recognition rate under occlusion and lighting changes, but also acquires the target's three-dimensional position and motion state, helping to determine the interaction between the target and restricted areas and promptly identify potential safety hazards. Furthermore, addressing the technical problem in existing crane safety detection processes where different types of targets exhibit significantly different motion patterns, making it difficult to accurately predict their future movement trends and leading to delayed dynamic risk identification and high misjudgment rates, this solution creatively employs a multi-type target trajectory modeling method for spatiotemporal trajectory modeling. To address the movement characteristics of different target types, a trajectory prediction model is matched to accurately predict the positional change trends of various targets. Furthermore, trajectory distribution modeling characterizes the potential future movement area of ​​the targets, providing forward-looking support for risk identification and thus enhancing the dynamic risk perception capability in lifting operation scenarios. In existing lifting safety inspection processes, traditional safety risk assessment methods largely rely on static spatial location judgments and have a single risk assessment dimension, resulting in insufficient perception of dynamic risks at the work site, difficulty in effectively identifying approach conflict behaviors between targets, and targets illegally entering high-risk work areas, leading to weak risk prediction capabilities at lifting operation sites. This solution creatively adopts a risk scoring method that integrates two risk factors for safety risk assessment. Simultaneously, it comprehensively assesses the overall safety risks at the lifting site from two key dimensions: conflict risk factors and violation risk factors, thereby significantly improving the accident prevention level in lifting operation environments.

[0005] The technical solution adopted by this invention is as follows: The lifting safety detection method based on artificial intelligence provided by this invention includes the following steps:

[0006] Step S1: Panoramic field of view perception;

[0007] Step S2: Target extraction;

[0008] Step S3: Spatiotemporal trajectory modeling;

[0009] Step S4: Security risk assessment.

[0010] Further, in step S1, the panoramic field of view perception specifically involves deploying a multi-view RGB camera and a depth camera at the middle position of the crane boom to simultaneously acquire a dual-modal multi-view image stream and obtain the intrinsic parameters of the depth camera. Then, image frames are extracted from the dual-modal multi-view image stream, and frame synchronization alignment is performed according to the timestamp. Finally, image preprocessing is performed to obtain RGB panoramic image frames and depth panoramic image frames.

[0011] Further, in step S2, the target extraction is used to extract key objects in the lifting operation scene. Specifically, based on the intrinsic parameters of the depth camera, RGB panoramic image frames, and depth panoramic image frames, a multimodal structured target extraction method is used to extract targets and obtain a set of target structured features, including the following steps:

[0012] Step S21: Image detection, specifically, by using the target detection network as the backbone network and introducing the feature pyramid network and path aggregation network for feature fusion, setting a dual-branch structure, constructing an image target detection model, and then using the image target detection model to perform image detection on RGB panoramic image frames to obtain image detection results. The image detection results include target bounding boxes, target type labels, restricted region bounding boxes, and restricted region type labels.

[0013] The dual-branch head structure is specifically designed by using the multi-scale detection layer of the target detection network as the detection head and the dilated convolutional block and decoding block of the semantic segmentation network as the segmentation head.

[0014] The target type labels include categories such as suspended objects, hooks, personnel, and vehicles;

[0015] The restricted area type labels include prohibited areas, robotic arm activity areas, and high-altitude suspended object areas;

[0016] Step S22: Spatial position estimation, specifically, the depth panoramic image frame and the RGB panoramic image frame are aligned at the pixel level, the pixel depth values ​​in the target area are extracted, and the target center is converted into three-dimensional coordinates by depth projection using the intrinsic parameters of the depth camera to obtain the target center position; the three-dimensional size of the target is obtained by performing boundary statistics on the depth distribution of the target area.

[0017] Step S23: Motion state modeling, specifically by calculating the difference in the target center position at adjacent time points and combining it with the time interval, the target velocity vector is obtained, and the pixel-level motion pattern of the target region is modeled by a deep learning optical flow network to obtain the target displacement pattern features. Then, the target velocity vector and the target displacement pattern features are concatenated to obtain the target motion state features.

[0018] Step S24: Target library generation, specifically, involves concatenating the target type label, target center position, target 3D size, and target motion state features of each target to obtain target structured features, and summarizing the structured features of each target to generate a set of structured target features.

[0019] Further, in step S3, the spatiotemporal trajectory modeling, used to predict the target's motion trajectory, specifically involves using a multi-class target trajectory modeling method based on a structured target feature set to perform spatiotemporal trajectory modeling, obtaining the target center position trajectory prediction distribution, including the following steps:

[0020] Step S31: Construction of the state trajectory sequence, specifically, extracting the structured features of each target from continuous historical time points and splicing them in chronological order to form the target state trajectory sequence;

[0021] Step S32: Classification prediction modeling, used to predict the future center position trajectory changes of targets for different target types. Specifically, based on the target type label, a trajectory prediction model is set, and the target state trajectory sequence is classified and predicted to obtain the target center position predicted trajectory sequence.

[0022] Specifically, when the target type label is a suspended object or a hook, a trajectory prediction model based on a standard transformer architecture is used; when the target type label is a person or a vehicle, a trajectory prediction model based on a social long short-term memory network is used.

[0023] Step S33: Position trajectory distribution generation. Specifically, first estimate the mean and covariance matrix of the predicted trajectory sequence of each target center position, and then generate the target center position trajectory prediction distribution by performing multidimensional Gaussian distribution modeling on the target center position predicted trajectory sequence.

[0024] Further, in step S4, the safety risk assessment specifically involves using a risk scoring method that integrates two risk factors to assess the safety risk based on the structured target feature set and the predicted distribution of the target center location trajectory, to obtain a comprehensive target risk value, including the following steps:

[0025] Step S41: Trajectory conflict assessment, specifically, calculates the probability that the distance between any two target center locations is less than the safe distance threshold at a future time point, and takes the maximum value as the trajectory conflict risk value;

[0026] Step S42: Spatial violation detection, specifically, calculating the probability of each target's center location entering the restricted area at a future time point, and taking the maximum value as the target violation risk value;

[0027] Step S43: Comprehensive risk modeling, specifically, first obtain the maximum trajectory conflict risk value between each target and other targets, then combine the target violation risk value with weighted fusion, integrate the two risk factors, and generate the target comprehensive risk value.

[0028] The lifting safety detection system based on artificial intelligence provided by the present invention includes: a panoramic field of view perception module, a target extraction module, a spatiotemporal trajectory modeling module, and a safety risk assessment module;

[0029] The panoramic field of view perception module is used for panoramic field of view perception. Through panoramic field of view perception, it obtains RGB panoramic image frames, depth panoramic image frames and depth camera intrinsic parameters, and sends the RGB panoramic image frames, the depth panoramic image frames and the depth camera intrinsic parameters to the target extraction module.

[0030] The target extraction module is used for target extraction. Through target extraction, a set of structured features of the target is obtained, and the set of structured features of the target is sent to the spatiotemporal trajectory modeling module and the safety risk assessment module.

[0031] The spatiotemporal trajectory modeling module is used for spatiotemporal trajectory modeling. Through spatiotemporal trajectory modeling, the predicted distribution of the target center position trajectory is obtained, and the predicted distribution of the target center position trajectory is sent to the safety risk assessment module.

[0032] The security risk assessment module is used for security risk assessment, and through the security risk assessment, the target comprehensive risk value is obtained.

[0033] The beneficial effects achieved by the present invention using the above solution are as follows:

[0034] (1) In view of the technical problems of limited monitoring field of view, inaccurate target recognition and difficulty in capturing the spatial position of the target in the existing crane safety inspection process, this solution creatively adopts a multimodal structured target extraction method to extract the target, integrating RGB image and depth information. This not only overcomes the shortcomings of the single camera in terms of the decrease in recognition rate under occlusion and light change, but also obtains the three-dimensional position and motion state of the target, which helps to judge the interaction relationship between the target and the restricted area and to discover potential safety hazards in a timely manner.

[0035] (2) In response to the technical problem that there are significant differences in the movement patterns of different types of targets in the existing crane safety inspection process, making it difficult to accurately predict their future movement trends, resulting in lagging dynamic risk identification and high misjudgment rate, this solution creatively adopts a multi-type target trajectory modeling method to perform spatiotemporal trajectory modeling. For the movement characteristics of different types of targets, a trajectory prediction model is matched to accurately predict the position change trend of various targets. Through trajectory distribution modeling, the range of possible future movement areas of targets is depicted, providing forward-looking support for risk identification, thereby improving the dynamic risk perception capability in crane operation scenarios.

[0036] (3) In view of the technical problems in the existing crane safety inspection process, the traditional safety risk assessment methods mostly rely on static spatial location judgment and have a single risk assessment dimension, thus lacking the perception of dynamic risks at the work site, making it difficult to effectively identify the approach conflict behavior between targets and the illegal entry of targets into high-risk work areas, resulting in weak risk prediction ability at the crane operation site, this solution creatively adopts a risk scoring method that integrates dual risk factors for safety risk assessment. At the same time, it comprehensively assesses the overall safety risks at the crane site from two key dimensions: conflict risk factors and violation risk factors, thereby significantly improving the level of accident prevention in the crane operation environment. Attached Figure Description

[0037] Figure 1 A flowchart illustrating the artificial intelligence-based lifting safety detection method provided by this invention;

[0038] Figure 2 A schematic diagram of the artificial intelligence-based lifting safety detection system provided by the present invention;

[0039] Figure 3 This is a flowchart illustrating step S2;

[0040] Figure 4 This is a flowchart illustrating step S3;

[0041] Figure 5 This is a flowchart illustrating step S4.

[0042] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are used together with the embodiments of the invention to explain the invention and do not constitute a limitation thereof. Detailed Implementation

[0043] The technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative effort are within the scope of protection of the present invention.

[0044] In the description of this invention, it should be understood that the terms "upper", "lower", "front", "rear", "left", "right", "top", "bottom", "inner", "outer", etc., indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings. They are only for the convenience of describing this invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, they should not be construed as limitations on this invention.

[0045] Example 1, see Figure 1 The present invention provides an artificial intelligence-based lifting safety detection method, which includes the following steps:

[0046] Step S1: Panoramic field of view perception;

[0047] Step S2: Target extraction;

[0048] Step S3: Spatiotemporal trajectory modeling;

[0049] Step S4: Security risk assessment.

[0050] Example 2, see Figure 1 This embodiment is based on the above embodiment. In step S1, the panoramic field of view perception specifically involves deploying a multi-view RGB camera and a depth camera at the middle position of the crane boom to simultaneously acquire a dual-modal multi-view image stream and obtain the intrinsic parameters of the depth camera. Then, image frames are extracted from the dual-modal multi-view image stream, and frame synchronization alignment is performed according to the timestamp. Finally, image preprocessing is performed to obtain RGB panoramic image frames and depth panoramic image frames.

[0051] The dual-modal multi-view image stream includes an RGB multi-view image stream and a depth multi-view image stream;

[0052] The image preprocessing includes multi-view image stitching, distortion correction, and illumination equalization.

[0053] Example 3, see Figure 1 and Figure 3This embodiment is based on the above embodiment. In step S2, the target extraction is used to extract key objects in the lifting operation scene. Specifically, it is based on the intrinsic parameters of the depth camera, RGB panoramic image frames, and depth panoramic image frames, and uses a multimodal structured target extraction method to extract targets, obtaining a target structured feature set, including the following steps:

[0054] Step S21: Image detection, specifically, by using the target detection network as the backbone network and introducing the feature pyramid network and path aggregation network for feature fusion, setting a dual-branch structure, constructing an image target detection model, and then using the image target detection model to perform image detection on RGB panoramic image frames to obtain image detection results. The image detection results include target bounding boxes, target type labels, restricted region bounding boxes, and restricted region type labels.

[0055] The dual-branch head structure is specifically designed by using the multi-scale detection layer of the target detection network as the detection head and the dilated convolutional block and decoding block of the semantic segmentation network as the segmentation head.

[0056] The target type labels include categories such as suspended objects, hooks, personnel, and vehicles;

[0057] The restricted area type labels include prohibited areas, robotic arm activity areas, and high-altitude suspended object areas;

[0058] Preferably, the target detection network is a YOLOv8 network, and the semantic segmentation network is a DeepLabv3+ network;

[0059] Step S22: Spatial position estimation, specifically, the depth panoramic image frame and the RGB panoramic image frame are aligned at the pixel level, the pixel depth values ​​in the target area are extracted, and the target center is converted into three-dimensional coordinates by depth projection using the intrinsic parameters of the depth camera to obtain the target center position; the three-dimensional size of the target is obtained by performing boundary statistics on the depth distribution of the target area.

[0060] Step S23: Motion state modeling, specifically by calculating the difference in the target center position at adjacent time points and combining it with the time interval, the target velocity vector is obtained, and the pixel-level motion pattern of the target region is modeled by a deep learning optical flow network to obtain the target displacement pattern features. Then, the target velocity vector and the target displacement pattern features are concatenated to obtain the target motion state features.

[0061] The formula for calculating the target velocity vector is:

[0062] ;

[0063] In the formula, It is the velocity vector of the i-th target at time t, where i is the first index of the target and t is the index of the time point. It is the center position of the i-th target at time t. It is the center position of the i-th target at time t-1. It is a time interval;

[0064] The calculation formula for the target displacement mode characteristics is as follows:

[0065] ;

[0066] In the formula, F is the target displacement pattern feature at time t. flow (·) is a function of a deep learning optical flow network. It is the RGB panoramic image frame at time point t-1. It is the RGB panoramic image frame at time t. It is the bounding box of the i-th target at time t;

[0067] Preferably, the deep learning optical flow network is a PWC optical flow network;

[0068] Step S24: Target library generation, specifically, involves concatenating the target type label, target center position, target 3D size, and target motion state features of each target to obtain target structured features, and summarizing the structured features of each target to generate a set of structured target features;

[0069] By performing the above operations, this solution addresses the technical problems in existing crane safety inspection processes, such as limited monitoring field of view, inaccurate target identification, and difficulty in capturing the spatial position of targets in dynamic construction environments. It creatively adopts a multimodal structured target extraction method to extract targets, integrating RGB images and depth information. This not only overcomes the shortcomings of a single camera in terms of reduced recognition rate under occlusion and changes in lighting, but also obtains the three-dimensional position and motion state of the target, which helps to determine the interaction relationship between the target and the restricted area and to promptly detect potential safety hazards.

[0070] Example 4, see Figure 1 and Figure 4 This embodiment is based on the above embodiment. In step S3, the spatiotemporal trajectory modeling is used to predict the target's motion trajectory. Specifically, it involves using a multi-class target trajectory modeling method based on a structured target feature set to perform spatiotemporal trajectory modeling, thereby obtaining the target center position trajectory prediction distribution. This includes the following steps:

[0071] Step S31: State trajectory sequence construction, specifically, extracting the structured features of each target from continuous historical time points and concatenating them in chronological order to form the target state trajectory sequence. The calculation formula is as follows:

[0072] ;

[0073] In the formula, It is the trajectory sequence of the i-th target state at time t. It is the i-th target structured feature at time point tK, where K is the number of backtracking time points. It is the i-th target structured feature at time t;

[0074] Step S32: Classification prediction modeling, used to predict the future center position trajectory changes of targets for different target types. Specifically, based on the target type label, a trajectory prediction model is set, and the target state trajectory sequence is classified and predicted to obtain the target center position predicted trajectory sequence.

[0075] The trajectory prediction model is specifically designed as follows: when the target type label is "suspended object" or "hook," a trajectory prediction model based on a standard transformer architecture is used; when the target type label is "personnel" or "vehicle," a trajectory prediction model based on a social long short-term memory network is used. The calculation formula is as follows:

[0076] ;

[0077] In the formula, It is the predicted trajectory sequence of the i-th target center location within a future time point. It predicts the time span, F trans (·) is a trajectory prediction model function based on a standard transformer architecture, where obj is the object class, hook is the hook class, and F soc (·) is a trajectory prediction model function based on social long short-term memory network, where per is the people class and car is the vehicle class;

[0078] Step S33: Generating the location trajectory distribution. Specifically, first, the mean and covariance matrix of the predicted trajectory sequence for each target center location are estimated. Then, a multidimensional Gaussian distribution model is performed on the predicted trajectory sequence for the target center location to generate the predicted trajectory distribution for the target center location. The calculation formula is as follows:

[0079] ;

[0080] In the formula, It is the predicted distribution of the trajectory of the i-th target center location at a future time point. It is a multidimensional Gaussian distribution operator. It is the mean of the predicted trajectory sequence of the i-th target center location within a future time point. It is the covariance matrix of the predicted trajectory sequence of the i-th target center position at a future time point;

[0081] By performing the above operations, this solution addresses the technical problem in existing crane safety inspection processes where the movement patterns of different types of targets differ significantly, making it difficult to accurately predict their future movement trends, resulting in lagging dynamic risk identification and a high misjudgment rate. It creatively employs a multi-target trajectory modeling method for spatiotemporal trajectory modeling. For the movement characteristics of different types of targets, a trajectory prediction model is matched to accurately predict the positional change trends of various targets. Furthermore, through trajectory distribution modeling, the potential future movement area of ​​the target is depicted, providing forward-looking support for risk identification and thereby enhancing the dynamic risk perception capability in crane operation scenarios.

[0082] Example 5, see Figure 1 and Figure 5 This embodiment is based on the above embodiment. In step S4, the security risk assessment specifically involves using a risk scoring method that integrates two risk factors to conduct a security risk assessment based on the structured target feature set and the target center location trajectory prediction distribution, to obtain the target comprehensive risk value. This includes the following steps:

[0083] Step S41: Trajectory conflict assessment, specifically calculating the probability that the distance between the center positions of any two targets is less than the safe distance threshold at a future time point, and taking the maximum value as the trajectory conflict risk value. The calculation formula is as follows:

[0084] ;

[0085] In the formula, This is the trajectory conflict risk value between the i-th and j-th targets, used to represent the conflict risk factor, where j is the second target index, which is not equal to the first target index. It is a future time point index, max is the maximum value function, and p(·) is the probability calculation function. It is in the The center position of the i-th target at a future time point. It is in the The j-th target center position at a future time point, ||·|| is the L2 norm operator used to evaluate the distance between any two target center positions, d safe It is the safe distance threshold;

[0086] Step S42: Spatial violation detection, specifically calculating the probability that the center position of each target will enter the restricted area at a future time point, and taking the maximum value as the target violation risk value. The calculation formula is as follows:

[0087] ;

[0088] In the formula, R is the i-th target violation risk value, used to represent the violation risk factor. m It is a set of restricted regions;

[0089] Step S43: Comprehensive risk modeling. Specifically, first, obtain the maximum trajectory conflict risk value between each target and other targets. Then, combine the target violation risk value with the target risk value for weighted fusion, integrate the two risk factors, and generate the target comprehensive risk value. The calculation formula is as follows:

[0090] ;

[0091] In the formula, It is the comprehensive risk value of the i-th target at time t. It is the conflict risk factor weight. It is the weight of the violation risk factor;

[0092] By performing the above operations, this solution addresses the technical problems in existing crane safety inspection processes. Traditional safety risk assessment methods mostly rely on static spatial location judgments and have a single risk assessment dimension, resulting in insufficient perception of dynamic risks at the work site, difficulty in effectively identifying approach conflict behaviors between targets, and targets illegally entering high-risk work areas, leading to weak risk prediction capabilities at crane operation sites. This solution creatively adopts a risk scoring method that integrates two risk factors for safety risk assessment. It comprehensively assesses the overall safety risks at the crane site from two key dimensions: conflict risk factors and violation risk factors, thereby significantly improving the level of accident prevention in crane operation environments.

[0093] Example 6, see Figure 2 Based on the above embodiments, the lifting safety detection system based on artificial intelligence provided by the present invention includes: a panoramic field of vision perception module, a target extraction module, a spatiotemporal trajectory modeling module, and a safety risk assessment module;

[0094] The panoramic field of view perception module is used for panoramic field of view perception. Through panoramic field of view perception, it obtains RGB panoramic image frames, depth panoramic image frames and depth camera intrinsic parameters, and sends the RGB panoramic image frames, the depth panoramic image frames and the depth camera intrinsic parameters to the target extraction module.

[0095] The target extraction module is used for target extraction. Through target extraction, a set of structured features of the target is obtained, and the set of structured features of the target is sent to the spatiotemporal trajectory modeling module and the safety risk assessment module.

[0096] The spatiotemporal trajectory modeling module is used for spatiotemporal trajectory modeling. Through spatiotemporal trajectory modeling, the predicted distribution of the target center position trajectory is obtained, and the predicted distribution of the target center position trajectory is sent to the safety risk assessment module.

[0097] The security risk assessment module is used for security risk assessment, and through the security risk assessment, the target comprehensive risk value is obtained.

[0098] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus.

[0099] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention.

[0100] The present invention and its embodiments have been described above. This description is not restrictive, and the accompanying drawings are only one embodiment of the present invention; the actual structure is not limited thereto. In conclusion, if those skilled in the art are inspired by this description and design similar structures and embodiments without departing from the spirit of the invention, such designs should fall within the protection scope of the present invention.

Claims

1. A lifting safety detection method based on artificial intelligence, characterized in that: The method includes the following steps: Step S1: Panoramic field of view perception, obtaining depth camera intrinsic parameters, RGB panoramic image frames and depth panoramic image frames; Step S2: Target extraction, specifically based on the intrinsic parameters of the depth camera, RGB panoramic image frames, and depth panoramic image frames, a multimodal structured target extraction method is used to extract targets and obtain a set of structured target features. This includes the following steps: Step S21: Image detection; Step S22: Spatial position estimation; Step S23: Motion state modeling; Step S24: Target library generation. Step S3: Spatiotemporal trajectory modeling, specifically based on the structured target feature set, uses a multi-class target trajectory modeling method to perform spatiotemporal trajectory modeling, obtaining the target center position trajectory prediction distribution, including the following steps: Step S31: State trajectory sequence construction; Step S32: Classification prediction modeling; Step S33: Position trajectory distribution generation; Step S4: Safety risk assessment, specifically based on the structured target feature set and the target center location trajectory prediction distribution, adopts a risk scoring method that integrates two risk factors to conduct a safety risk assessment and obtain the target comprehensive risk value, including the following steps: Step S41: Trajectory conflict assessment; Step S42: Spatial violation detection; Step S43: Comprehensive risk modeling.

2. The lifting safety detection method based on artificial intelligence according to claim 1, characterized in that: In step S21, the image detection specifically involves using a target detection network as the backbone network and introducing a feature pyramid network and a path aggregation network for feature fusion, setting a dual-branch structure, constructing an image target detection model, and then using the image target detection model to perform image detection on RGB panoramic image frames to obtain image detection results. The image detection results include target bounding boxes, target type labels, restricted region bounding boxes, and restricted region type labels. The dual-branch head structure is specifically designed by using the multi-scale detection layer of the target detection network as the detection head and the dilated convolutional block and decoding block of the semantic segmentation network as the segmentation head. The target type labels include categories such as suspended objects, hooks, personnel, and vehicles; The restricted area type labels include prohibited areas, robotic arm activity areas, and high-altitude suspended object areas; In step S22, the spatial position estimation specifically involves pixel-level alignment of the depth panoramic image frame and the RGB panoramic image frame, extracting the pixel depth values ​​within the target area, and using the intrinsic parameters of the depth camera to convert the target center into three-dimensional coordinates through depth projection to obtain the target center position; and obtaining the target three-dimensional size by performing boundary statistics on the depth distribution of the target area. In step S23, the motion state modeling specifically involves calculating the difference in the target center position at adjacent time points and combining it with the time interval to obtain the target velocity vector, and then using a deep learning optical flow network to model the pixel-level motion pattern of the target region to obtain the target displacement pattern features. Finally, the target velocity vector and the target displacement pattern features are concatenated to obtain the target motion state features. In step S24, the target library is generated by concatenating the target type label, target center position, target three-dimensional size, and target motion state features of each target to obtain target structured features, and then summarizing each target structured feature to generate a set of structured target features.

3. The lifting safety detection method based on artificial intelligence according to claim 2, characterized in that: In step S31, the construction of the state trajectory sequence specifically involves extracting the structured features of each target from continuous historical time points and splicing them together in chronological order to form a target state trajectory sequence. In step S32, the classification prediction modeling is used to predict the future center position trajectory changes of the target for different target types. Specifically, based on the target type label, a trajectory prediction model is set, and the target state trajectory sequence is classified and predicted to obtain the target center position predicted trajectory sequence. Specifically, when the target type label is a suspended object or a hook, a trajectory prediction model based on a standard transformer architecture is used; when the target type label is a person or a vehicle, a trajectory prediction model based on a social long short-term memory network is used. In step S33, the location trajectory distribution generation specifically involves first estimating the mean and covariance matrix of the predicted trajectory sequence for each target center location, and then generating the target center location trajectory prediction distribution by performing multidimensional Gaussian distribution modeling on the predicted trajectory sequence for the target center location.

4. The lifting safety detection method based on artificial intelligence according to claim 3, characterized in that: In step S41, the trajectory conflict assessment specifically involves calculating the probability that the distance between any two target center locations is less than a safe distance threshold at a future time point, and taking the maximum value as the trajectory conflict risk value. In step S42, the spatial violation detection specifically involves calculating the probability that each target center location will enter the restricted area at a future time point, and taking the maximum value as the target violation risk value. In step S43, the comprehensive risk modeling specifically involves first obtaining the maximum trajectory conflict risk value between each target and other targets, and then combining the target violation risk value with a weighted fusion to integrate the two risk factors and generate a comprehensive target risk value.

5. The lifting safety detection method based on artificial intelligence according to claim 4, characterized in that: In step S1, the panoramic field of view perception specifically involves deploying a multi-view RGB camera and a depth camera at the middle position of the crane boom to simultaneously acquire a dual-modal multi-view image stream and obtain the intrinsic parameters of the depth camera. Then, image frames are extracted from the dual-modal multi-view image stream, and frame synchronization alignment is performed according to the timestamp. Finally, image preprocessing is performed to obtain RGB panoramic image frames and depth panoramic image frames.

6. An artificial intelligence-based lifting safety detection system, used to implement the artificial intelligence-based lifting safety detection method as described in any one of claims 1-5, characterized in that: It includes a panoramic field of view perception module, a target extraction module, a spatiotemporal trajectory modeling module, and a safety risk assessment module.

7. The lifting safety detection system based on artificial intelligence according to claim 6, characterized in that: The panoramic field of view perception module is used for panoramic field of view perception. Through panoramic field of view perception, it obtains RGB panoramic image frames, depth panoramic image frames and depth camera intrinsic parameters, and sends the RGB panoramic image frames, the depth panoramic image frames and the depth camera intrinsic parameters to the target extraction module. The target extraction module is used for target extraction. Through target extraction, a set of structured features of the target is obtained, and the set of structured features of the target is sent to the spatiotemporal trajectory modeling module and the safety risk assessment module. The spatiotemporal trajectory modeling module is used for spatiotemporal trajectory modeling. Through spatiotemporal trajectory modeling, the predicted distribution of the target center position trajectory is obtained, and the predicted distribution of the target center position trajectory is sent to the safety risk assessment module. The security risk assessment module is used for security risk assessment, and through the security risk assessment, the target comprehensive risk value is obtained.

Citation Information

Cited By

  • Monitoring and real-time alarm method and system for operation safety of mobile mechanical arm

    CN122265908A