A personnel safety monitoring method and system for hoisting in a forklift cooperative working environment

CN122821087APending Publication Date: 2026-09-25ANHUI ZHIZHI ENG TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610936700.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-26
Publication Date
2026-09-25

AI Technical Summary

Benefits of technology

1、本发明采用一阶段模型全局检测大目标,结合自适应膨胀裁剪与二阶段模型精细检测小目标,有效解决了远距离小目标易漏检的问题。同时,基于大目标检测框自适应生成非对称扩展监控区域,实现动态作业区域的人员侵入精准判断,配合连续多帧确认机制,极大降低了误报与漏报率。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122821087A_ABST
    Figure CN122821087A_ABST
Patent Text Reader

Abstract

The application discloses a personnel safety monitoring method for hoisting and forklift cooperative working environment, and relates to the technical field of deep learning image processing, and comprises the following steps: S1, acquiring real-time video stream of a working area and decoding into continuous frame images; S2, target detection is performed on the images, large targets and small targets are identified, and a first detection frame of the large targets and a second detection frame of the small targets are output; and S3, an adaptive expansion region is generated in the images according to the position and size of the first detection frame; the application adopts a one-stage model to globally detect large targets, combines adaptive inflation cropping with a two-stage model to finely detect small targets, and effectively solves the problem that small targets at a long distance are prone to be missed. Meanwhile, based on the adaptive generation of an asymmetric expansion monitoring region from the large target detection frame, precise judgment of personnel intrusion in a dynamic working area is realized, and in combination with a continuous multi-frame confirmation mechanism, the false alarm and missed alarm rates are greatly reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of deep learning-based image processing technology, and in particular to a personnel safety monitoring method and system for a hoisting and forklift cooperative working environment. Background Art

[0002] In a modern cement production process, the finished product packaging and warehousing links are key nodes connecting the production line and logistics transportation. Finished products are automatically packed into bags by a packaging machine, transferred to a designated temporary storage area by a forklift, stacked manually, and finally transported over long distances and stored at high positions through hoisting. This series of operations involves complex interweaving of forklifts, hoisting equipment and stacking personnel, forming a high-risk working environment.

[0003] At present, safety management for such scenarios mainly relies on traditional physical isolation, ground markings, manual supervision and basic acousto-optic alarm devices. However, these methods have significant limitations: 1. Passive defense and delayed response: Warning lines and sign boards can only play a reminding role, and cannot prevent personnel from entering in violation of regulations. Manual supervision is limited by blind spots, fatigue and response speed, making it difficult to achieve all-weather, no-dead-angle real-time monitoring; 2. Lack of intelligent recognition capability: Existing basic monitoring cameras are usually only used for post-event video recording and traceability, and do not have front-end intelligent analysis capability. The system cannot automatically distinguish personnel from materials, so when personnel appear in the forklift transportation path or hoisting operation area, it cannot immediately trigger accurate early warning or emergency braking; 3. High risk of cooperative operation: In the handover area of forklift transfer and crane hoisting, the working space overlaps, and both dynamic and static risks coexist. There are blind spots during the driving of forklifts, while hoisting operations face risks such as heavy object swinging and falling objects from high altitude. The prior art can hardly delimit the dynamic working areas of these two devices in real time and detect personnel intrusion.

[0004] In view of the above deficiencies in the prior art, the present invention aims to provide an intelligent visual personnel safety monitoring system for a hoisting and forklift cooperative working environment. The system captures real-time images through a camera and sends them to a deep learning detection model, completes alarm judgment through post-processing logic, and realizes immediate triggering of accurate early warning or emergency braking through linkage with forklifts and hoisting equipment. Summary of the Invention

[0005] The object of the present invention is to propose a personnel safety monitoring method and system for a hoisting and forklift cooperative working environment to solve the problems in the prior art.

[0006] A personnel safety monitoring method for a hoisting and forklift cooperative working environment, comprising the following steps: S1. Acquire the real-time video stream of the work area and decode it into continuous frame images; S2. Perform target detection on the image, identify large and small targets, and output the first detection box of the large target and the second detection box of the small target; the large target is a forklift or lifting equipment, and the small target is a person; S3. Generate an adaptive expansion region in the image based on the position and size of the first detection box; S4. Determine whether the second detection box falls within the adaptive expansion area, and generate a violation signal based on the judgment results of multiple consecutive frames; S5. Trigger audible and visual alarms based on violation signals and generate equipment control commands to perform deceleration or emergency stop operations on forklifts or lifting equipment through the industrial control system.

[0007] Preferably, step S2 specifically includes: S21. Use a one-stage target detection model to perform global detection on the image and output the first detection box of the large target. S22. Perform adaptive dilation processing based on the position of the first detection box, determine the dilation region, and crop the dilation region out of the image; S23. A two-stage target detection model is used to perform local detection on the cropped region image, and the second detection box of the small target is output. S24. Map the local coordinates of the second detection box back to the original coordinates of the image.

[0008] Preferably, the adaptive dilation process includes: Calculate the lateral and vertical expansion of the first detection frame; The lateral expansion amount is equal to the width of the first detection frame multiplied by the lateral expansion coefficient and then rounded down; The longitudinal expansion amount is equal to the height of the first detection frame multiplied by the longitudinal expansion coefficient and then rounded down; Based on the lateral and longitudinal expansion amounts, the boundary of the first detection frame is expanded outward to obtain the expansion region.

[0009] Preferably, when multiple expansion regions overlap, the minimum bounding rectangle strategy is used to merge the overlapping expansion regions into a single region before cropping.

[0010] Preferably, step S3 specifically includes: S31. Filter the first detection box based on the detection target box score and the detection box size threshold, and filter out low-quality detection results; S32. Calculate the coordinates of the center point of the first detection frame after filtering; S33. Perform asymmetric expansion based on the center point position of the first detection frame to generate an adaptive expansion region; The asymmetric expansion includes: the left and right boundaries of the adaptive expansion region are consistent with the left and right boundaries of the first detection frame; the upper boundary of the adaptive expansion region is set at the middle position of the upper half of the first detection frame, and a basic expansion amount is added; the lower boundary of the adaptive expansion region is set inside the lower boundary of the first detection frame.

[0011] Preferably, step S4 specifically includes: S41. Determine whether the large target is in motion based on the displacement of the center point of the first detection box in multiple consecutive frames of images. S42. When the large target is in motion, determine whether the center point of the second detection box is located within the adaptive expansion area; S43. When the second detection box is located within the adaptive expansion area for multiple consecutive frames, a violation signal is generated.

[0012] Preferably, the method for determining whether a large target is in motion in step S41 is as follows: calculate the Euclidean distance from the center point of the first detection box in the current frame to the average position of the center points of multiple historical frames, and determine that the large target is in motion when the distance is greater than a preset motion threshold.

[0013] Preferably, in step S5, the industrial control system is a distributed control system, which sends equipment control commands to the actuators of forklifts or hoisting equipment through the interface of the distributed control system.

[0014] A personnel safety monitoring system for collaborative lifting and forklift operations is also proposed to implement the aforementioned monitoring method, including: The image acquisition module is used to acquire real-time video streams of the work area and decode them into continuous frame images; The AI ​​detection module is used to detect objects in images and output a first detection box containing large objects. The first detection box is adaptively expanded and cropped. The cropped region image is then detected by a two-stage object detection model, and a second detection box containing small objects is output. The coordinate transformation module is used to convert the coordinates of the second detection box relative to the cropped image obtained by the two-stage detection model into the coordinates of the original image, and merge the transformed coordinate results with the one-stage inference results as the inference results of the current frame. The data analysis module is used to filter the inference results, obtain target detection information, and design alarm logic to judge violations. The automatic alarm module is used to upload the alarm logic information to the PC or mobile terminal for alarm confirmation via the backend service; The device linkage module is used to combine effective alarms with device linkage to control the operation of the device.

[0015] Compared with existing technologies, the advantages of this invention are: 1. This invention employs a one-stage model for global detection of large targets, combined with adaptive dilation and pruning, and a two-stage model for fine-grained detection of small targets, effectively solving the problem of missed detection of small targets at long distances. Simultaneously, based on the adaptive generation of asymmetric expanded monitoring areas from the large target detection bounding box, it achieves accurate judgment of personnel intrusion in dynamic work areas. Coupled with a continuous multi-frame confirmation mechanism, it greatly reduces false alarms and false negatives.

[0016] 2. This invention deeply integrates visual detection results with industrial control systems. When personnel are detected entering the forklift or hoisting operation area in violation of regulations, it not only pushes alarm information in real time, but also automatically sends equipment control commands through the DCS interface to perform deceleration or emergency stop operations on the forklift or hoisting equipment. This realizes intrinsic safety protection from passive monitoring to active intervention, effectively avoiding the occurrence of safety accidents. Attached Figure Description

[0017] Figure 1 This is a flowchart illustrating the monitoring method in this invention.

[0018] Figure 2 This is a schematic diagram of the monitoring system in this invention. Detailed Implementation

[0019] To facilitate understanding of this application and to make the above-mentioned objectives, features and advantages of this application more apparent, the specific embodiments of this application will be described in detail below with reference to the accompanying drawings.

[0020] Reference Figure 1-2 As shown, a method for personnel safety monitoring in a lifting and forklift collaborative operation environment includes the following steps: S1. Acquire the real-time video stream of the work area and decode it into continuous frame images; Image acquisition utilizes a 4-megapixel high-definition camera specifically designed for complex indoor and outdoor environments. Its core features a 1 / 2.7-inch progressive scan CMOS sensor, supporting... With full HD resolution, it outputs clear and smooth video footage. The device features excellent day / night switching capabilities and infrared night vision, accurately capturing details even in complete darkness or low-light conditions, ensuring 24 / 7 surveillance without blind spots. In actual deployment, the camera is securely mounted in the warehouse, high off the ground. Above the 10-meter-high beam, the camera is precisely aimed at the core work area, achieving comprehensive monitoring and efficient detection of targets within the work area through optimized viewing angle coverage.

[0021] In the image acquisition module, in order to meet the training requirements, it is necessary not only to acquire video stream data in real time, but also to collect a large number of on-site operation images in advance from the surveillance cameras and to perform detailed annotation work on these images, specifically by marking large and small targets in the images with rectangular boxes.

[0022] S2. Perform target detection on the image, identify large targets (forklifts or lifting equipment) and small targets (people), and output the first detection box of the large target and the second detection box of the small target; The object detection model framework adopts The object detection framework is described below: The framework mainly consists of three parts, namely the backbone network (…). ), converged network ( ) and prediction networks ( ),in: Backbone network: by input size of The image was processed Through successive downsampling operations, the spatial resolution of the image is gradually reduced while semantic information is enhanced, ultimately constructing three feature levels with different receptive fields. Specifically, the network output... Feature diagrams for three specifications. Among them, The feature map retains rich spatial details, making it suitable for capturing small targets; while The feature maps contain deep semantic information and are mainly used to identify large objects. Then, the feature maps of different sizes are fused in batches through a fusion network. Fusion Network: To overcome the limitations of single-scale features, the fusion network adopts a dual-path feature pyramid structure, that is, through a bidirectional aggregation strategy of top-down and bottom-up, the feature maps of three different levels output by the backbone network are deeply interacted, which greatly enhances the model's ability to represent targets of different sizes. Finally, each fused feature map is connected to the prediction network. Prediction Network: As the output of the network, the prediction network receives the fused high-dimensional feature map and uses a decoupled head design to output the target class probability and bounding box coordinates of each feature point independently. The model can more accurately locate target objects within the working area and output high-confidence class information.

[0023] The specific construction and training of the object detection model are as follows: Obtained through the image acquisition module The size is picture And filter the images to obtain Zhang Youxiao's training pictures The valid images are labeled and then scaled and padded to unify them. The image size, and sent. The target detection model is trained to obtain a one-stage detection model. The output of this model yields the raw detection box data, which includes the first detection boxes corresponding to multiple independent large targets. An adaptive dilation method is used to dilate the regions of these first detection boxes, scaling them to a fixed size, and then cropping them. The cropped images containing smaller targets are then cleaned, labeled, and fed into the system. The target detection model is trained to obtain a two-stage detection model for detecting small targets.

[0024] Among them, adaptive expansion calculation: For the first detection box Its new coordinates after expansion The calculation is as follows:

[0025] in, This is the lateral expansion coefficient. The longitudinal expansion coefficient is... This indicates rounding down to the nearest integer.

[0026] Calculate the new coordinates after dilation and utilize and The function ensures that the coordinates do not exceed the boundaries of the original image, that is:

[0027] in, The function's purpose is to retrieve... and The larger one, The function's purpose is to retrieve... and The smaller one.

[0028] To eliminate duplicate detections caused by overlapping cropping or close arrangement of multiple targets, a merging strategy is adopted: During the merging process, determine the two arbitrary rectangles. and Whether two rectangles overlap is determined using the inverse operation of boundary exclusion logic. The necessary and sufficient condition for two rectangles not to overlap is:

[0029] Therefore, the condition for two rectangles to overlap is the inverse of the above logic:

[0030] When overlap is detected, the minimum bounding rectangle strategy is adopted, that is, the two boxes... and Merge into a new rectangle : .

[0031] In the AI ​​detection module, a first-stage detection model is used to detect consecutive frames of video streams from cameras in the work area to obtain the first detection box. After adaptive dilation and boundary constraints, the detection boxes are merged to optimize redundant overlapping detection boxes. The clipped boxes are then fed into the second-stage detection model for small target inference.

[0032] Since the inference results of the two-stage detection model are relative to the cropped image, in order to accurately locate the target in the original image, it is necessary to convert the relative coordinates to the original image coordinates. This requires the use of a coordinate transformation module, the specific working process of which is as follows: Get the coordinates of the original image The corresponding top-left corner coordinates of the cropping area The top-left corner coordinates of the inference results from the two-stage detection model bottom right corner coordinates The coordinates of the detection box in the upper left corner of the original image are calculated according to the translation transformation formula. .

[0033] The translation transformation formula is:

[0034] Similarly, the coordinates of the detection box in the lower right corner of the original image are calculated using the translation transformation formula. .

[0035] The translation transformation formula is:

[0036] Finally, the transformed coordinate results are merged with the first-stage inference results to form the inference result for the current frame.

[0037] S3. Generate an adaptive expansion region in the image based on the position and size of the first detection box; After obtaining the reasoning results, a data analysis module is needed to analyze them. First, the targets are filtered by suppressing targets whose scores are below a threshold, and by suppressing targets whose width and height are below a threshold. The filtered results are then obtained.

[0038] in, This represents the score of the target bounding box; Indicates the score threshold; Indicate the coordinates of the top left and bottom right corners of the target; , Indicates the width and height thresholds of the target.

[0039] After obtaining the screening results, since the target object's position is constantly updated throughout the work area, an adaptive expansion region generation strategy is adopted. This allows the model to focus only on the work area surrounding the target object, reducing computation on irrelevant scenes and processing the detection box coordinates. ,in The coordinates of the top left and bottom right corners are given respectively; then the target center point is calculated from the coordinates of the top left and bottom right corners. .

[0040] in,

[0041] Adaptive expansion region set to Based on the horizontal movement characteristics of the target object during on-site operations, it is assumed that its horizontal space occupation is mainly determined by the target width, requiring no additional expansion. Therefore, the left and right boundaries of the detection frame remain unchanged.

[0042] Based on the center point of the target detection box, asymmetric expansion is performed, positioning the upper boundary at the middle of the upper half of the detection box, and adding a basic expansion amount to cover the upper range of motion of the target.

[0043] in, Indicates taking the absolute value; Indicates the preset extended pixel base; The lower boundary is contracted upwards to move it upwards, eliminating background interference such as the ground and focusing on the effective monitoring area of ​​the target subject.

[0044] Then, boundary constraints limit the effective width of the image. and height Within the designated area, ensure that the area does not exceed the boundary. That is:

[0045] Finally, the four corner points will be expanded. This serves as the final output for the monitored area.

[0046] S4. Determine whether the second detection box falls within the adaptive expansion area, and generate a violation signal based on the judgment results of multiple consecutive frames; When a target is operating within the work area, a violation can only be determined if the large target is in motion. To determine whether a large target is in motion, a continuous multi-frame judgment strategy is employed, saving multiple frames of data. By calculating each center point to the mean center point Euclidean distance .

[0047] in,

[0048]

[0049] Obtain the Euclidean distance Later, when At this time, it is determined that the target is in a running state.

[0050] When the current target is in motion, determine the center point of the smaller target. ,Right now:

[0051] Whether it is within the adaptive expansion detection region of a large target, i.e.:

[0052] When a small target is within the adaptively expanded detection area, the number of effective frames is counted. Add one, that is:

[0053] When small goals are continuous If the frame meets the condition, an alarm is triggered, namely:

[0054] in, This indicates the threshold number of times an alarm will be triggered.

[0055] S5. Trigger audible and visual alarms based on violation signals and generate equipment control commands to perform deceleration or emergency stop operations on forklifts or lifting equipment through the industrial control system.

[0056] When detecting large targets, if multiple frames of small targets appear within the adaptively expanded detection area, the algorithm will automatically identify and trigger an alarm. In the automatic alarm module, alarm information is simultaneously pushed to mobile devices or PCs in text and image format via the early warning platform, directly reaching management personnel. Management personnel can then quickly locate the alarm position on-site and make targeted adjustments to the equipment based on detailed alarm data. This mechanism effectively ensures that on-site management personnel can receive and handle alarms anytime, anywhere, and in a timely manner, improving response efficiency and management flexibility.

[0057] The timeliness issue in handling violations of small targets from alarm to resolution can be effectively controlled by linking the detection results with large target equipment, and by connecting the automatic alarm module with the equipment linkage module.

[0058] As an important part of the construction of intelligent factories, the DCS system enables real-time monitoring of equipment parameters and decentralized remote control within the factory.

[0059] Therefore, the DCS system interface can be connected to the video surveillance system and the equipment control system.

[0060] When the video surveillance system detects a violation of operating procedures, it not only pushes an alarm to management personnel but also transmits the equipment anomaly signal to the DCS system via the DCS interface. The distributed control system can then automatically stop equipment operation based on the received anomaly signal, while management personnel can simultaneously intervene manually through the DCS system. This reduces the time delay between AI-generated automatic alarms and on-site management intervention, thereby improving production efficiency.

[0061] Example Step 1: collection The images were cleaned to obtain... The size is To effectively train and annotate images, a semi-automatic annotation method is adopted. First, a general large model is used to assist in annotation to quickly obtain the basic outline of large targets. Then, the preliminary annotation results are fine-tuned by manual intervention to ensure the accuracy and quality of the annotation. In the entire annotation process, there are only three types of target recognition objects: large targets: "forklift" and "crane", and small targets: "people".

[0062] Step 2: The image is detected using a one-stage detection model to obtain a first detection box containing a large target. Through adaptive dilation and boundary constraints, the detection boxes are then merged to optimize redundant overlapping detection boxes. Finally, the boxes are cropped and fed into a two-stage detection model to obtain a second detection box containing a small target.

[0063] In this embodiment, a method of first shrinking and then filling is adopted, because First, reduce the width of the image proportionally. Then, the height is fixed by filling with fixed pixels at the top and bottom. Finally, it was unified to Image size.

[0064] For one of the first detection boxes Its new coordinates after expansion The calculation is as follows:

[0065] in, This is the lateral expansion coefficient. The longitudinal expansion coefficient is... This indicates rounding down to the nearest integer.

[0066] Calculate the new coordinates after dilation and utilize and The function ensures that the coordinates do not exceed the boundaries of the original image.

[0067]

[0068] For images with two first detection boxes, determine the two first detection boxes. and Does it overlap?

[0069] After determining overlap, the minimum bounding rectangle strategy is adopted, which involves dividing the two first detection boxes... and Merge into a new rectangle :

[0070] Step 3: Since the inference results of the two-stage detection model are relative to the cropped image coordinates, in order to accurately locate the target in the original image, it is necessary to convert the relative coordinates to the original image coordinates.

[0071] Obtain the detection bounding box information of the original image, including one of the detection bounding boxes. The coordinates of the top left corner of the corresponding cropping area The top-left corner coordinates of the inference results from the two-stage detection model bottom right corner The coordinates of the detection box in the upper left corner of the original image are calculated according to the translation transformation formula. .

[0072] The translation transformation formula is:

[0073] Similarly, the coordinates of the detection box in the lower right corner of the original image are calculated using the translation transformation formula. .

[0074] The translation transformation formula is:

[0075] Finally, the transformed coordinate results are merged with the first-stage inference results to form the inference result for the current frame.

[0076] Step 4: After obtaining the reasoning results, the target objects are screened, that is:

[0077] After obtaining the screening results, since the target object's position is constantly updated throughout the work area, an adaptive expansion region generation strategy is adopted. This allows the model to focus only on the work area surrounding the target object, reducing computation on irrelevant scenes, and processing the detection box coordinates of one frame. ,in Given the coordinates of the top left and bottom right corners, calculate the target center point. .

[0078] in,

[0079] The adaptive expansion region is set to [ Based on the horizontal movement characteristics of the target object during operation, it is assumed that its horizontal space occupation is mainly determined by the target width, requiring no additional expansion. Therefore, the left and right boundaries of the detection frame remain unchanged.

[0080] Based on the height and center point position of the target detection box, asymmetric expansion is performed, positioning the upper boundary at the middle of the upper half of the detection box, and adding a basic expansion amount to cover the upper range of motion of the target.

[0081] Among them, take =30; The lower boundary is contracted upwards to move it upwards, eliminating background interference such as the ground and focusing on the effective monitoring area of ​​the target subject.

[0082] Then, boundary constraints limit the effective width of the image. and height Within the designated area, ensure that the area does not exceed the boundary. That is:

[0083] The final four corner points after expansion:

[0084] Right now The monitored area is output as the final raw video stream size.

[0085] Step 5: This example By calculating each center point to the mean center point Euclidean distance Obtain the Euclidean distance Later, when At that time, it is determined that the target is in motion.

[0086] In this example, 5 consecutive frames are saved, that is... ,Right now:

[0087]

[0088]

[0089]

[0090]

[0091] Once the current target is in motion, determine the center point of the smaller target. ,Right now:

[0092] Whether it is within the adaptive expansion detection region of a large target, i.e.:

[0093] When a small target is within the adaptively expanded detection area, the number of effective frames is counted. Add one, that is:

[0094] When small goals are continuous If the frame meets the condition, an alarm is triggered, namely:

[0095] Step 6: When detecting large targets, if multiple frames of small targets appear within the adaptively expanded detection area, the algorithm will automatically identify and trigger an alarm. In this embodiment, two methods are used for alarm generation: one is to display the alarm information on a large screen on a PC, and the other is to upload the alarm information to a mobile device for display.

[0096] Simultaneously, abnormal equipment signals can be transmitted to the DCS system via the DCS interface. The DCS system can immediately perform deceleration or emergency stop operations on the forklift and lifting equipment based on the received abnormal equipment signals.

[0097] As is known from common technical knowledge, this invention can be implemented through other embodiments that do not depart from its spirit or essential characteristics. Therefore, the disclosed embodiments described above are merely illustrative in all respects and are not the only ones. All modifications within the scope of this invention or its equivalents are included in this invention.

Claims

1. A method for personnel safety monitoring in a lifting and forklift collaborative operation environment, characterized in that, Includes the following steps: S1. Acquire the real-time video stream of the work area and decode it into continuous frame images; S2. Perform target detection on the image, identify large and small targets, and output the first detection box of the large target and the second detection box of the small target; the large target is a forklift or lifting equipment, and the small target is a person; S3. Generate an adaptive expansion region in the image based on the position and size of the first detection box; S4. Determine whether the second detection box falls within the adaptive expansion area, and generate a violation signal based on the judgment results of multiple consecutive frames; S5. Trigger audible and visual alarms based on violation signals and generate equipment control commands to perform deceleration or emergency stop operations on forklifts or lifting equipment through the industrial control system.

2. The method for personnel safety monitoring in a lifting and forklift collaborative operation environment according to claim 1, characterized in that: Step S2 specifically includes: S21. Use a one-stage target detection model to perform global detection on the image and output the first detection box of the large target. S22. Perform adaptive dilation processing based on the position of the first detection box, determine the dilation region, and crop the dilation region out of the image; S23. A two-stage target detection model is used to perform local detection on the cropped region image, and the second detection box of the small target is output. S24. Map the local coordinates of the second detection box back to the original coordinates of the image.

3. The method for personnel safety monitoring in a lifting and forklift collaborative operation environment according to claim 2, characterized in that: The adaptive expansion process includes: Calculate the lateral and vertical expansion of the first detection frame; The lateral expansion amount is equal to the width of the first detection frame multiplied by the lateral expansion coefficient and then rounded down; The longitudinal expansion amount is equal to the height of the first detection frame multiplied by the longitudinal expansion coefficient and then rounded down; Based on the lateral and longitudinal expansion amounts, the boundary of the first detection frame is expanded outward to obtain the expansion region.

4. A method for personnel safety monitoring in a lifting and forklift collaborative operation environment according to claim 3, characterized in that: When multiple expansion regions overlap, the minimum bounding rectangle strategy is used to merge the overlapping expansion regions into a single region before cropping.

5. A method for personnel safety monitoring in a lifting and forklift collaborative operation environment according to claim 1, characterized in that: Step S3 specifically includes: S31. Filter the first detection box based on the detection target box score and the detection box size threshold, and filter out low-quality detection results; S32. Calculate the coordinates of the center point of the first detection frame after filtering; S33. Perform asymmetric expansion based on the center point position of the first detection frame to generate an adaptive expansion region; The asymmetric expansion includes: the left and right boundaries of the adaptive expansion region are consistent with the left and right boundaries of the first detection frame; the upper boundary of the adaptive expansion region is set at the middle position of the upper half of the first detection frame, and a basic expansion amount is added; the lower boundary of the adaptive expansion region is set inside the lower boundary of the first detection frame.

6. A method for personnel safety monitoring in a lifting and forklift collaborative operation environment according to claim 1, characterized in that: Step S4 specifically includes: S41. Determine whether the large target is in motion based on the displacement of the center point of the first detection box in multiple consecutive frames of images. S42. When the large target is in motion, determine whether the center point of the second detection box is located within the adaptive expansion area; S43. When the second detection box is located within the adaptive expansion area for multiple consecutive frames, a violation signal is generated.

7. A method for personnel safety monitoring in a lifting and forklift collaborative operation environment according to claim 6, characterized in that: The method for determining whether a large target is in motion in step S41 is as follows: calculate the Euclidean distance from the center point of the first detection box in the current frame to the average position of the center points in historical multiple frames. When the distance is greater than a preset motion threshold, it is determined that the large target is in motion.

8. A method for personnel safety monitoring in a lifting and forklift collaborative operation environment according to claim 1, characterized in that: In step S5, the industrial control system is a distributed control system, which sends equipment control commands to the actuators of forklifts or hoisting equipment through the interface of the distributed control system.

9. A personnel safety monitoring system for use in a lifting and forklift collaborative operation environment, used to implement the monitoring method according to any one of claims 1-8, characterized in that: include: The image acquisition module is used to acquire real-time video streams of the work area and decode them into continuous frame images; The AI ​​detection module is used to detect objects in images and output a first detection box containing large objects. The first detection box is adaptively expanded and cropped. The cropped region image is then detected by a two-stage object detection model, and a second detection box containing small objects is output. The coordinate transformation module is used to convert the coordinates of the second detection box relative to the cropped image obtained by the two-stage detection model into the coordinates of the original image, and merge the transformed coordinate results with the one-stage inference results as the inference results of the current frame. The data analysis module is used to filter the inference results, obtain target detection information, and design alarm logic to judge violations. The automatic alarm module is used to upload the alarm logic information to the PC or mobile terminal for alarm confirmation via the backend service; The device linkage module is used to combine effective alarms with device linkage to control the operation of the device.