Method, device and equipment for identifying falling of personnel in distribution center and storage medium
Through the combination of the lightweight RT-DETR network and the federated learning framework, the confidence threshold is dynamically adjusted, and the timely identification of fall accidents in personnel in the distribution center is solved, the detection accuracy and sensitivity are improved, and safe operation is ensured.
Patent Information
- Application Number
- CN202510486034.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-17
- Publication Date
- 2025-08-12
AI Technical Summary
The lack of real-time monitoring of personnel behavior in the distribution center has led to the inability to detect fall accidents in a timely manner, affecting safety management and operational efficiency.
A lightweight RT-DETR object detection network is adopted, combining multi-scale feature fusion and self-attention mechanism, distributed training is carried out through the federated learning framework, and the scene complexity score is generated using the pre-trained population density estimation model, dynamically adjusting the confidence threshold of the personnel fall recognition model, and triggering the alarm mechanism.
The detection accuracy of personnel with different distances and postures has been improved, misjudgment has been reduced, detection sensitivity has been improved, alarms have been triggered in a timely manner, and personnel safety has been ensured.
Smart Images

Figure CN120472526A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of logistics technology, and in particular to a method, device, equipment and storage medium for identifying a person falling in a distribution center. Background Art
[0002] In today's booming logistics industry, efficient and safe operations of distribution centers, key hubs for cargo distribution and transshipment, are crucial. Distribution centers are complex environments, with personnel frequently performing activities such as cargo handling and equipment operation. Existing safety management measures have significant limitations.
[0003] Currently, distribution center safety management focuses primarily on monitoring equipment operating conditions and the cargo storage environment. Regarding equipment operation, various sensors and monitoring systems provide real-time visibility into equipment performance, promptly identifying potential equipment failures and ensuring smooth cargo handling, sorting, and other operations. Regarding the cargo storage environment, devices such as temperature and humidity sensors and smoke alarms are widely used to ensure that the storage environment meets requirements and prevent damage.
[0004] However, these measures neglect real-time monitoring and analysis of human behavior. Distribution centers are characterized by frequent human activity, creating a working environment with numerous safety risks, such as slippery floors, cluttered cargo, and improper equipment operation. These factors contribute to frequent accidents, including falls. Due to the lack of real-time monitoring of human behavior, accidents often go undetected. Currently, manual inspections are the primary means of identifying problems, but these inspections have time gaps and lack comprehensive, all-encompassing coverage, easily delaying the optimal time for rescue. While surveillance cameras can be helpful, they are typically reviewed after the fact, making it difficult to issue alerts and initiate rescue efforts immediately after an incident.
[0005] If an accident such as a fall occurs and timely rescue is not provided, serious casualties are likely to occur, causing immense pain and loss to the employee and their family, and negatively impacting the normal operation of the distribution center. Employee injuries can lead to decreased work efficiency, labor shortages, and even trigger a chain reaction, disrupting the normal flow of goods and the overall operation of the distribution center, ultimately damaging the company's economic benefits and reputation.
[0006] Therefore, the existing technology still needs to be improved and developed. Summary of the Invention
[0007] The present invention provides a method, device, equipment and storage medium for identifying a person falling in a distribution center, which are used to detect whether a person has fallen in the distribution center and trigger an alarm mechanism when a person falls.
[0008] A first aspect of the present invention provides a method for identifying falls of people in a distribution center, the method comprising: collecting personnel image data in different scenarios of the distribution center, and preparing a local training set based on the personnel image data; constructing a target detection network based on a lightweight RT-DETR, and introducing a multi-scale feature fusion mechanism and a self-attention mechanism in the target detection network to obtain a model to be trained; based on the local training set, introducing a federated learning framework, and jointly performing distributed training on the model to be trained by edge computing nodes of multiple distribution centers to obtain a personnel fall recognition model; regularly obtaining real-time images of the distribution center, inputting the real-time images into a pre-trained crowd density estimation model, obtaining a density heat map output by the crowd density estimation model, and calculating a scene complexity score based on the density heat map; adjusting the confidence threshold of the personnel fall recognition model based on the scene complexity score, and inputting the real-time images into the adjusted personnel fall recognition model to detect whether there is a person fall behavior in the distribution center, and triggering an alarm mechanism when the detection result shows that there is a person fall behavior.
[0009] Optionally, in a first implementation method of the first aspect of the present invention, the collecting of personnel image data in different scenarios of the distribution center and the production of a local training set based on the personnel image data include: collecting personnel image data in different scenarios of the distribution center, and preprocessing the personnel image data to obtain preprocessed data; labeling personnel postures and falling behaviors in the preprocessed data to obtain sample data; and dynamically enhancing the sample data using a generative adversarial network to generate a local training set.
[0010] Optionally, in a second implementation method of the first aspect of the present invention, the target detection network based on lightweight RT-DETR is constructed, and a multi-scale feature fusion mechanism and a self-attention mechanism are introduced into the target detection network to obtain a model to be trained, including: constructing a target detection network based on lightweight RT-DETR; adding a multi-scale feature fusion module to the target detection network to obtain an adjusted target detection network; introducing a self-attention mechanism in the encoder or decoder of the adjusted target detection network to obtain a model to be trained.
[0011] Optionally, in a third implementation of the first aspect of the present invention, a federated learning framework is introduced based on the local training set, and the edge computing nodes of multiple distribution centers are combined to perform distributed training on the model to be trained to obtain a personnel fall recognition model, including: using the edge computing node of the distribution center, using the local training set to perform initial training on the model to be trained, and generating local parameters; encrypting the local parameters and uploading them to the central server so that the central server aggregates the local parameters of the edge computing nodes of all distribution centers, and using the federated averaging algorithm to optimize the local parameters of the edge computing nodes of all distribution centers to obtain optimized parameters; obtaining the optimized parameters issued by the central server, and using the optimized parameters to optimize the model after the initial training to obtain a personnel fall recognition model.
[0012] Optionally, in a fourth implementation of the first aspect of the present invention, the real-time image of the distribution center is periodically obtained, the real-time image is input into a pre-trained crowd density estimation model, a density heat map output by the crowd density estimation model is obtained, and a scene complexity score is calculated based on the density heat map, including: periodically obtaining a real-time image of the distribution center, and inputting the real-time image into a pre-trained crowd density estimation model, and a density heat map output by the crowd density estimation model; performing an integration operation on the density heat map to obtain a real-time total number of people estimate, and calculating the real-time total number of people ratio based on the real-time total number of people estimate and the preset maximum safe carrying capacity of the distribution center, and calculating the entropy value standard value of the density heat map; and calculating the scene complexity score based on the total number of people ratio and the entropy value standard value.
[0013] Optionally, in a fifth implementation of the first aspect of the present invention, the confidence threshold of the personnel fall recognition model is adjusted based on the scene complexity score, and the real-time image is input into the adjusted personnel fall recognition model to detect whether there is a personnel fall behavior in the distribution center. When the detection result is that there is a personnel fall behavior, an alarm mechanism is triggered, including: calculating the confidence threshold to be updated based on the scene complexity score, and adjusting the confidence threshold of the personnel fall recognition model according to the confidence threshold to be updated to obtain the adjusted personnel fall recognition model; preprocessing the real-time image, and inputting the preprocessed real-time image into the adjusted personnel fall recognition model, and obtaining the detection result output by the adjusted personnel fall recognition model; when the detection result is that there is a personnel fall behavior, the alarm mechanism is triggered.
[0014] Optionally, in a sixth implementation of the first aspect of the present invention, the confidence threshold to be updated is calculated based on the scene complexity score, and the confidence threshold of the person fall recognition model is adjusted according to the confidence threshold to be updated to obtain the adjusted person fall recognition model, including: using exponentially weighted moving average to smooth the scene complexity score to obtain a smoothed scene complexity score; calculating the confidence threshold to be updated based on the maximum and minimum values of the confidence threshold of the person fall recognition model and the smoothed scene complexity score; adjusting the confidence threshold of the person fall recognition model according to the confidence threshold to be updated to obtain the adjusted person fall recognition model.
[0015] A second aspect of the present invention provides a device for identifying falls in a distribution center, comprising: a production module for collecting personnel image data in different scenarios of the distribution center and producing a local training set based on the personnel image data; a construction module for constructing a target detection network based on a lightweight RT-DETR, and introducing a multi-scale feature fusion mechanism and a self-attention mechanism into the target detection network to obtain a model to be trained; a training module for introducing a federated learning framework based on the local training set, and performing distributed training on the model to be trained in conjunction with edge computing nodes of multiple distribution centers to obtain a personnel fall recognition model; a calculation module for regularly acquiring real-time images of the distribution center, inputting the real-time images into a pre-trained crowd density estimation model, acquiring a density heat map output by the crowd density estimation model, and calculating a scene complexity score based on the density heat map; a detection module for adjusting the confidence threshold of the personnel fall recognition model based on the scene complexity score, and inputting the real-time images into the adjusted personnel fall recognition model to detect whether there is a person fall behavior in the distribution center, and triggering an alarm mechanism when the detection result shows that there is a person fall behavior.
[0016] Optionally, in a first implementation method of the second aspect of the present invention, the production module includes: a collection unit, used to collect personnel image data in different scenarios of the distribution center, and preprocess the personnel image data to obtain preprocessed data; a labeling unit, used to label personnel postures and falling behaviors in the preprocessed data to obtain sample data; a production unit, used to dynamically enhance the sample data using a generative adversarial network, and generate a local training set.
[0017] Optionally, in a second implementation of the second aspect of the present invention, the construction module includes: a construction unit for constructing a target detection network based on a lightweight RT-DETR; an adding unit for adding a multi-scale feature fusion module to the target detection network to obtain an adjusted target detection network; and an updating unit for introducing a self-attention mechanism in the encoder or decoder of the adjusted target detection network to obtain a model to be trained.
[0018] Optionally, in a third implementation of the second aspect of the present invention, the training module includes: a training unit, used to perform initial training on the model to be trained using the local training set through the edge computing node of the distribution center, and generate local parameters; an encryption unit, used to encrypt the local parameters and upload them to the central server so that the central server aggregates the local parameters of the edge computing nodes of all distribution centers, and uses a federal averaging algorithm to optimize the local parameters of the edge computing nodes of all distribution centers to obtain optimized parameters; an optimization unit, used to obtain the optimized parameters issued by the central server, and use the optimized parameters to optimize the model after initial training to obtain a personnel fall recognition model.
[0019] Optionally, in a fourth implementation of the second aspect of the present invention, the calculation module includes: an acquisition unit, used to periodically acquire real-time images of the distribution center, and input the real-time images into a pre-trained crowd density estimation model to obtain a density heat map output by the crowd density estimation model; a first calculation unit, used to perform an integral operation on the density heat map to obtain a real-time total number estimate, and calculate the real-time total number ratio based on the real-time total number estimate and the preset maximum safe carrying capacity of the distribution center, and calculate the entropy value standard value of the density heat map; a second calculation unit, used to calculate a scene complexity score based on the total number ratio and the entropy value standard value.
[0020] Optionally, in a fifth implementation of the second aspect of the present invention, the detection module includes: an adjustment unit, used to calculate a confidence threshold to be updated based on the scene complexity score, and adjust the confidence threshold of the person fall recognition model according to the confidence threshold to be updated to obtain an adjusted person fall recognition model; a detection unit, used to preprocess the real-time image, and input the preprocessed real-time image into the adjusted person fall recognition model, and obtain the detection result output by the adjusted person fall recognition model; an alarm unit, used to trigger an alarm mechanism when the detection result is that a person falls.
[0021] The third aspect of the present invention provides a device for identifying falls of people in a distribution center, comprising: a memory and at least one processor, wherein the memory stores computer-readable instructions, and the memory and the at least one processor are interconnected via a line; the at least one processor calls the computer-readable instructions in the memory so that the device for identifying falls of people in a distribution center executes the various steps of the method for identifying falls of people in a distribution center as described above.
[0022] A fourth aspect of the present invention provides a computer-readable storage medium, which stores computer-readable instructions. When the computer-readable storage medium is run on a computer, it enables the computer to execute the various steps of the method for identifying falls of people in a distribution center as described above.
[0023] In the technical solution provided by the present invention, a lightweight RT-DETR target detection network is adopted, and a multi-scale feature fusion mechanism and a self-attention mechanism are introduced. While ensuring detection efficiency, the detection accuracy of people at different distances and postures and the feature extraction capability of complex scenes are significantly improved; moreover, during the training process, a federated learning framework is introduced for distributed model training, which not only protects the data security of each distribution center, but also makes full use of multi-party data and computing resources, accelerates the training speed and improves model performance; in addition, a pre-trained crowd density estimation model is used to generate a scene complexity score, and the confidence threshold of the personnel fall recognition model is dynamically adjusted according to the score, so that the model can flexibly adjust the detection strategy according to different scenarios, effectively reduce misjudgment and improve detection sensitivity, and promptly trigger the alarm mechanism when a fall behavior is detected, which can quickly respond to fall accidents and provide strong protection for the safety of personnel in the distribution center. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] Figure 1 A first flow chart of a method for identifying a person falling in a distribution center provided by an embodiment of the present invention;
[0025] Figure 2 A second flow chart of the method for identifying a person falling in a distribution center provided by an embodiment of the present invention;
[0026] Figure 3 A third flow chart of the method for identifying a person falling in a distribution center provided by an embodiment of the present invention;
[0027] Figure 4 A fourth flow chart of the method for identifying a person falling in a distribution center provided by an embodiment of the present invention;
[0028] Figure 5 A fifth flow chart of the method for identifying a person falling in a distribution center provided by an embodiment of the present invention;
[0029] Figure 6A sixth flow chart of the method for identifying a person falling in a distribution center provided by an embodiment of the present invention;
[0030] Figure 7 A schematic diagram of the structure of a device for identifying a person falling in a distribution center provided by an embodiment of the present invention;
[0031] Figure 8 A schematic diagram of the structure of a device for identifying falls in a distribution center provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0032] The embodiment of the present invention provides a method, device, equipment and storage medium for identifying a person falling in a distribution center. The method is used to detect whether a person has fallen in the distribution center and trigger an alarm mechanism when a person falls. The method includes: collecting personnel image data in different scenarios of a distribution center, and preparing a local training set based on the personnel image data; constructing a target detection network based on a lightweight RT-DETR, and introducing a multi-scale feature fusion mechanism and a self-attention mechanism into the target detection network to obtain a model to be trained; based on the local training set, introducing a federated learning framework, and jointly performing distributed training on the model to be trained by edge computing nodes of multiple distribution centers to obtain a personnel fall recognition model; regularly obtaining real-time images of the distribution center, inputting the real-time images into a pre-trained crowd density estimation model, obtaining a density heat map output by the crowd density estimation model, and calculating a scene complexity score based on the density heat map; based on the scene complexity score, adjusting the confidence threshold of the personnel fall recognition model, and inputting the real-time images into the adjusted personnel fall recognition model to detect whether there is a personnel fall behavior in the distribution center, and triggering an alarm mechanism when the detection result shows that there is a personnel fall behavior.
[0033] The terms "first," "second," "third," "fourth," and the like (if any) in the description and claims of the present invention and in the accompanying drawings are used to distinguish similar objects and are not necessarily used to describe a particular order or precedence. It should be understood that the terms used in this manner are interchangeable where appropriate so that the embodiments described herein can be implemented in an order other than that illustrated or described herein. In addition, the terms "including" or "having" and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, method, system, product, or apparatus that includes a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units that are not explicitly listed or that are inherent to these processes, methods, products, or apparatus.
[0034] For ease of understanding, the specific process of the embodiment of the present invention is described below. Figure 1In a first embodiment of a method for identifying a person falling in a distribution center according to an embodiment of the present invention, the method includes:
[0035] S101. Collect personnel image data in different scenarios of the distribution center, and create a local training set based on the personnel image data.
[0036] It is understandable that the execution subject of the present invention can be a personnel fall identification device in a distribution center, or a terminal or a server, which is not limited here. The embodiment of the present invention is described by taking a server as the execution subject as an example.
[0037] In this example, the distribution center features a variety of complex scenarios, such as loading and unloading areas, sorting areas, and aisles. The behaviors and postures of people in these scenarios vary significantly. Collecting image data of people in these scenarios comprehensively covers various possible situations within the distribution center, providing a rich set of samples for model training.
[0038] Surveillance cameras deployed in the distribution center are used to collect images of people in different scenarios of the distribution center (such as high density, low light, and complex backgrounds), covering normal behavior, carrying goods, and falling movements, and accurately annotating them (location, behavior tags).
[0039] During the collection process, it is necessary to ensure data diversity (such as different perspectives and occlusion conditions) and annotation consistency (such as clearly distinguishing between "falling" and "squatting").
[0040] S102. Construct a target detection network based on lightweight RT-DETR, and introduce a multi-scale feature fusion mechanism and a self-attention mechanism into the target detection network to obtain a model to be trained.
[0041] In this embodiment, RT-DETR is a target detection network. Its lightweight design enables it to run efficiently on edge devices with limited computing resources, making it suitable for actual application scenarios in distribution centers.
[0042] In this embodiment, the sizes and distances of people in the distribution center vary. The multi-scale feature fusion mechanism adjusts the feature maps of different levels to the same resolution through upsampling and downsampling operations and then fuses them. This enables the model to capture the features of targets of different sizes and improves the detection accuracy of people at different distances and postures.
[0043] In this embodiment, the self-attention mechanism allows the model to automatically focus on key areas in the image, suppress redundant information, and enhance the model's ability to extract features of people in complex scenes, thereby improving detection accuracy.
[0044] S103. Based on the local training set, a federated learning framework is introduced to combine the edge computing nodes of multiple distribution centers to perform distributed training on the training model to obtain a person fall recognition model.
[0045] In this embodiment, with the help of a federated learning framework, distributed model training is carried out in conjunction with the edge computing nodes of multiple distribution centers. Each edge computing node uses a local training set to independently train the model, and selects a suitable optimization algorithm (such as stochastic gradient descent, Adam algorithm, etc.) to update the model parameters based on the characteristics of the local data. After the training is completed, the encrypted model parameters are uploaded to the central server. The central server uses aggregation strategies such as the federal averaging algorithm to comprehensively update the model parameters of each node, generate optimized global model parameters, and send them to each edge computing node. Through multiple rounds of such training iterations, a personnel fall recognition model with excellent performance is finally obtained.
[0046] S104. Regularly obtain real-time images of the distribution center, input the real-time images into a pre-trained crowd density estimation model, obtain a density heat map output by the crowd density estimation model, and calculate a scene complexity score based on the density heat map.
[0047] In this example, surveillance cameras deployed within the distribution center regularly capture real-time images of the distribution center. These images are fed into a pre-trained crowd density estimation model, which outputs a density heatmap reflecting the distribution of people. The density heatmap is analyzed according to specific calculation rules, taking into account factors such as the distribution center's spatial layout, aisle widths, and cargo stacking conditions. This generates a scene complexity score that quantifies the scene's complexity.
[0048] S105. Based on the scene complexity score, adjust the confidence threshold of the personnel fall recognition model, and input the real-time image into the adjusted personnel fall recognition model to detect whether there is a person falling in the distribution center. When the detection result shows that there is a person falling, trigger the alarm mechanism.
[0049] In this embodiment, the confidence threshold of the personnel fall recognition model is intelligently adjusted based on the generated scene complexity score. When the scene is complex and there are many interference factors, the confidence threshold is appropriately increased to reduce the false alarm rate; when the scene is relatively simple, the confidence threshold is lowered to improve the detection sensitivity, and the real-time image is input into the adjusted personnel fall recognition model for detection. Once the model determines that there is a person falling in the image, the alarm mechanism is immediately triggered, and relevant personnel are notified through sound and light alarms, message push, etc. to deal with it in time to ensure the safety of personnel in the distribution center.
[0050] This embodiment provides a method for identifying falls in distribution centers. The method uses a lightweight RT-DETR target detection network and introduces a multi-scale feature fusion mechanism and a self-attention mechanism. While ensuring detection efficiency, it significantly improves the detection accuracy of people at different distances and postures and the feature extraction capability for complex scenes. Moreover, during the training process, a federated learning framework is introduced for distributed model training, which not only protects the data security of each distribution center, but also fully utilizes multi-party data and computing resources, accelerates training speed and improves model performance. In addition, a pre-trained crowd density estimation model is used to generate a scene complexity score, and the confidence threshold of the person fall recognition model is dynamically adjusted based on the score. This allows the model to flexibly adjust the detection strategy according to different scenarios, effectively reducing misjudgments and improving detection sensitivity. When a fall is detected, an alarm mechanism is triggered in a timely manner, which can quickly respond to fall accidents and provide strong protection for the safety of people in the distribution center.
[0051] See also Figure 2 The second embodiment of the method for identifying a person falling in a distribution center according to the present invention includes:
[0052] S201: Collect personnel image data in different scenarios of the distribution center, and pre-process the personnel image data to obtain pre-processed data.
[0053] In this embodiment, the personnel image data is preprocessed, including image denoising to remove noise generated by camera shooting, transmission, and other processes to improve image quality; image normalization to unify image parameters such as size, brightness, and contrast to make model training more stable; cropping and scaling to adjust the image size according to model input requirements, focus on key areas, remove irrelevant background, and improve model training efficiency.
[0054] S202: Label the posture and falling behavior of the person in the pre-processed data to obtain sample data.
[0055] In this embodiment, tools such as LabelImg and CVAT can be used to annotate person positions (bounding boxes) and behavior labels (normal, fall, carrying, etc.). Annotating person postures can accurately mark the position and angle of various body parts, helping the model learn the characteristic differences between normal and abnormal behaviors. When annotating falls, the time, direction, and movement process of the fall are clearly defined, providing accurate training samples for the model to recognize falls.
[0056] S203: Use a generative adversarial network to dynamically enhance the sample data and generate a local training set.
[0057] In this embodiment, the generative adversarial network (GAN) consists of a generator and a discriminator. The generator generates new image data that is similar but not identical by learning the features of the labeled data, such as simulating different lighting conditions, human occlusion, and fall scenes at different angles. The discriminator is responsible for distinguishing between the generated fake data and the real labeled data, and the two are mutually adversarial and trained collaboratively. Through GAN data enhancement, the diversity of training data can be greatly expanded, allowing the model to be exposed to more fall samples in different scenarios, improving the generalization ability of the model, and enabling it to more accurately identify fall behaviors when facing complex and changeable distribution center scenarios in actual applications.
[0058] In this example, by collecting image data of personnel in various distribution center scenarios, a rich and diverse source material was provided for model training, covering possible behaviors of personnel in various working conditions, ensuring the model's broad adaptability. The detailed annotation of personnel postures and falls provided the model with accurate learning samples, enabling it to accurately identify the key characteristics of falls. The use of a generative adversarial network for dynamic data augmentation further enriched the diversity of training data, simulated various complex situations that may arise in real-world scenarios, and significantly enhanced the model's generalization capabilities.
[0059] See also Figure 3 A third embodiment of a method for identifying a person falling in a distribution center according to an embodiment of the present invention includes:
[0060] S301. Build a target detection network based on lightweight RT-DETR.
[0061] In this example, RT-DETR (Real-Time Detection Transformer) is an object detection network based on the Transformer architecture. It demonstrates high accuracy and speed in object detection tasks. The Transformer architecture's advantages lie in its ability to effectively capture global information in an image and its excellent handling of long-range dependencies, which are crucial for accurately detecting people in various positions and postures within distribution centers.
[0062] In this embodiment, the backbone network of the original RT-DETR (such as ResNet) is replaced with a lightweight network (such as MobileNetV3 or EfficientNet-Lite), which significantly reduces the number of parameters and computational complexity. The lightweight design takes into account the computing resource limitations in the actual application scenarios of the distribution center. Using a lightweight version of RT-DETR can reduce the computational complexity and parameter quantity of the model, and improve the operating efficiency of the model without losing too much detection performance, so that it can run quickly on edge computing devices or servers with limited resources to meet the needs of real-time detection.
[0063] S302: Add a multi-scale feature fusion module to the target detection network to obtain an adjusted target detection network.
[0064] In this example, the distances between people in the distribution center and the camera vary, resulting in differences in the size and resolution of the people in the image. A single-scale feature map cannot simultaneously capture the overall features of distant people and the detailed features of close-up people.
[0065] The multi-scale feature fusion module processes feature maps at different levels, combining low-resolution, semantically rich deep feature maps with high-resolution, detail-rich shallow feature maps. Lower-resolution feature maps are upsampled to increase their size; higher-resolution feature maps are downsampled as needed to maintain the same resolution across feature maps of different scales. Fusion is then performed through element-by-element addition and concatenation, enabling the model to fully utilize information from different scales, enhancing detection capabilities and improving accuracy for individuals at varying distances and postures.
[0066] S303. Introduce a self-attention mechanism into the encoder or decoder of the adjusted target detection network to obtain a model to be trained.
[0067] As you can understand, the self-attention mechanism allows the model to automatically focus on key areas in the image and suppress redundant information. After introducing the self-attention mechanism in the encoder or decoder, the model will dynamically assign different attention weights to features at different locations based on the content of the input image.
[0068] In this embodiment, local window self-attention (e.g., Swin Transformer) is introduced into the encoder to divide the image into multiple windows. Attention is calculated only within the window, reducing computational complexity. A shifted window mechanism is used to enable information exchange between windows, compensating for the field of view limitations of local attention.
[0069] In this embodiment, the lightweight RT-DETR is used as the cornerstone to build a target detection network. While ensuring detection accuracy, the operating efficiency of the model is greatly improved, enabling it to achieve real-time detection in the limited computing resource environment of the distribution center. The addition of the multi-scale feature fusion module cleverly solves the problem of feature differences in images of people at different distances. By fusing feature maps of different scales, the model can capture the overall information of distant people and pay attention to the details of close people, significantly improving the detection accuracy of people in various postures and positions; and the introduction of the self-attention mechanism enables the model to intelligently focus on key areas in the image and automatically ignore irrelevant background information. In complex distribution center scenarios, such as cargo occlusion and dense crowds, it can also accurately identify people, further enhancing the model's adaptability and detection capabilities to complex scenarios.
[0070] See also Figure 4 A fourth embodiment of a method for identifying a person falling in a distribution center according to an embodiment of the present invention includes:
[0071] S401: Use the local training set to perform initial training on the to-be-trained model through the edge computing node of the distribution center and generate local parameters.
[0072] In this embodiment, the distribution center's edge computing nodes have the advantage of being close to the data source, enabling local data processing and reducing data transmission latency and bandwidth pressure. Each distribution center's local training set contains image data of people in specific scenarios within that area, reflecting local environmental characteristics, occupant behavior patterns, and other information.
[0073] The edge computing node initially trains the target model using a local training set. During training, the model continuously adjusts its weights, biases, and other parameters based on the input data. After training, it generates a set of local parameters that reflect the data characteristics of the distribution center. This training method fully utilizes the data diversity of each distribution center, enabling the model to learn the characteristics of fall behavior in different scenarios.
[0074] In this embodiment, sensitive data (such as facial information) can be desensitized before training, or differential noise can be added during the parameter update stage.
[0075] S402: Encrypt the local parameters and upload them to the central server so that the central server aggregates the local parameters of the edge computing nodes of all distribution centers, and optimizes the local parameters of the edge computing nodes of all distribution centers using a federated averaging algorithm to obtain optimized parameters.
[0076] In this embodiment, homomorphic encryption (such as Paillier encryption) or secure multi-party computation (MPC) can be used to encrypt local parameter updates to ensure that the central server cannot reversely infer the original data.
[0077] Encrypting local parameters ensures data security. In federated learning, the data in each distribution center is sensitive and should not be disclosed. By encrypting local parameters and uploading them, even if the data is intercepted during transmission, attackers cannot obtain valuable information.
[0078] After the central server collects the encrypted local parameters uploaded by each edge computing node, it first decrypts them. It then uses a federated averaging algorithm to aggregate and optimize the local parameters of all nodes. This algorithm takes a weighted average of the local parameters based on the amount of training data or other weighting factors for each node. By combining the training results from each distribution center, it derives a set of optimized global parameters, known as the optimized parameters. This aggregation approach fully leverages data from multiple distribution centers, improving the model's generalization capabilities.
[0079] S403: Obtain optimization parameters sent by the central server, and use the optimization parameters to optimize the model after initial training to obtain a person fall recognition model.
[0080] In this embodiment, the edge computing node obtains optimized parameters from the central server and applies them to the initially trained model. By updating the model's parameters, the model incorporates the training experience of each distribution center, further optimizing its performance. Through multiple iterations of training, uploading, aggregation, and updating, the model continuously learns and adapts to the scenarios of different distribution centers, ultimately achieving a high-performance model for human fall recognition.
[0081] In this embodiment, edge computing nodes are used for local training, fully leveraging the uniqueness of the data from each distribution center, enabling the model to learn the characteristics of falls in different scenarios, thereby improving the model's adaptability and generalization capabilities. Furthermore, the encrypted upload of local parameters effectively protects the data security of each distribution center and addresses security concerns in data sharing. Furthermore, the central server uses a federated averaging algorithm to aggregate and optimize local parameters, integrating the training results of multiple distribution centers to obtain better global parameters, further improving the performance of the model. By continuously iteratively updating model parameters, the model can continuously learn and improve, resulting in an accurate and reliable model for identifying human falls, providing strong protection for the safety of personnel in distribution centers.
[0082] See also Figure 5 A fifth embodiment of a method for identifying a person falling in a distribution center according to an embodiment of the present invention includes:
[0083] S501. Regularly obtain real-time images of the distribution center, and input the real-time images into a pre-trained crowd density estimation model to obtain a density heat map output by the crowd density estimation model.
[0084] In this embodiment, the distribution of personnel in the distribution center is in dynamic change. Regular acquisition of real-time images (for example, every few minutes) can timely capture the real-time status of personnel. The pre-trained crowd density estimation model is trained based on a large amount of crowd image data and has the ability to identify the distribution of people in the image. After the real-time image is input into the pre-trained crowd density estimation model, the model will output a density heat map based on the pixel features and position information of the people in the image. In the heat map, the color of densely populated areas is darker, and the color of sparsely populated areas is lighter, thereby intuitively presenting the distribution of people in different areas of the distribution center.
[0085] S502: Integrate the density heat map to obtain a real-time total number of people estimate, calculate the real-time total number of people ratio based on the real-time total number of people estimate and the preset maximum safe carrying capacity of the distribution center, and calculate the entropy standard value of the density heat map.
[0086] In this embodiment, the density of each area of the density heat map can be accumulated through integration, so as to estimate the real-time total number of people.
[0087] In this embodiment, the real-time total number of people accounts for R N Expressed as:
[0088]
[0089] Where N curren Represents the real-time total number of people estimated, N max Indicates the maximum safe carrying capacity of the distribution center, which is a preset value.
[0090] In this embodiment, entropy is used in information theory to measure the uncertainty or disorder of data. In the context of a density heat map, the standard entropy value can reflect the uniformity of the distribution of people. Calculating the entropy value of a heat map can quantify the disorder of the distribution of people. If the distribution of people is relatively uniform, the entropy value is high; if people are concentrated in certain areas, the entropy value is low.
[0091] Entropy standard value R of density heat map S Expressed as:
[0092]
[0093] Where S represents the entropy value of the density heat map, S min and S max Represent the minimum and maximum historical entropy values respectively.
[0094] S503. Calculate the scene complexity score based on the proportion of the total number of people and the entropy standard value.
[0095] In this embodiment, the scene complexity score C is represented as:
[0096] C=α·R N +β·R S
[0097] Where α and β are weight coefficients, which must satisfy α + β = 1, such as α is 0.7 and β is 0.3.
[0098] In this example, the total headcount ratio and the entropy standard value describe the distribution of personnel in the distribution center from different perspectives. The total headcount ratio reflects the degree of crowding, while the entropy standard value reflects the uniformity of personnel distribution. Combining these two values to calculate the scenario complexity score provides a more comprehensive assessment of the complexity of the distribution center scenario.
[0099] In this embodiment, real-time images of the distribution center are regularly acquired and a density heat map is generated using a pre-trained model. On this basis, the real-time total number of people estimate, the real-time total number of people ratio, and the entropy standard value are calculated to quantitatively analyze the personnel distribution of the distribution center from two key dimensions: the number of people and the uniformity of distribution. These quantitative indicators are comprehensively calculated to obtain a scene complexity score, which can comprehensively and accurately evaluate the complexity of the distribution center scene.
[0100] See also Figure 6 A sixth embodiment of a method for identifying a person falling in a distribution center according to an embodiment of the present invention includes:
[0101] S601: Smoothing the scene complexity score by using an exponentially weighted moving average to obtain a smoothed scene complexity score.
[0102] In this embodiment, a confidence threshold to be updated is calculated based on the scene complexity score, and the confidence threshold of the person fall recognition model is adjusted according to the confidence threshold to be updated to obtain the adjusted person fall recognition model, specifically including: using exponentially weighted moving average to smooth the scene complexity score to obtain a smoothed scene complexity score; calculating the confidence threshold to be updated based on the maximum and minimum values of the confidence threshold of the person fall recognition model and the smoothed scene complexity score; adjusting the confidence threshold of the person fall recognition model according to the confidence threshold to be updated to obtain the adjusted person fall recognition model.
[0103] In this embodiment, the smooth scene complexity score is expressed as:
[0104]
[0105] Where, represents the smooth scene complexity score at the current time t, represents the smoothed scene complexity score at the previous time t-1, and λ is the smoothing factor (such as 0.2).
[0106] In this embodiment, the confidence threshold θ to be updated is expressed as:
[0107]
[0108] Where θ min Represents the minimum value of the confidence threshold of the person fall recognition model, θ max Indicates the maximum value of the confidence threshold of the person fall recognition model, represents the smooth scene complexity score at the current time t.
[0109] For example, if the original confidence threshold is 0.5, when the scene complexity is high, it will be raised to 0.7. Only when the model predicts that the probability of falling exceeds 0.7, it will be considered a fall; when the scene complexity is low, it will be lowered to 0.3. When the model predicts that the probability of falling exceeds 0.3, it will be considered a fall.
[0110] S602: Preprocess the real-time image, input the preprocessed real-time image into the adjusted person fall recognition model, and obtain the detection result output by the adjusted person fall recognition model.
[0111] In this example, the real-time images captured by the distribution center may contain various issues, such as noise, uneven lighting, and image blur. The preprocessing process aims to improve image quality and enhance the model detection performance. Preprocessing operations include denoising, which removes salt and pepper noise and Gaussian noise from the image to improve image clarity; grayscale transformation, which adjusts the image's brightness and contrast to enhance the individual's features; and image normalization, which maps the image's pixel values to a specific range to unify the data scale for easier model processing.
[0112] In this embodiment, the pre-processed real-time image is input into the adjusted person fall recognition model. The model extracts, analyzes, and judges the image features and outputs a prediction result. The model predicts the state of the person in the image based on the learned fall behavior characteristics and determines whether there is a fall behavior. The prediction result is usually presented in the form of a probability value. For example, if the predicted probability of falling is 0.8, it means that the model believes that there is a high possibility of falling behavior in the image; if the probability is 0.2, it means that the model believes that there is a high possibility of no falling behavior.
[0113] S603: When the detection result shows that a person has fallen, an alarm mechanism is triggered.
[0114] In this embodiment, the prediction results output by the model are evaluated. When the prediction results indicate that a person has fallen (i.e., the predicted probability exceeds the adjusted confidence threshold), the subsequent alarm process is initiated. For example, if the adjusted confidence threshold is 0.6, when the model predicts a fall probability of 0.7, the condition for triggering an alarm is met.
[0115] Once a person falls, an alarm is triggered immediately. This alarm can be activated in a variety of ways, such as by emitting a high-decibel audible alarm to attract the attention of nearby personnel; by flashing a light alarm to attract attention in noisy or dimly lit areas; and by sending a notification message to the manager's mobile phone, computer, or other terminal device, informing them of the location and time of the fall, so that rescue can be organized in a timely manner.
[0116] In this embodiment, the confidence threshold of the human fall recognition model is dynamically adjusted based on the complexity of the scene, which significantly improves the accuracy of the model in identifying fall behaviors in different environments; moreover, through real-time image preprocessing, the image quality of the input model is effectively improved, and the model's ability to detect fall behaviors is enhanced; in addition, once a fall behavior is detected, the alarm mechanism is quickly triggered, and relevant personnel can be notified in time to take rescue measures, thereby minimizing the damage caused by the fall accident and ensuring the life safety of the distribution center staff.
[0117] The above describes the method for identifying a person falling in a distribution center according to an embodiment of the present invention. The following describes the device according to an embodiment of the present invention. Figure 7 , the implementation of the device for identifying a person falling in a distribution center in an embodiment of the present invention includes:
[0118] Production module 701, used to collect personnel image data in different scenes of the distribution center and produce a local training set based on the personnel image data;
[0119] A construction module 702 is used to construct an object detection network based on lightweight RT-DETR, and introduce a multi-scale feature fusion mechanism and a self-attention mechanism into the object detection network to obtain a model to be trained;
[0120] A training module 703 is configured to introduce a federated learning framework based on the local training set and coordinate edge computing nodes of multiple distribution centers to perform distributed training on the to-be-trained model to obtain a person fall recognition model;
[0121] A calculation module 704 is configured to periodically acquire real-time images of the distribution center, input the real-time images into a pre-trained crowd density estimation model, obtain a density heat map output by the crowd density estimation model, and calculate a scene complexity score based on the density heat map.
[0122] The detection module 705 is used to adjust the confidence threshold of the personnel fall recognition model based on the scene complexity score, and input the real-time image into the adjusted personnel fall recognition model to detect whether there is a person fall in the distribution center. When the detection result shows that there is a person fall, the alarm mechanism is triggered.
[0123] In this embodiment, the production module 701 includes: a collection unit 7011, which is used to collect personnel image data in different scenarios of the distribution center, and preprocess the personnel image data to obtain preprocessed data; a labeling unit 7012, which is used to label the personnel posture and falling behavior in the preprocessed data to obtain sample data; a production unit 7013, which is used to dynamically enhance the sample data using a generative adversarial network and generate a local training set.
[0124] In this embodiment, the construction module 702 includes: a construction unit 7021, which is used to construct a target detection network based on the lightweight RT-DETR; an adding unit 7022, which is used to add a multi-scale feature fusion module to the target detection network to obtain an adjusted target detection network; an updating unit 7023, which is used to introduce a self-attention mechanism in the encoder or decoder of the adjusted target detection network to obtain a model to be trained.
[0125] In this embodiment, the training module 703 includes: a training unit 7031, which is used to use the local training set to perform initial training on the model to be trained through the edge computing node of the distribution center and generate local parameters; an encryption unit 7032, which is used to encrypt the local parameters and upload them to the central server so that the central server aggregates the local parameters of the edge computing nodes of all distribution centers, and uses the federal averaging algorithm to optimize the local parameters of the edge computing nodes of all distribution centers to obtain optimized parameters; an optimization unit 7033, which is used to obtain the optimized parameters sent by the central server, and use the optimized parameters to optimize the model after the initial training to obtain a personnel fall recognition model.
[0126] In this embodiment, the calculation module 704 includes: an acquisition unit 7041, which is used to periodically acquire real-time images of the distribution center and input the real-time images into a pre-trained crowd density estimation model to obtain a density heat map output by the crowd density estimation model; a first calculation unit 7042, which is used to perform an integral operation on the density heat map to obtain a real-time total number of people estimate, and calculate the real-time total number of people ratio based on the real-time total number of people estimate and the preset maximum safe carrying capacity of the distribution center, and calculate the entropy value standard value of the density heat map; a second calculation unit 7043, which is used to calculate a scene complexity score based on the total number of people ratio and the entropy value standard value.
[0127] In this embodiment, the detection module 705 includes: an adjustment unit 7051, which is used to calculate the confidence threshold to be updated based on the scene complexity score, and adjust the confidence threshold of the person fall recognition model according to the confidence threshold to be updated to obtain the adjusted person fall recognition model; a detection unit 7052, which is used to preprocess the real-time image, and input the preprocessed real-time image into the adjusted person fall recognition model, and obtain the detection result output by the adjusted person fall recognition model; an alarm unit 7053, which is used to trigger an alarm mechanism when the detection result is that a person falls.
[0128] In this embodiment, a lightweight RT-DETR target detection network is adopted, and a multi-scale feature fusion mechanism and a self-attention mechanism are introduced. While ensuring detection efficiency, the detection accuracy of people at different distances and postures and the feature extraction capability of complex scenes are significantly improved. Moreover, during the training process, a federated learning framework is introduced for distributed model training, which not only protects the data security of each distribution center, but also makes full use of multi-party data and computing resources, accelerates the training speed and improves model performance. In addition, a pre-trained crowd density estimation model is used to generate a scene complexity score, and the confidence threshold of the person fall recognition model is dynamically adjusted according to the score, so that the model can flexibly adjust the detection strategy according to different scenarios, effectively reduce misjudgment and improve detection sensitivity, and trigger the alarm mechanism in time when a fall behavior is detected, which can quickly respond to fall accidents and provide strong protection for the safety of personnel in the distribution center.
[0129] Figure 7 The structure of the device for identifying a person falling in a distribution center shown does not constitute a limitation on the device, and can implement the steps of the method for identifying a person falling in a distribution center provided by the above-mentioned method embodiments.
[0130] above Figure 7 The apparatus for identifying falls of persons in a distribution center in an embodiment of the present invention is described in detail from the perspective of modular functional entities. The apparatus for identifying falls of persons in a distribution center in an embodiment of the present invention is described in detail from the perspective of hardware processing.
[0131] Figure 8This is a structural diagram of a device for identifying falls in distribution centers provided by an embodiment of the present invention. The device 800 may have relatively large differences due to different configurations or performances, and may include one or more processors (central processing units, CPU) 810 (for example, one or more processors) and a memory 820, and one or more storage media 830 (for example, one or more massive storage devices) for storing applications 833 or data 832. Among them, the memory 820 and the storage medium 830 may be temporary storage or permanent storage. The program stored in the storage medium 830 may include one or more modules (not shown), and each module may include a series of instruction operations on the device 800. Furthermore, the processor 810 may be configured to communicate with the storage medium 830 to execute a series of instruction operations in the storage medium on the device 800.
[0132] The device 800 may also include one or more power supplies 840, one or more wired or wireless network interfaces 850, one or more input and output interfaces 860, and / or one or more operating systems 831, such as Windows Server, Mac OS X, Unix, Linux, FreeBSD, etc.
[0133] An embodiment of the present invention also provides a computer-readable storage medium, which may be a non-volatile computer-readable storage medium or a volatile computer-readable storage medium. Instructions are stored in the computer-readable storage medium. When the instructions are executed on a computer, the computer executes the steps of the method for identifying falls of people in a distribution center.
[0134] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the above-described systems, devices, and units can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0135] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM), random access memory (RAM), magnetic disk or optical disk, etc., various media that can store program code.
[0136] As described above, the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that the technical solutions described in the above embodiments can still be modified, or some of the technical features thereof can be replaced by equivalents. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for identifying a person falling in a distribution center, characterized in that: The method for identifying a person falling in a distribution center includes: Collect personnel image data from different scenarios in the distribution center and create a local training set based on the personnel image data; Construct an object detection network based on lightweight RT-DETR, and introduce a multi-scale feature fusion mechanism and a self-attention mechanism into the object detection network to obtain a model to be trained; Based on the local training set, a federated learning framework is introduced to jointly perform distributed training on the model to be trained by edge computing nodes of multiple distribution centers to obtain a person fall recognition model; Regularly obtain real-time images of the distribution center, input the real-time images into a pre-trained crowd density estimation model, obtain the density heat map output by the crowd density estimation model, and calculate the scene complexity score based on the density heat map; Based on the scene complexity score, the confidence threshold of the personnel fall recognition model is adjusted, and the real-time image is input into the adjusted personnel fall recognition model to detect whether there is a person falling in the distribution center. When the detection result shows that there is a person falling, the alarm mechanism is triggered.
2. The method for identifying a person falling in a distribution center according to claim 1, characterized in that: The collecting of personnel image data in different scenes of the distribution center and the preparation of a local training set based on the personnel image data include: Collecting personnel image data in different scenarios of the distribution center and preprocessing the personnel image data to obtain preprocessed data; Annotating the posture and falling behavior of the person in the preprocessed data to obtain sample data; A generative adversarial network is used to dynamically enhance the sample data and generate a local training set.
3. The method for identifying a person falling in a distribution center according to claim 1, characterized in that: The object detection network based on the lightweight RT-DETR is constructed, and a multi-scale feature fusion mechanism and a self-attention mechanism are introduced into the object detection network to obtain a model to be trained, including: Build a target detection network based on lightweight RT-DETR; Adding a multi-scale feature fusion module to the target detection network to obtain an adjusted target detection network; In the encoder or decoder of the adjusted object detection network, a self-attention mechanism is introduced to obtain the model to be trained.
4. The method for identifying a person falling in a distribution center according to claim 1, characterized in that: Based on the local training set, a federated learning framework is introduced to jointly perform distributed training on the to-be-trained model with edge computing nodes of multiple distribution centers to obtain a person fall recognition model, including: Using the local training set, the model to be trained is initially trained by the edge computing node of the distribution center, and local parameters are generated; Encrypting the local parameters and uploading them to a central server so that the central server aggregates the local parameters of the edge computing nodes of all distribution centers, and optimizing the local parameters of the edge computing nodes of all distribution centers using a federated averaging algorithm to obtain optimized parameters; Obtain the optimization parameters sent by the central server, and use the optimization parameters to optimize the model after the initial training to obtain a person fall recognition model.
5. The method for identifying a person falling in a distribution center according to claim 1, characterized in that: The method includes regularly acquiring real-time images of the distribution center, inputting the real-time images into a pre-trained crowd density estimation model, obtaining a density heat map output by the crowd density estimation model, and calculating a scene complexity score based on the density heat map, including: Regularly obtain real-time images of the distribution center and input the real-time images into a pre-trained crowd density estimation model to obtain a density heat map output by the crowd density estimation model; Performing an integration operation on the density heat map to obtain a real-time total number of people estimate, and calculating the real-time total number of people ratio based on the real-time total number of people estimate and the preset maximum safe carrying capacity of the distribution center, and calculating the entropy standard value of the density heat map; The scene complexity score is calculated based on the total number of people and the entropy standard value.
6. The method for identifying a person falling in a distribution center according to claim 5, characterized in that: The confidence threshold of the personnel fall recognition model is adjusted based on the scene complexity score, and the real-time image is input into the adjusted personnel fall recognition model to detect whether there is a personnel fall in the distribution center. When the detection result shows that there is a personnel fall, an alarm mechanism is triggered, including: Calculating a confidence threshold to be updated based on the scene complexity score, and adjusting the confidence threshold of the person fall recognition model according to the confidence threshold to be updated to obtain an adjusted person fall recognition model; Preprocessing the real-time image, inputting the preprocessed real-time image into the adjusted person fall recognition model, and obtaining the detection result output by the adjusted person fall recognition model; When the detection result indicates that a person has fallen, an alarm mechanism is triggered.
7. The method for identifying a person falling in a distribution center according to claim 6, characterized in that: The step of calculating a confidence threshold to be updated based on the scene complexity score, and adjusting the confidence threshold of the person fall recognition model according to the confidence threshold to be updated to obtain an adjusted person fall recognition model includes: The scene complexity score is smoothed using exponentially weighted moving average to obtain a smoothed scene complexity score; Calculating a confidence threshold to be updated based on the maximum and minimum values of the confidence threshold of the personnel fall recognition model and the smooth scene complexity score; The confidence threshold of the person fall recognition model is adjusted according to the confidence threshold to be updated to obtain an adjusted person fall recognition model.
8. A device for identifying people falling in a distribution center, characterized in that: include: The production module is used to collect personnel image data in different scenarios of the distribution center and create a local training set based on the personnel image data; A construction module is used to construct an object detection network based on lightweight RT-DETR, and introduce a multi-scale feature fusion mechanism and a self-attention mechanism into the object detection network to obtain a model to be trained; A training module is configured to introduce a federated learning framework based on the local training set and jointly perform distributed training on the to-be-trained model with edge computing nodes of multiple distribution centers to obtain a person fall recognition model; A calculation module is used to regularly obtain real-time images of the distribution center, input the real-time images into a pre-trained crowd density estimation model, obtain a density heat map output by the crowd density estimation model, and calculate a scene complexity score based on the density heat map; The detection module is used to adjust the confidence threshold of the personnel fall recognition model based on the scene complexity score, and input the real-time image into the adjusted personnel fall recognition model to detect whether there is a person fall in the distribution center. When the detection result shows that there is a person fall, an alarm mechanism is triggered.
9. A device for identifying people falling in a distribution center, characterized in that: comprising a memory and at least one processor, wherein the memory has computer-readable instructions stored therein; The at least one processor calls the computer-readable instructions in the memory to execute the various steps of the method for identifying a person falling in a distribution center as recited in any one of claims 1 to 7.
10. A computer-readable storage medium having computer-readable instructions stored thereon, characterized in that: When the computer-readable instructions are executed by a processor, the steps of the method for identifying a person falling in a distribution center are implemented.