Child schoolbag safety positioning method and system based on deep learning

Through the improved YOLOv5 target detection network and multi-sensor fusion technology, the problem of inaccurate positioning of children's backpacks was solved, accurate detection and active warning were achieved in complex environments, and the effectiveness of child safety monitoring was improved.

CN120689406APending Publication Date: 2025-09-23ZHEJIANG CAARANY BUSINESS LEISURE PRODS
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510682076.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-26
Publication Date
2025-09-23

AI Technical Summary

Technical Problem

The accuracy of existing child positioning and monitoring technology decreases when the indoor signal is weak or blocked. Traditional target detection algorithms have difficulty accurately identifying children's backpacks, resulting in inaccurate positioning. In addition, the safe area determination method lacks flexibility and cannot provide timely warnings of potential risks.

Method used

An improved YOLOv5 target detection network combined with multi-sensor fusion technology is used. By annotating and pre-processing children's schoolbag images, a high-quality training dataset is constructed. A spatial attention mechanism and an additional small target detection head are added. The Kalman filter algorithm is used to fuse inertial measurement unit data, and a multi-level warning is generated by combining the ray method safety area determination.

Benefits of technology

It improves the accuracy of children's backpack detection and tracking stability, realizes the transition from passive monitoring to active early warning, enhances the effectiveness of child safety protection, and can accurately identify backpacks of different sizes, colors and shapes in complex environments, and identify potential risks before danger occurs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120689406A_ABST
    Figure CN120689406A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of data processing, and discloses a children schoolbag safety positioning method and system based on deep learning. The method comprises the following steps: marking and enhancing a children schoolbag image, and training an improved YOLOv5 model; then combining a visual detection result and inertial measurement data of the model, and realizing multi-sensor fusion positioning through a Kalman filtering algorithm; whether the schoolbag is in a safe area or not is judged through a ray method, and safety early warning information of the corresponding level is generated in combination with the motion parameters. According to the invention, through fusion of the improved YOLOv5 network and multiple sensors, the detection precision and tracking stability of the child schoolbag small target are improved. The multi-level safety area and ray method judgment are combined with motion trend analysis, conversion from passive monitoring to active early warning is achieved, and the effectiveness of child safety guarantee is enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of data processing technology, and in particular to a method and system for safely locating children's school bags based on deep learning. Background Art

[0002] With growing concerns about child safety, the demand for child activity monitoring and location-based technologies continues to grow. Currently, common child location-based monitoring technologies include GPS-based location tracking, RFID-based campus card recognition systems, and video surveillance systems. These technologies utilize satellite positioning, electronic tag recognition, and image analysis to enable real-time monitoring of children's locations. Deep learning-based computer vision technology has made significant progress in object detection. Object detection algorithms such as YOLOv5 excel in both speed and accuracy, providing new technical solutions for child safety monitoring.

[0003] However, existing technologies have significant shortcomings. Single-use GPS positioning suffers from poor accuracy when indoor signals are weak or obstructed. RFID systems can only identify objects at fixed detection points and lack continuous tracking capabilities. Video-based surveillance systems are susceptible to occlusion, lighting fluctuations, and other factors in complex environments, resulting in low accuracy for small objects such as children's backpacks. Traditional object detection algorithms struggle to accurately identify small objects like children's backpacks, especially in crowded campus environments, leading to inaccurate positioning. Furthermore, the limitations of single sensors create blind spots, reducing security.

[0004] At the same time, traditional safety zone determination methods lack flexibility, making it difficult to accurately delineate safety boundaries and dynamically predict risks. Existing systems often rely on fixed thresholds, failing to incorporate target motion trends for early warning. This results in passive and delayed security monitoring, making it difficult to detect and prevent potential safety risks in a timely manner. Summary of the Invention

[0005] This application provides a deep learning-based method and system for safely locating children's backpacks. This system improves the detection accuracy and tracking stability of small targets in children's backpacks through an improved YOLOv5 network and multi-sensor fusion. Multi-layered safety zones and ray-guided determination, combined with motion trend analysis, enable a shift from passive monitoring to active early warning, enhancing the effectiveness of child safety assurance.

[0006] In the first aspect, the present application provides a children's backpack safety positioning method based on deep learning, and the children's backpack safety positioning method based on deep learning includes: labeling and enhancing preprocessing of the collected children's backpack images to obtain a training data set and a verification data set; inputting the training data set into an improved YOLOv5 target detection network for training, the improved YOLOv5 target detection network includes a spatial attention mechanism and an additional small target detection head for small target detection, and using the verification data set to evaluate and optimize the training process to obtain a children's backpack detection model; obtaining a children's backpack detection model; based on the children's backpack detection model and inertial measurement unit data, multi-sensor data is fused through a Kalman filtering algorithm to obtain the real-time position coordinates, movement speed and movement direction of the children's backpack; performing ray method safety area determination processing on the real-time position coordinates of the children's backpack, combining the movement speed and movement direction, and generating multi-level safety warning information according to a preset safety area boundary distance threshold and departure time threshold.

[0007] Optionally, the labeling and enhancement preprocessing of the collected children's schoolbag images to obtain a training dataset and a verification dataset includes: Images of children's schoolbags are acquired through campus monitoring equipment and scene acquisition devices, including scenes of indoor classrooms, corridors, playgrounds, and outdoor public areas, to obtain raw image data; Performing bounding box annotation on the children's backpack in the original image data, recording the backpack position coordinates and backpack type label, and obtaining backpack annotation data; Normalizing the image size of the backpack annotation data, maintaining the aspect ratio while adjusting the resolution, to obtain a standardized image set; performing illumination adjustment and color balancing processing on the standardized image set to obtain an illumination-enhanced image; Performing multi-view synthesis processing on the illumination-enhanced image to generate a composite scene and obtain a scene-enriched image; The scene enriched images are divided according to a training and verification ratio to obtain a training data set and a verification data set.

[0008] Optionally, the training dataset is input into an improved YOLOv5 object detection network for training, wherein the improved YOLOv5 object detection network includes a spatial attention mechanism for small object detection and an additional small object detection head, and the training process is evaluated and optimized using the validation dataset to obtain a children's schoolbag detection model; obtaining the children's schoolbag detection model includes: Build a YOLOv5 infrastructure with a CSPDarknet53 backbone network, which consists of alternating convolutional layers and residual blocks to extract features from the input image and generate multi-scale feature maps. Adding a spatial attention module to the feature extraction layer of the backbone network, performing channel weight calculation and spatial position weighting on the feature map to obtain an attention-enhanced feature map; Constructing a feature fusion neck structure including a feature pyramid network FPN and a path aggregation network PAN, performing upsampling and downsampling operations on the attention enhancement feature map to obtain a multi-level fusion feature; On the basis of the multi-level fusion features, an additional detection head dedicated to small target detection is added, and the downsampling multiple of the feature map is reduced to obtain a small target enhanced feature map; Performing anchor box cluster analysis on the labeled data in the training dataset, calculating the optimal anchor box size configuration, and assigning anchor box parameters of different scales to the three detection layers to obtain anchor box settings for the size distribution of children's school bags; The network is iteratively trained based on the anchor frame setting and the CIOU loss function. During the training process, a cosine annealing learning rate strategy is used and the model performance is regularly evaluated on the validation dataset. When the performance indicator on the validation dataset no longer improves for a preset number of consecutive times, the optimal model weight is saved to obtain a children's backpack detection model.

[0009] Optionally, the constructing includes a feature fusion neck structure comprising a feature pyramid network FPN and a path aggregation network PAN, performing upsampling and downsampling operations on the attention enhancement feature map to obtain a multi-level fusion feature, including: Extracting three attention-enhanced feature maps of different scales from the backbone network, including a shallow feature map, a mid-level feature map, and a deep feature map, to obtain a multi-scale feature set; Performing convolution processing on the deep feature map to obtain semantic compression features; The semantic compression feature is enlarged to the same spatial size as the middle-layer feature map through an upsampling operation, and feature channel splicing is performed with the middle-layer feature map to obtain a first fused feature map; Performing convolution processing on the first fused feature map, and enlarging it to the same spatial size as the shallow feature map through an upsampling operation, performing feature channel splicing with the shallow feature map, completing feature pyramid network processing, and obtaining a second fused feature map; Based on the second fused feature map, extract key information and reduce the resolution through convolution and maximum pooling operations, and perform feature channel splicing with the first fused feature map to obtain a third fused feature map; Convolution and maximum pooling operations are performed on the third fusion feature map, and feature channel splicing is performed with the semantic compression feature to complete path aggregation network processing to obtain multi-level fusion features.

[0010] Optionally, the method further comprises adding an additional detection head dedicated to small target detection on the basis of the multi-level fusion features, reducing the downsampling multiple of the feature map, and obtaining a small target enhanced feature map, including: Selecting a shallow feature map from the multi-level fusion features, performing channel compression convolution processing on the shallow feature map, retaining spatial resolution information, and obtaining small target basic features; Performing a depth-separable convolution operation on the basic features of the small target to obtain a deep feature representation; Inputting the deep feature representation into the attention guidance module, calculating the spatial position importance weight map, and enhancing the potential children's backpack area in the feature map to obtain regional enhancement features; Performing an upsampling operation on the region enhancement feature through bilinear interpolation to enlarge the feature map resolution to one-fourth of the original input image to obtain a high-resolution feature map; The high-resolution feature map is fused with the shallow features transmitted through the jump connection, and multi-level information is integrated by element-by-element addition to obtain multi-scale fused features; A detection head structure consisting of a series of convolutional layers and batch normalization layers is applied to the multi-scale fusion features, and three branches are output: a bounding box coordinate prediction branch, an object confidence prediction branch, and a category prediction branch, to generate a small object enhanced feature map suitable for small-sized children's schoolbag detection.

[0011] Optionally, the multi-sensor data is fused and processed based on the children's backpack detection model and the inertial measurement unit data through a Kalman filter algorithm to obtain the real-time position coordinates, movement speed and movement direction of the children's backpack, including: Detecting the position and bounding box of the child's school bag in each frame of the image sequence using the child's school bag detection model to obtain visual detection data; Acceleration and angular velocity data are collected from the backpack's built-in inertial measurement unit, and noise filtering and coordinate conversion are performed to obtain motion sensor data; Establish a Kalman filter state space model, define the state vector including position and velocity and its state transfer matrix, and obtain the prediction model; Performing state prediction calculation on the state vector at the previous moment, and using the prediction model to infer the state at the current moment to obtain a predicted state; Combining the visual detection data and the motion sensor data, calculating the deviation between the measured value and the predicted value, updating the predicted state, and obtaining a fusion state; The real-time position coordinates, movement speed and movement direction of the children's backpack are extracted from the fusion state.

[0012] Optionally, the ray method safety zone determination processing is performed on the real-time position coordinates of the children's backpack, and multi-level safety warning information is generated according to a preset safety zone boundary distance threshold and a departure time threshold in combination with the movement speed and movement direction, including: Define the multi-level security area boundaries of the core security area, general security area and extended security area. Each security area is described by a set of polygonal coordinate points to obtain the security area boundary data; Applying a ray method algorithm to the real-time location coordinates of the child's backpack, emitting a ray from the real-time location coordinates in any direction and counting the number of intersections with the safety zone boundary, determining whether the backpack is within the safety zone based on the parity of the number of intersections, and obtaining an inside-outside determination result; Based on the determination result of whether the backpack is inside or outside the area and the real-time location coordinates of the child's backpack, the distance between the backpack and the nearest safe area boundary is calculated. In combination with the movement speed and direction, the possibility of the backpack leaving or entering the safe area is predicted to obtain a safety risk level. Based on the security risk level, combined with the safety zone boundary distance threshold and the departure time threshold, an early warning level is determined, including attention level, warning level and emergency level, to obtain a graded early warning sign; According to the graded warning signs, a warning content including the backpack's location information, the time of leaving the safe area and the predicted trajectory is generated to generate a complete warning message; The complete warning message is sent to the administrator or guardian through multiple notification channels, and the warning event information is recorded at the same time to generate multi-level security warning information.

[0013] In a second aspect, the present application provides a children's schoolbag safety positioning system based on deep learning, the children's schoolbag safety positioning system based on deep learning comprising: A processing module is used to perform annotation processing and enhancement preprocessing on the collected children's schoolbag images to obtain a training data set and a verification data set; A training module is configured to input the training dataset into an improved YOLOv5 object detection network for training, wherein the improved YOLOv5 object detection network includes a spatial attention mechanism and an additional small object detection head for small object detection, and evaluate and optimize the training process using the validation dataset to obtain a children's schoolbag detection model; and obtain a children's schoolbag detection model; A fusion module is used to fuse the multi-sensor data using a Kalman filter algorithm based on the children's backpack detection model and the inertial measurement unit data to obtain the real-time position coordinates, movement speed and movement direction of the children's backpack; A generation module is used to perform ray method safety area determination processing on the real-time position coordinates of the children's backpack, combine the movement speed and movement direction, and generate multi-level safety warning information according to the preset safety area boundary distance threshold and departure time threshold.

[0014] In a third aspect, a children's backpack safety positioning device based on deep learning is provided, comprising: a memory and at least one processor, wherein the memory stores instructions; the at least one processor calls the instructions in the memory so that the children's backpack safety positioning device based on deep learning executes the above-mentioned children's backpack safety positioning method based on deep learning.

[0015] In a fourth aspect, a computer-readable storage medium is provided, wherein instructions are stored in the computer-readable storage medium. When the computer-readable storage medium is run on a computer, the computer executes the above-mentioned method for safely locating children's backpacks based on deep learning.

[0016] In the technical solution provided by this application, by professionally labeling and enhancing the pre-processing of the collected images of children's backpacks, high-quality training datasets and verification datasets are created, providing sufficient learning materials for the deep learning model, thereby significantly improving the recognition accuracy of children's backpacks. The improved YOLOv5 target detection network is specially optimized for the characteristics of small targets such as children's backpacks. The addition of a spatial attention mechanism enables the model to focus on key areas in the image and effectively extract the visual features of children's backpacks; the design of an additional small target detection head reduces the downsampling multiple of the feature map, retaining more spatial detail information, and significantly improving the network's detection ability for small-sized backpacks. The introduction of the verification dataset realizes real-time evaluation and optimization of the training process, ensuring that the model parameters are adjusted in a direction that is more conducive to backpack detection. The resulting children's backpack detection model can accurately identify backpacks of different sizes, colors, and shapes in complex backgrounds.

[0017] The multi-sensor fusion based on the children's backpack detection model and inertial measurement unit data is an important innovation of the present invention. The Kalman filter algorithm effectively integrates the visual detection results and motion sensor data, overcomes the limitations of a single sensor in situations such as occlusion and lighting changes, and achieves continuous and stable tracking of the position of the children's backpack. Even when the sensor data is temporarily lost or interfered with, it can still maintain accurate positioning. The acquired real-time position coordinates, movement speed and movement direction provide comprehensive dynamic information for safety monitoring, far exceeding the capabilities of traditional static position monitoring. The ray method safety area determination processing converts abstract safety rules into specific spatial determination algorithms, and combines the motion parameters of the children's backpack for predictive analysis, enabling the system to identify potential risks before danger occurs. The preset safety area boundary distance threshold and departure time threshold provide quantitative standards for judgment, and the generation of multi-level safety warning information realizes the refinement and differentiated processing of warnings.

[0018] This application improves the algorithm for small target detection, making the deep learning model more suitable for the specific identification object of children's backpacks. At the same time, the multi-sensor fusion algorithm overcomes the limitations of a single sensor, and the safe area determination algorithm makes the early warning mechanism more forward-looking. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0020] Figure 1 This is a schematic diagram of an embodiment of a method for safely locating a children's schoolbag based on deep learning in an embodiment of the present application; Figure 2 This is a schematic diagram of an embodiment of a children's schoolbag safety positioning system based on deep learning in an embodiment of the present application; Figure 3 This is a schematic block diagram of the structure of a children's schoolbag safety positioning device based on deep learning in an embodiment of the present invention. DETAILED DESCRIPTION

[0021] The embodiments of the present application provide a method and system for safely locating children's school bags based on deep learning. The terms "first", "second", "third", "fourth", etc. (if any) in the specification and claims of this application and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments described herein can be implemented in an order other than that illustrated or described herein. In addition, the terms "including" or "having" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units that are clearly listed, but may include other steps or units that are not clearly listed or that are inherent to these processes, methods, products or devices.

[0022] For ease of understanding, the specific process of the embodiment of the present application is described below. Figure 1 In the embodiments of the present application, an embodiment of a method for safely locating a children's schoolbag based on deep learning includes: Step S101: performing annotation processing and enhancement preprocessing on the collected children's schoolbag images to obtain a training data set and a verification data set; Step S102: Input the training dataset into the improved YOLOv5 object detection network for training. The improved YOLOv5 object detection network includes a spatial attention mechanism and an additional small object detection head for small object detection. The training process is evaluated and optimized using the validation dataset to obtain a children's schoolbag detection model. Step S103: Based on the children's backpack detection model and the inertial measurement unit data, the multi-sensor data is fused and processed by the Kalman filter algorithm to obtain the real-time position coordinates, movement speed and movement direction of the children's backpack; Step S104: Perform ray method safety zone determination processing on the real-time position coordinates of the child's backpack, combine the movement speed and movement direction, and generate multi-level safety warning information according to the preset safety zone boundary distance threshold and departure time threshold.

[0023] It is understandable that the execution subject of this application can be a children's schoolbag safety positioning system based on deep learning, or a terminal or a server, which is not limited here. The embodiment of this application is described by taking the server as the execution subject as an example.

[0024] Specifically, the collected images of children's backpacks are annotated and pre-processed for enhancement to produce training and validation datasets. Images of children's backpacks are acquired using campus surveillance equipment and scene acquisition devices, encompassing a variety of scenes such as indoor classrooms, corridors, playgrounds, and outdoor public areas. These raw image data are annotated with manual bounding boxes, recording the location coordinates and type label of each backpack. After annotation, the images are size-normalized to maintain their aspect ratio and adjust to a uniform resolution, forming a standardized image set. The standardized image set is then subjected to lighting adjustment and color balancing to address image differences under varying lighting conditions, resulting in lighting-enhanced images. To increase data diversity, the lighting-enhanced images are synthesized from multiple perspectives, and a composite scene containing multiple backpack instances is generated using stitching technology to form scene-enriched images. Finally, the scene-enriched images are divided proportionally into training and validation datasets.

[0025] The training dataset was fed into a modified YOLOv5 object detection network for training. This modified network first constructs a basic architecture consisting of a CSPDarknet53 backbone network. It extracts features from the input image through alternating convolutional layers and residual blocks, generating multi-scale feature maps. A spatial attention module is added to the backbone network's feature extraction layer to calculate channel weights and spatial position weights on the feature maps, focusing the model on the children's backpacks and generating attention-enhanced feature maps. Next, a fusion neck structure is constructed, combining a feature pyramid network (FPN) and a path aggregation network (PAN). Features from different levels are fused through upsampling and downsampling operations to generate multi-level fused features. An additional detection head dedicated to small object detection is added to the multi-level fused features. The downsampling factor of the feature maps is reduced to extract features of small backpacks and generate enhanced feature maps for small objects. Anchor box clustering analysis is then performed on the annotated data in the training dataset to calculate the anchor box parameter settings that best fit the size distribution of children's backpacks. Finally, network training is performed based on anchor box settings and the CIOU loss function. A cosine annealing learning rate strategy is used during training. Model performance is regularly evaluated using a validation dataset. The optimal model weights are retained when performance indicators no longer improve. This results in a deep learning model specifically designed for detecting children's backpacks. Based on the children's backpack detection model and inertial measurement unit data, multi-sensor data fusion is performed using a Kalman filter algorithm. First, the children's backpack detection model processes image sequences captured by the camera, detecting the backpack's position and bounding box in each frame. Simultaneously, acceleration and angular velocity data are collected from the backpack's built-in inertial measurement unit. After noise filtering and coordinate transformation, the backpack's motion state parameters are obtained. A Kalman filter state-space model is then established, using the backpack's position and velocity as state variables. A state transition matrix is ​​defined to describe the motion. State prediction calculations are performed on the previous state vector to infer the backpack's possible position and motion state at the current moment. Next, the deviation between the actual measured and predicted values ​​is calculated by combining visual inspection data with motion sensor data, and the state prediction value is updated. Finally, the real-time position coordinates, velocity, and direction of the backpack are extracted from the fused state data.

[0026] The real-time location coordinates of a child's backpack are used to determine the safety zone. First, a multi-level safety zone boundary is defined, including a core safety zone, a general safety zone, and an extended safety zone. Each zone is described by a set of polygonal coordinate points. A ray method algorithm is applied to the backpack's real-time location coordinates. A ray is emitted from these coordinates in any direction and the number of intersections with the safety zone boundary is calculated. If the number of intersections is odd, the backpack is within the zone; if it is even, it is outside. Based on this determination result and the backpack's real-time location coordinates, the distance from the backpack to the nearest safety zone boundary is calculated. The probability of the backpack leaving or entering the safety zone is predicted based on the speed and direction of movement. Based on this calculation result and preset safety zone boundary distance and exit time thresholds, an alert level is determined, including caution, warning, and emergency. Based on the alert level, an alert is generated containing the backpack's location information, the time it left the safety zone, and the predicted trajectory. This alert is sent to the administrator or guardian through various notification channels, and the alert event information is recorded, forming a multi-level safety alert message.

[0027] In a specific embodiment, the process of executing step S101 may specifically include the following steps: Images of children's schoolbags are acquired through campus monitoring equipment and scene acquisition devices, including scenes of indoor classrooms, corridors, playgrounds, and outdoor public areas, to obtain raw image data; Perform bounding box annotation on the children's backpack in the original image data, record the backpack's location coordinates and backpack type label, and obtain the backpack annotation data; Normalize the image size of the backpack annotation data, maintain the aspect ratio while adjusting the resolution to obtain a standardized image set; Performing lighting adjustment and color balancing on the standardized image set to obtain lighting-enhanced images; Perform multi-view synthesis processing on the illumination-enhanced image to generate a composite scene and obtain a scene-enriched image; The scene enriched images are divided into training and validation ratios to obtain training data sets and validation data sets.

[0028] Specifically, images of children's backpacks are acquired through campus surveillance equipment and scene acquisition devices, including images in various scenes such as classrooms, corridors, and playgrounds, to form raw image data. These images are then annotated with bounding boxes, marking the location coordinates and type label (such as backpack, shoulder bag, etc.) of each backpack, and recorded as annotated data in XML or JSON format. The annotated images are then size-normalized, and the original images of different sizes are uniformly adjusted to the same resolution (such as 640×640 pixels). Pixel values ​​are calculated using an interpolation algorithm, while the original aspect ratio is maintained through padding to avoid distortion of the backpack shape, resulting in a standardized image set. The standardized images are then subjected to lighting adjustment and color balance processing. Histogram equalization technology is used to enhance image contrast, and color cast problems are corrected in the HSV color space to make the image color distribution more balanced, resulting in a lighting-enhanced image. The illuminated images were synthesized from multiple perspectives using the Mosaic data augmentation method. Four different images were stitched together into a single image at a preset ratio. Affine transformations, such as random rotations, scaling, and translations, were also applied to simulate the changes in the backpack's appearance from different perspectives, generating scene-enriched images. Finally, these images were divided into a training dataset and a validation dataset at a ratio of approximately 8:2 to ensure a similar distribution of backpacks and scenes in both datasets.

[0029] In a specific embodiment, the process of executing step S102 may specifically include the following steps: Build a YOLOv5 infrastructure with a CSPDarknet53 backbone network, which consists of alternating convolutional layers and residual blocks to extract features from the input image and generate multi-scale feature maps. A spatial attention module is added to the feature extraction layer of the backbone network to calculate the channel weights and spatial position weights of the feature map to obtain an attention-enhanced feature map; Construct a feature fusion neck structure consisting of a feature pyramid network (FPN) and a path aggregation network (PAN), perform upsampling and downsampling operations on the attention enhancement feature map, and obtain multi-level fusion features; Specifically, three attention-enhanced feature maps of different scales are extracted from the backbone network, including shallow feature maps, middle feature maps and deep feature maps, to obtain a multi-scale feature set; the deep feature map is convolved to obtain semantic compression features; the semantic compression features are enlarged to the same spatial size as the middle feature map through an upsampling operation, and feature channels are spliced ​​with the middle feature map to obtain a first fused feature map; the first fused feature map is convolved and enlarged to the same spatial size as the shallow feature map through an upsampling operation, and feature channels are spliced ​​with the shallow feature map to complete feature pyramid network processing to obtain a second fused feature map; based on the second fused feature map, key information is extracted and the resolution is reduced through convolution and maximum pooling operations, and feature channels are spliced ​​with the first fused feature map to obtain a third fused feature map; the third fused feature map is convolved and maximum pooled, and feature channels are spliced ​​with the semantic compression features to complete path aggregation network processing to obtain multi-level fused features.

[0030] On the basis of multi-level fusion features, an additional detection head dedicated to small target detection is added, the downsampling multiple of the feature map is reduced, and an enhanced feature map of small targets is obtained; Specifically, a shallow feature map is selected from the multi-level fusion features, and the shallow feature map is subjected to channel compression convolution processing to retain the spatial resolution information to obtain the basic features of the small target; a depth-wise separable convolution operation is performed on the basic features of the small target to obtain a deep feature representation; the deep feature representation is input into the attention guidance module, the spatial position importance weight map is calculated, and the potential children's backpack area in the feature map is enhanced to obtain the regional enhancement feature; the regional enhancement feature is upsampled by bilinear interpolation to enlarge the feature map resolution to one-quarter of the original input image to obtain a high-resolution feature map; the high-resolution feature map is fused with the shallow features transmitted through the jump connection, and the multi-level information is integrated by element-by-element addition to obtain the multi-scale fusion feature; a detection head structure consisting of a series of convolutional layers and batch normalization layers is applied to the multi-scale fusion feature, and three branches are output: a bounding box coordinate prediction branch, a target confidence prediction branch, and a category prediction branch to generate a small target enhanced feature map suitable for small-sized children's backpack detection.

[0031] Anchor box clustering analysis is performed on the labeled data in the training dataset to calculate the optimal anchor box size configuration. Anchor box parameters of different scales are assigned to the three detection layers to obtain anchor box settings tailored to the size distribution of children's backpacks. The network is iteratively trained based on the anchor box setting and CIOU loss function. The cosine annealing learning rate strategy is used during the training process, and the model performance is regularly evaluated on the validation dataset. When the performance indicator on the validation dataset no longer improves for a preset number of consecutive times, the optimal model weights are saved to obtain the children's backpack detection model.

[0032] Specifically, a YOLOv5 infrastructure was constructed with the CSPDarknet53 backbone network. CSPDarknet53 is a feature extraction network designed to reduce computational redundancy through a cross-stage partial network (CSP) design. This backbone network consists of alternating convolutional layers and residual blocks. The convolutional layers extract image features, while the residual blocks use skip connections to address the vanishing gradient problem in deep networks. The network takes a 640×640 pixel image as input. After the initial convolutional layer, it passes through five CSP modules, each containing multiple residual units. This progressively reduces the spatial size of the feature map and increases the number of channels, resulting in three feature maps of different sizes: 80×80×256, 40×40×512, and 20×20×1024, corresponding to shallow, mid-level, and deep features. A spatial attention module was added to the backbone network's feature extraction layer to enhance detection of small objects such as children's backpacks. The spatial attention module first calculates channel weights on the feature map. Using global average pooling and maximum pooling, it generates two vectors describing channel importance. These two vectors are concatenated and passed through a 1×1 convolution to generate the channel weight coefficients. Spatial position weights are also calculated. Average pooling and maximum pooling are performed on the feature map in the spatial dimension. After concatenation, a 7×7 convolution is performed to generate a spatial attention map. Finally, the channel weights and spatial weights are multiplied and applied to the original feature map to generate an attention-enhanced feature map. This feature map retains the original information while being more responsive to the characteristics of the area where the backpack is located.

[0033] A feature fusion neck structure consisting of a feature pyramid network (FPN) and a path aggregation network (PAN) is constructed to perform upsampling and downsampling operations on the attention-enhanced feature maps. First, three attention-enhanced feature maps of different scales are extracted from the backbone network: a deep feature map of 20×20×1024, a mid-level feature map of 40×40×512, and a shallow feature map of 80×80×256, forming a multi-scale feature set. The deep feature map is subjected to channel compression using 1×1 convolution, reducing the number of channels from 1024 to 512, thus obtaining semantically compressed features. Using the FPN top-down pathway, the semantically compressed features are upsampled by a factor of 2 to 40×40 and concatenated with the mid-level feature map along the channel dimension, resulting in a first fused feature map of 40×40×1024. A 1×1 convolution is applied to this fused feature map to further compress the channels to 512. It is then upsampled by a factor of 2 to 80×80 and concatenated with the shallow feature map to form a second fused feature map of 80×80×768, completing FPN processing. Next, a PAN bottom-up pathway is implemented to extract key information from the second fused feature map using 3×3 convolution and max pooling with a stride of 2. The resolution is then reduced to 40×40 and concatenated with the first fused feature map along the channel dimension to produce a third fused feature map of 40×40×1536. The third fused feature map is again convolved and max pooled with a stride of 2 to reduce its size to 20×20. This is then concatenated with the semantically compressed features to produce a feature map of 20×20×1536. In this way, high-level semantic information is transmitted through the top-down path of FPN, and the positioning accuracy is enhanced through the bottom-up path of PAN, forming multi-level fusion features of three sizes: 80×80×256, 40×40×512 and 20×20×1024, which have both positioning accuracy and semantic information.

[0034] Based on the multi-level fusion features, an additional detection head dedicated to small object detection is added. First, a shallow feature map of 80×80×256 is selected from the multi-level fusion features. This size preserves more spatial details and is beneficial for small object detection. A 1×1 convolution is applied to the shallow feature map to perform channel compression, reducing computational effort while preserving key information. This results in an 80×80×128 basic feature representation for small objects. A depthwise separable convolution is then applied to the basic feature. This convolution first performs channel-by-channel spatial convolution and then performs point-wise convolution to combine channel information, significantly reducing the number of parameters compared to standard convolution. This results in an 80×80×128 deep feature representation. The deep feature representation is then fed into an attention-guided module, which computes the importance weight of each spatial position. By applying global context encoding to the feature map, the relationship between different spatial positions is captured, generating a spatial position importance weight map. This weight map is then multiplied with the original feature map to enhance the feature values ​​of the backpack region and weaken those of the background region, resulting in an 80×80×128 region-enhanced feature. The regional enhancement features are upsampled by a factor of 2 using bilinear interpolation, increasing the feature map resolution from 80×80 to 160×160, a quarter of the original input image size, generating a high-resolution feature map of 160×160×128. The high-resolution feature map is then fused with features transferred from the shallow layers of the backbone network via skip connections, integrating multi-level information using element-by-element addition to obtain a multi-scale fused feature map of 160×160×128.

[0035] The detection head architecture, which applies a series of convolutional and batch normalization layers to the multi-scale fused features, consists of three parallel branches: 1) a bounding box coordinate prediction branch, which outputs four channels representing the center coordinates, width, and height of the bounding box; 2) an object confidence prediction branch, which outputs a single channel representing the confidence level of a backpack detection; and 3) a category prediction branch, which outputs N channels representing the probability distribution of N backpack categories. The feature maps from these three branches are concatenated to form an enhanced feature map for small objects, providing precise location and category information for subsequent backpack detection.

[0036] Anchor box clustering analysis was performed on the annotated data in the training dataset to calculate the most optimal anchor box configuration for the size distribution of children's backpacks. Specifically, the aspect ratios of all backpacks were extracted from the annotated data. These ratios were then clustered using the K-means algorithm into nine categories, corresponding to three anchor boxes for each of the three detection layers. Based on the size of the cluster centers, small, medium, and large anchor boxes were assigned to the 80×80, 40×40, and 20×20 detection layers, respectively. Smaller anchor boxes were assigned to high-resolution feature maps, while larger anchor boxes were assigned to low-resolution feature maps, forming an anchor box configuration tailored to the size distribution of children's backpacks. The network was iteratively trained based on this anchor box configuration and the CIOU (Complete Intersection over Union) loss function. This loss function not only considers the overlapping area of ​​the bounding boxes, but also the center point distance and aspect ratio, resulting in more accurate localization of small objects. A cosine annealing learning rate strategy was used during training. The initial learning rate was set to 0.01 and gradually decreased with each training round, slowly decreasing in the middle and late stages to prevent the model from oscillating in local optima. During training, performance indicators such as mAP (meanAverage Precision) are regularly calculated on the validation dataset. When the indicators no longer improve for a preset number of consecutive times (such as 20 times), the model is considered to have converged, and the current model weights are saved as the children's backpack detection model.

[0037] In a specific embodiment, the process of executing step S103 may specifically include the following steps: The children's backpack detection model detects the position and bounding box of the children's backpack in each frame of the image sequence to obtain visual detection data; Acceleration and angular velocity data are collected from the backpack's built-in inertial measurement unit, and noise filtering and coordinate conversion are performed to obtain motion sensor data; Establish a Kalman filter state space model, define the state vector including position and velocity and its state transfer matrix, and obtain the prediction model; Perform state prediction calculation on the state vector of the previous moment, use the prediction model to infer the current state, and obtain the predicted state; Combine visual detection data and motion sensor data, calculate the deviation between the measured value and the predicted value, update the predicted state, and obtain the fusion state; The real-time position coordinates, movement speed and movement direction of the children's backpack are extracted from the fusion state.

[0038] Specifically, a children's backpack detection model processes image sequences captured by a camera. This detection model receives video frames as input and performs forward inference on each frame to identify and locate the child's backpack in the image. It then outputs the bounding box coordinates containing the backpack's position and the detection confidence score. The detection results are integrated into visual detection data, containing the backpack's position on the two-dimensional image plane. Acceleration and angular velocity data are collected from the backpack's built-in inertial measurement unit (IMU). An IMU is a sensor that integrates an accelerometer and a gyroscope. The accelerometer measures linear acceleration along three axes, while the gyroscope measures angular velocity along three axes. Because raw IMU data often contains noise, it must first be filtered using a low-pass filter to remove high-frequency noise. A coordinate system transformation is then performed, converting the data from the sensor coordinate system to the global coordinate system. The motion parameters of the backpack in three-dimensional space are obtained, forming the motion sensor data. A Kalman filter state-space model is then established. The Kalman filter is a recursive estimation algorithm that alternates between prediction and update phases to optimally estimate the state of a dynamic system. First, define the state vector, which consists of the backpack's position coordinates (x, y, z) and velocity (vx, vy, vz), for a total of six components. Then, establish a state transition matrix to describe the change in state from the previous moment to the current moment. Based on the uniform motion model, the position is equal to the previous position plus the product of the velocity and the time interval, and the velocity remains constant. Also, define the process noise covariance matrix and the observation noise covariance matrix to represent the uncertainty of state prediction and measurement, respectively. These matrices together form the Kalman filter prediction model.

[0039] Based on the constructed prediction model, a state prediction calculation is performed on the state vector at the previous moment. First, the state transition matrix is ​​multiplied by the previous state vector to calculate a prior estimate of the current state. Simultaneously, the state covariance matrix is ​​updated to reflect the increase in uncertainty during the prediction process. This step is equivalent to predicting the current position and velocity of the backpack based on its previous position and velocity, resulting in a predicted state. When new measurement data is acquired, the deviation between the measured and predicted values ​​is calculated by combining the visual inspection data and the motion sensor data, and the predicted state is updated. The Kalman gain is first calculated, which determines the degree of confidence in the new measurement. The confidence is inversely proportional to the measurement noise and directly proportional to the prediction uncertainty. The deviation between the measured and predicted values ​​is then calculated and multiplied by the Kalman gain to obtain a state correction. This correction is added to the predicted state to obtain an updated state estimate. Simultaneously, the state covariance matrix is ​​updated to reflect the reduction in uncertainty after the introduction of the new measurement. This process fuses multi-sensor data, resulting in a fused state. The real-time position coordinates, velocity, and direction of the backpack are extracted from the fused state. The position coordinates are directly obtained from the first three components of the state vector, the velocity magnitude is calculated by taking the square root of the sum of the squares of the three velocity components, and the direction of motion is calculated by taking the inverse tangent of the velocity components.

[0040] In a specific embodiment, the process of executing step S104 may specifically include the following steps: Define the multi-level security area boundaries of the core security area, general security area and extended security area. Each security area is described by a set of polygonal coordinate points to obtain the security area boundary data; Apply the ray method algorithm to the real-time location coordinates of the child's backpack. Shoot a ray from the real-time location coordinates in any direction and calculate the number of intersections with the safety zone boundary. Determine whether the backpack is within the safety zone based on the parity of the intersection number, and obtain the result of whether it is inside or outside the zone. Based on the determination results of whether the backpack is inside or outside the area and the real-time location coordinates of the child's backpack, the distance between the backpack and the nearest safe area boundary is calculated. Combined with the movement speed and direction, the possibility of the backpack leaving or entering the safe area is predicted to obtain the safety risk level. Based on the security risk level, combined with the safety zone boundary distance threshold and the departure time threshold, the warning level is determined, including attention level, warning level and emergency level, and a graded warning sign is obtained; Based on the graded warning signs, a warning message is generated containing the backpack's location information, the time it left the safe area, and the predicted trajectory, generating a complete warning message. The complete warning message is sent to the administrator or guardian through multiple notification channels, and the warning event information is recorded at the same time to generate multi-level security warning information.

[0041] Specifically, a multi-layered safety zone boundary is defined, including a core safety zone, a general safety zone, and an extended safety zone. The core safety zone is typically the primary activity area for children, such as classrooms and activity rooms; the general safety zone includes areas like the school playground and corridors; and the extended safety zone is the area surrounding the school. Each safety zone is described using a set of polygonal coordinate points. For example, a core safety zone might be represented by a rectangular area consisting of four coordinate points: {(100, 200), (100, 400), (300, 400), (300, 200)}. This set of coordinate points constitutes the safety zone boundary data, which is stored in the system as a basis for decision making. After obtaining the real-time location coordinates of the child's backpack, a ray casting algorithm is used to determine whether the backpack is within the safety zone. The ray casting algorithm is a basic algorithm for determining whether a point is inside or outside a polygon. The implementation process involves emitting an infinitely long ray from the backpack's location in any fixed direction (usually horizontally to the right) and counting the number of intersections between the ray and the polygon boundary. If the number of intersections is odd, the point is inside the polygon; if it is even, the point is outside. For example, a ray is sent from the backpack's location (150, 300) to the right, and the number of intersections with the core safety zone boundary is counted. If only one intersection is found, the backpack is considered to be within the safety zone. This method is repeated for all safety zones to determine the backpack's exact location.

[0042] Based on the in-zone determination results and the backpack's real-time location coordinates, the distance between the backpack and the nearest safe zone boundary is calculated. This calculation method finds the shortest distances between the backpack's current location and each line segment of the safe zone boundary, taking the minimum value as the distance from the backpack to the boundary. A dot product operation is performed using the backpack's velocity vector and the boundary normal vector to determine whether the backpack is moving toward the boundary. A negative dot product indicates the backpack is approaching the boundary; a positive one indicates it is moving away. Combining distance and movement trends, the system predicts how long it will take the backpack to reach the boundary at its current speed, thereby determining the safety risk level. Risk levels are categorized as low, medium, and high, corresponding to when the backpack is inside the safe zone, approaching the boundary, or about to or has already crossed the boundary. Based on the safety risk level, combined with preset safety zone boundary distance thresholds (e.g., 10 meters) and exit time thresholds (e.g., 30 seconds), a specific alert level is determined. Alert levels are categorized as attention, warning, and emergency. When a backpack is within the safe zone but within a distance threshold, a caution alert is triggered. When the backpack is about to leave the safe zone within the predicted time threshold, a warning alert is triggered. When the backpack has already left the safe zone or has exceeded a significant distance, an emergency alert is triggered. Different alert levels correspond to different handling strategies and notification frequencies.

[0043] When generating alert content, multiple pieces of information are integrated to form a complete warning message. The alert content includes the backpack's current location (e.g., "southeast corner of the playground"), the predicted time of departure from the safe zone (e.g., "expected to leave the school area in 10 seconds"), the predicted movement trajectory (e.g., "moving toward the school gate"), and the alert level indicator. Messages with different formats and urgency levels are generated for different alert levels. High-level alert messages have more eye-catching fonts and more concise and direct content. The complete alert message is sent to relevant personnel via multiple notification channels. Notification channels include mobile app push notifications, SMS messages, and emails. The appropriate notification method is selected based on the alert level. For example, attention-level alerts are sent only via in-app notifications, warning-level alerts are sent via both app push notifications and SMS messages, and emergency-level alerts initiate phone notifications and multiple channels simultaneously. Detailed information for each alert event is also recorded, including time, location, backpack ID, and alert level, forming a complete alert record.

[0044] The above describes the method for safely locating a children's schoolbag based on deep learning in the embodiment of the present application. The following describes the system for safely locating a children's schoolbag based on deep learning in the embodiment of the present application. Figure 2 In the embodiments of the present application, an embodiment of a children's schoolbag safety positioning system based on deep learning includes: Processing module 201, for performing labeling and enhancement preprocessing on the collected children's schoolbag images to obtain a training data set and a verification data set; A training module 202 is configured to input the training dataset into a modified YOLOv5 object detection network for training, wherein the modified YOLOv5 object detection network includes a spatial attention mechanism for small object detection and an additional small object detection head, and use the validation dataset to evaluate and optimize the training process to obtain a children's schoolbag detection model; and obtain a children's schoolbag detection model. A fusion module 203 is configured to fuse the multi-sensor data using a Kalman filter algorithm based on the children's backpack detection model and the inertial measurement unit data to obtain the real-time position coordinates, movement speed, and movement direction of the children's backpack; The generation module 204 is used to perform ray method safety zone determination processing on the real-time position coordinates of the children's backpack, and generate multi-level safety warning information based on the preset safety zone boundary distance threshold and departure time threshold in combination with the movement speed and direction.

[0045] above Figure 2 The children's backpack safety positioning system based on deep learning in the embodiment of the present invention is described in detail from the perspective of modular functional entities. The children's backpack safety positioning device based on deep learning in the embodiment of the present invention is described in detail from the perspective of hardware processing.

[0046] Figure 3 This is a schematic diagram of the structure of a deep learning-based children's backpack safety locating device provided by an embodiment of the present invention. This deep learning-based children's backpack safety locating device 300 may vary significantly depending on configuration or performance. It may include one or more central processing units (CPUs) 310 (e.g., one or more processors), memory 320, and one or more storage media 330 (e.g., one or more mass storage devices) storing applications 333 or data 332. The memory 320 and storage medium 330 may be either transient or persistent storage. The program stored in the storage medium 330 may include one or more modules (not shown), each of which may include a series of instructions and operations within the deep learning-based children's backpack safety locating device 300. Furthermore, the processor 310 may be configured to communicate with the storage medium 330, executing the series of instructions and operations stored in the storage medium 330 on the deep learning-based children's backpack safety locating device 300 to implement the steps of the deep learning-based children's backpack safety locating method described above.

[0047] The deep learning-based children's backpack safety positioning device 300 may also include one or more power supplies 340, one or more wired or wireless network interfaces 350, one or more input and output interfaces 360, and / or one or more operating systems 331, such as Windows Server, Mac OS X, Unix, Linux, FreeBSD, etc. It will be understood by those skilled in the art that Figure 3 The structure of the deep learning-based children's backpack safety positioning device shown does not constitute a limitation of the deep learning-based children's backpack safety positioning device provided by the present invention, and may include more or fewer components than shown in the figure, or a combination of certain components, or a different arrangement of components.

[0048] The present invention also provides a computer-readable storage medium, which may be a non-volatile computer-readable storage medium or a volatile computer-readable storage medium. The computer-readable storage medium stores instructions. When the instructions are executed on a computer, the computer executes the steps of the deep learning-based children's backpack safety positioning method.

[0049] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the above-described systems, systems and units can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.

[0050] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for enabling a deep learning-based children's backpack safety positioning device (which can be a personal computer, server, or network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM), random access memory (RAM), disk or optical disk, etc., various media that can store program code.

[0051] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention.

Claims

1. A method for safely locating children's schoolbags based on deep learning, characterized in that: The method comprises: The collected children's schoolbag images are labeled and enhanced to obtain training and validation datasets. Inputting the training dataset into an improved YOLOv5 object detection network for training, wherein the improved YOLOv5 object detection network includes a spatial attention mechanism and an additional small object detection head for small object detection, and using the validation dataset to evaluate and optimize the training process to obtain a children's schoolbag detection model; obtaining a children's schoolbag detection model; Based on the children's backpack detection model and inertial measurement unit data, the multi-sensor data is fused and processed through the Kalman filter algorithm to obtain the real-time position coordinates, movement speed and movement direction of the children's backpack; The real-time position coordinates of the children's backpack are processed for safety zone determination using the ray method, and multi-level safety warning information is generated based on the movement speed and direction and according to the preset safety zone boundary distance threshold and departure time threshold.

2. The method for safely locating children's schoolbags based on deep learning according to claim 1, characterized in that: The collected images of children's school bags are annotated and pre-processed to obtain a training dataset and a verification dataset, including: Images of children's schoolbags are acquired through campus monitoring equipment and scene acquisition devices, including scenes of indoor classrooms, corridors, playgrounds, and outdoor public areas, to obtain raw image data; Performing bounding box annotation on the children's backpack in the original image data, recording the backpack position coordinates and backpack type label, and obtaining backpack annotation data; Normalizing the image size of the backpack annotation data, maintaining the aspect ratio while adjusting the resolution, to obtain a standardized image set; performing illumination adjustment and color balancing processing on the standardized image set to obtain an illumination-enhanced image; Performing multi-view synthesis processing on the illumination-enhanced image to generate a composite scene and obtain a scene-enriched image; The scene enriched images are divided according to a training and verification ratio to obtain a training data set and a verification data set.

3. The method for safely locating children's schoolbags based on deep learning according to claim 1, characterized in that: The training dataset is input into an improved YOLOv5 target detection network for training, wherein the improved YOLOv5 target detection network includes a spatial attention mechanism for small target detection and an additional small target detection head, and the training process is evaluated and optimized using the validation dataset to obtain a children's schoolbag detection model; Get the children's schoolbag detection model, including: Build a YOLOv5 infrastructure with a CSPDarknet53 backbone network, which consists of alternating convolutional layers and residual blocks to extract features from the input image and generate multi-scale feature maps. Adding a spatial attention module to the feature extraction layer of the backbone network, performing channel weight calculation and spatial position weighting on the feature map to obtain an attention-enhanced feature map; Constructing a feature fusion neck structure including a feature pyramid network FPN and a path aggregation network PAN, performing upsampling and downsampling operations on the attention enhancement feature map to obtain a multi-level fusion feature; On the basis of the multi-level fusion features, an additional detection head dedicated to small target detection is added, and the downsampling multiple of the feature map is reduced to obtain a small target enhanced feature map; Performing anchor box cluster analysis on the labeled data in the training dataset, calculating the optimal anchor box size configuration, and assigning anchor box parameters of different scales to the three detection layers to obtain anchor box settings for the size distribution of children's school bags; The network is iteratively trained based on the anchor frame setting and the CIOU loss function. During the training process, a cosine annealing learning rate strategy is used and the model performance is regularly evaluated on the validation dataset. When the performance indicator on the validation dataset no longer improves for a preset number of consecutive times, the optimal model weight is saved to obtain a children's backpack detection model.

4. The method for safely locating children's schoolbags based on deep learning according to claim 3, characterized in that: The construction includes a feature fusion neck structure of a feature pyramid network FPN and a path aggregation network PAN, and performs upsampling and downsampling operations on the attention enhancement feature map to obtain multi-level fusion features, including: Extracting three attention-enhanced feature maps of different scales from the backbone network, including a shallow feature map, a mid-level feature map, and a deep feature map, to obtain a multi-scale feature set; Performing convolution processing on the deep feature map to obtain semantic compression features; The semantic compression feature is enlarged to the same spatial size as the middle-layer feature map through an upsampling operation, and feature channel splicing is performed with the middle-layer feature map to obtain a first fused feature map; Performing convolution processing on the first fused feature map, and enlarging it to the same spatial size as the shallow feature map through an upsampling operation, performing feature channel splicing with the shallow feature map, completing feature pyramid network processing, and obtaining a second fused feature map; Based on the second fused feature map, extract key information and reduce the resolution through convolution and maximum pooling operations, and perform feature channel splicing with the first fused feature map to obtain a third fused feature map; Convolution and maximum pooling operations are performed on the third fusion feature map, and feature channel splicing is performed with the semantic compression feature to complete path aggregation network processing to obtain multi-level fusion features.

5. The method for safely locating children's schoolbags based on deep learning according to claim 4, characterized in that: The method adds an additional detection head dedicated to small target detection on the basis of the multi-level fusion features, reduces the downsampling multiple of the feature map, and obtains a small target enhanced feature map, including: Selecting a shallow feature map from the multi-level fusion features, performing channel compression convolution processing on the shallow feature map, retaining spatial resolution information, and obtaining small target basic features; Performing a depth-separable convolution operation on the basic features of the small target to obtain a deep feature representation; Inputting the deep feature representation into the attention guidance module, calculating the spatial position importance weight map, and enhancing the potential children's backpack area in the feature map to obtain regional enhancement features; Performing an upsampling operation on the region enhancement feature through bilinear interpolation to enlarge the feature map resolution to one-fourth of the original input image to obtain a high-resolution feature map; The high-resolution feature map is fused with the shallow features transmitted through the jump connection, and multi-level information is integrated by element-by-element addition to obtain multi-scale fused features; A detection head structure consisting of a series of convolutional layers and batch normalization layers is applied to the multi-scale fusion features, and three branches are output: a bounding box coordinate prediction branch, an object confidence prediction branch, and a category prediction branch, to generate a small object enhanced feature map suitable for small-sized children's schoolbag detection.

6. The method for safely locating children's schoolbags based on deep learning according to claim 1, characterized in that: The method of fusing the multi-sensor data based on the children's backpack detection model and the inertial measurement unit data through the Kalman filter algorithm to obtain the real-time position coordinates, movement speed and movement direction of the children's backpack includes: Detecting the position and bounding box of the child's school bag in each frame of the image sequence using the child's school bag detection model to obtain visual detection data; Acceleration and angular velocity data are collected from the backpack's built-in inertial measurement unit, and noise filtering and coordinate conversion are performed to obtain motion sensor data; Establish a Kalman filter state space model, define the state vector including position and velocity and its state transfer matrix, and obtain the prediction model; Performing state prediction calculation on the state vector at the previous moment, and using the prediction model to infer the state at the current moment to obtain a predicted state; Combining the visual detection data and the motion sensor data, calculating the deviation between the measured value and the predicted value, updating the predicted state, and obtaining a fusion state; The real-time position coordinates, movement speed and movement direction of the children's backpack are extracted from the fusion state.

7. The method for safely locating children's schoolbags based on deep learning according to claim 1, characterized in that: The ray method safety zone determination processing is performed on the real-time position coordinates of the children's backpack, and in combination with the movement speed and movement direction, multi-level safety warning information is generated according to a preset safety zone boundary distance threshold and a departure time threshold, including: Define the multi-level security area boundaries of the core security area, general security area and extended security area. Each security area is described by a set of polygonal coordinate points to obtain the security area boundary data; Applying a ray method algorithm to the real-time location coordinates of the child's backpack, emitting a ray from the real-time location coordinates in any direction and counting the number of intersections with the safety zone boundary, determining whether the backpack is within the safety zone based on the parity of the number of intersections, and obtaining an inside-outside determination result; Based on the determination result of whether the backpack is inside or outside the area and the real-time location coordinates of the child's backpack, the distance between the backpack and the nearest safe area boundary is calculated. In combination with the movement speed and direction, the possibility of the backpack leaving or entering the safe area is predicted to obtain a safety risk level. Based on the security risk level, combined with the safety zone boundary distance threshold and the departure time threshold, an early warning level is determined, including attention level, warning level and emergency level, to obtain a graded early warning sign; According to the graded warning signs, a warning content including the backpack's location information, the time of leaving the safe area and the predicted trajectory is generated to generate a complete warning message; The complete warning message is sent to the administrator or guardian through multiple notification channels, and the warning event information is recorded at the same time to generate multi-level security warning information.

8. A children's schoolbag safety positioning system based on deep learning, characterized by: Used to implement the method for safely locating a children's schoolbag based on deep learning according to any one of claims 1 to 7, the children's schoolbag safety positioning system based on deep learning comprises: A processing module is used to perform annotation processing and enhancement preprocessing on the collected children's schoolbag images to obtain a training data set and a verification data set; A training module is configured to input the training dataset into an improved YOLOv5 object detection network for training, wherein the improved YOLOv5 object detection network includes a spatial attention mechanism and an additional small object detection head for small object detection, and evaluate and optimize the training process using the validation dataset to obtain a children's schoolbag detection model; and obtain a children's schoolbag detection model; A fusion module is used to fuse the multi-sensor data using a Kalman filter algorithm based on the children's backpack detection model and the inertial measurement unit data to obtain the real-time position coordinates, movement speed and movement direction of the children's backpack; A generation module is used to perform ray method safety area determination processing on the real-time position coordinates of the children's backpack, combine the movement speed and movement direction, and generate multi-level safety warning information according to the preset safety area boundary distance threshold and departure time threshold.

9. A children's schoolbag safety positioning device based on deep learning, characterized in that: It includes a memory and a processor, the memory stores a computer program that can be run on the processor, and when the processor executes the computer program, it implements the deep learning-based children's schoolbag safety positioning method described in any one of claims 1 to 7.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the processor causes the processor to execute the method for safely locating a children's schoolbag based on deep learning as claimed in any one of claims 1 to 7.