Drowning Recognition Method and System for Low Frame Rate SlowFast Model Based on Knowledge Distillation

Through knowledge distillation technology, Fast branching and label smoothing technology is trained, combined with the swimming pool environment data set, a SlowFast model adapted to low frame rates is built, which solves the accuracy problem of the drowning detection system under low light and high reflective conditions, and achieves efficient drowning recognition and accurate rescue.

CN119723683BActive Publication Date: 2025-07-29巨岩智能科技(杭州)有限公司 +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510243579.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-03
Publication Date
2025-07-29
Estimated Expiration
2045-03-03

AI Technical Summary

Technical Problem

The existing drowning detection system is not effective under special conditions such as low light and high reflection, and the high frame rate video input requirements are difficult to meet, resulting in a decrease in drowning recognition accuracy, especially in resource-constrained environments, which is difficult to achieve efficient drowning recognition.

Method used

The low-frame-rate SlowFast model based on knowledge distillation is adopted, and the Fast branch is trained through knowledge distillation technology, combined with the Slow branch parameters of the open-source SlowFast pre-trained model, transfer learning is performed using the swimming pool environment behavior dataset, and label smoothing technology is applied to build a drowning recognition model adapted to low frame rate conditions.

Benefits of technology

Maintaining high-accuracy drowning detection capabilities under low frame rate conditions improves the identification accuracy and adaptability of the model in the swimming pool environment, reduces false positives, and improves rescue efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119723683B_ABST
    Figure CN119723683B_ABST
Patent Text Reader

Abstract

The present invention discloses a drowning recognition method and system for a low-frame-rate SlowFast model based on knowledge distillation. The method includes: obtaining an image to be recognized; inputting the image to be recognized into a drowning recognition model for drowning behavior recognition to obtain a recognition result; wherein, the drowning recognition model includes a SlowFast model, and the drowning recognition model is a model obtained by training a Fast branch through knowledge distillation technology, combining the parameters of the trained Fast branch with the Slow branch of an open-source SlowFast pre-trained model to form a network, and performing transfer learning with a swimming pool environment behavior dataset; and outputting the recognition result. By implementing the method of the present invention, it is possible to adapt to lower frame rate conditions while maintaining a good recognition effect and improve the drowning recognition accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a drowning recognition method, and more specifically to a drowning recognition method and system based on a low-frame-rate SlowFast model with knowledge distillation. Background Art

[0002] In an indoor swimming pool, the pool environment presents various complex characteristics, which pose many challenges to object detection and tracking in a video surveillance system. The main background of the pool is water, and the movements of swimmers cause fluctuations on the water surface, further increasing the difficulty of object detection and tracking. When a swimmer dives underwater, the refraction, turbidity of the water, and the influence of waves make it difficult to clearly observe most body parts. Whether it is natural light or artificial lighting, reflections will occur on the water surface. Although certain elimination can be achieved through preprocessing techniques, due to the water surface fluctuations, the location of the reflection will change accordingly, thus exacerbating the difficulty of visual recognition of the object. In addition, swimmers come from different age groups, and they show various postures and behaviors in the water or on the shore, which places higher requirements on the detection algorithm. Besides the swimmers themselves, there are various fixed or moving objects such as grandstands, life-saving equipment, training equipment, and personal items in the pool and its surrounding environment, and these objects may become interference sources.

[0003] Swimming is not only a popular sport, but also one of the main causes of high incidence of drowning accidents. To ensure the safety of swimming pools, regulations require each pool to be equipped with a certain number of professional lifeguards, but this part of the cost accounts for a relatively large proportion in the operating expenses. During peak hours, the number of lifeguards needs to be increased to achieve full coverage monitoring. Despite being equipped with sufficient lifeguards, problems such as blind spots in the field of vision and inattentiveness may still occur in actual operation. Especially in emergency situations, missing the best rescue opportunity may threaten life safety. Currently, most drowning detection systems rely on visible light images captured by RGB cameras for analysis. Although this method can meet the requirements to a certain extent, its effect will be greatly reduced under special conditions such as low light and high reflectivity. In addition, in order to improve the action recognition effect, some advanced algorithms usually require a relatively high video frame rate input, which poses a significant challenge to the actual application scenario because many places cannot provide stable and high enough frame rate support.

[0004] Therefore, it is necessary to design a new method to adapt to lower frame rate conditions and improve the accuracy of drowning recognition while maintaining a good recognition effect. Summary of the Invention

[0005] The purpose of the present invention is to overcome the defects of the prior art and provide a drowning recognition method and system based on a low-frame-rate SlowFast model with knowledge distillation.

[0006] To achieve the above object, the present invention adopts the following technical solutions: A drowning recognition method based on a low-frame-rate SlowFast model using knowledge distillation, comprising:

[0007] Obtain the image to be recognized;

[0008] Input the image to be recognized into the drowning recognition model for drowning behavior recognition to obtain a recognition result;

[0009] Output the recognition result;

[0010] Wherein, the drowning recognition model includes:

[0011] Train the Fast branch through knowledge distillation technology;

[0012] Combine the parameters of the trained Fast branch with the Slow branch of the open-source SlowFast pre-trained model to form a network to obtain a new SlowFast model;

[0013] Obtain the pool environment behavior dataset;

[0014] Use the pool environment behavior dataset to train the new SlowFast model by applying label smoothing technology to obtain the drowning recognition model.

[0015] A further technical solution thereof is: The training of the Fast branch through knowledge distillation technology includes:

[0016] Using the open-source SlowFast model trained with the Kinetics-700 dataset as a basis, perform high-frame-rate pre-training on the Fast branch of the SlowFast model, and use the existing pre-trained weights of the Fast branch of the SlowFast model as a starting point to obtain a high-frame-rate version of the Fast branch;

[0017] Construct a low-frame-rate version of the Fast branch, and simulate the input format of the full frame by copying some frame images;

[0018] Use the high-frame-rate version of the Fast branch as the teacher model and the low-frame-rate version of the Fast branch as the student model, and transfer the information between the teacher model and the student model through knowledge distillation to obtain the trained Fast branch.

[0019] A further technical solution thereof is: The high-frame-rate pre-training is to perform pre-training on the Fast branch with a video frame sequence composed of 32 frames as the input.

[0020] A further technical solution thereof is: The simulation of the input format of the full frame by copying some frame images includes:

[0021] By selecting 4 frames of images and copying them 8 times respectively, a video frame sequence consisting of 32 frames of images is formed as the input for the training of the Fast branch.

[0022] Its further technical solution is: during the process of training the Fast branch through the knowledge distillation technology, the mean square error loss function is used for processing, so that the low-frame-rate version of the Fast branch approximates the high-frame-rate version of the Fast branch.

[0023] Its further technical solution is: using the label smoothing technology to train the new SlowFast model by using the pool environment behavior data set to obtain a drowning recognition model, including:

[0024] Copy each 4 frames of images in the pool environment behavior data set 8 times to form a 32-frame video input to obtain a sample set;

[0025] Use the sample set to apply the label smoothing technology to train the new SlowFast model to obtain a drowning recognition model.

[0026] Its further technical solution is: using the sample set to apply the label smoothing technology to train the new SlowFast model to obtain a drowning recognition model, including:

[0027] During the training process, the labels in the sample set are transformed, where represents the total number of categories, represents the smoothing parameter, , represents the current actual behavior category; is the label after smoothing.

[0028] The present invention also provides a drowning recognition system for a low-frame-rate SlowFast model based on knowledge distillation, including:

[0029] An image acquisition unit for acquiring an image to be recognized;

[0030] A recognition unit for inputting the image to be recognized into the drowning recognition model for drowning behavior recognition to obtain a recognition result; wherein, the drowning recognition model includes a SlowFast model, and the drowning recognition model is a network formed by training the Fast branch through the knowledge distillation technology and combining the parameters of the Slow branch of the open-source SlowFast pre-trained model, and is a model obtained by transfer learning from the pool environment behavior data set;

[0031] An output unit for outputting the recognition result.

[0032] Its further technical solution is: It further includes a training unit for:

[0033] Training the Fast branch through knowledge distillation technology; combining the parameters of the trained Fast branch with the Slow branch of the open-source SlowFast pre-trained model to form a network, so as to obtain a new SlowFast model; obtaining a pool environment behavior dataset; using the pool environment behavior dataset to train the new SlowFast model by applying label smoothing technology to obtain a drowning recognition model.

[0034] The beneficial effects of the present invention compared with the prior art are:

[0035] In the present invention, the image to be recognized is input into the SlowFast model, and its dual-branch structure is used to process information at different frame rates. The Fast branch is trained through knowledge distillation technology so that it can effectively obtain temporal information from the Slow branch, thereby improving the performance under low frame rate conditions; the parameters of the trained Fast branch are fused with the Slow branch of the open-source SlowFast pre-trained model to form an efficient network structure adapted to low frame rates; transfer learning is performed using the pool environment behavior dataset so that the model can accurately recognize drowning behavior in a specific pool scenario and improve the adaptability of the model to low frame rate images; finally, drowning behavior is recognized through this model and the recognition result is output, so as to still maintain a high-accuracy drowning detection ability under low frame rate conditions.

[0036] The present invention will be further described below in conjunction with the accompanying drawings and specific embodiments. Description of the Drawings

[0037] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings required for the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other accompanying drawings can be obtained based on these drawings without creative efforts.

[0038] Figure 1 It is a schematic flowchart of the drowning recognition method based on the knowledge distillation low frame rate SlowFast model provided by the embodiment of the present invention;

[0039] Figure 2 It is a schematic sub-flowchart of the drowning recognition method based on the knowledge distillation low frame rate SlowFast model provided by the embodiment of the present invention Figure 1 ;

[0040] Figure 3 It is a schematic sub-flowchart of the drowning recognition method based on the knowledge distillation low frame rate SlowFast model provided by the embodiment of the present invention Figure 2 ;

[0041] Figure 4 Schematic diagram of the sub - process of the drowning recognition method based on knowledge distillation for the low - frame - rate SlowFast model provided by the embodiments of the present invention Figure 3 ;

[0042] Figure 5 Schematic diagram of the training process of the Fast branch provided by the embodiments of the present invention;

[0043] Figure 6 Schematic diagram of the swimming pool behavior dataset provided by the embodiments of the present invention;

[0044] Figure 7 Schematic block diagram of the drowning recognition system based on knowledge distillation for the low - frame - rate SlowFast model provided by the embodiments of the present invention;

[0045] Figure 8 Schematic block diagram of the computer device provided by the embodiments of the present invention. Detailed implementation manners

[0046] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0047] It should be understood that when used in this specification and the appended claims, the terms "include" and "comprise" indicate the presence of the described features, wholes, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components, and / or their combinations.

[0048] It should also be understood that the terms used in this specification of the present invention are only for the purpose of describing specific embodiments and are not intended to limit the present invention. As used in this specification of the present invention and the appended claims, unless the context clearly indicates otherwise, the singular forms "a", "an", and "the" are intended to include the plural forms.

[0049] It should be further understood that the term "and / or" used in this specification of the present invention and the appended claims refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations.

[0050] Please refer to Figure 1 , Figure 1Schematic flowchart of a drowning recognition method based on a low-frame-rate SlowFast model with knowledge distillation provided by an embodiment of the present invention. The drowning recognition method based on the low-frame-rate SlowFast model with knowledge distillation is applied to a server. The server interacts with terminals and cameras for data. It uses the high-frame-rate Fast branch as the teacher model, transfers its knowledge to the low-frame-rate Fast branch, and thus improves the performance of the low-frame-rate model; by replicating some frame images, it simulates the full-frame input format to ensure that the low-frame-rate Fast branch can process input information similar to that of the high frame rate; it uses the pool environment behavior dataset for transfer learning and applies label smoothing technology to reduce overfitting and improve the model's recognition ability for unknown behaviors; during the knowledge distillation process, it uses the mean square error loss function to make the low-frame-rate Fast branch approximate the output of the high-frame-rate version and optimize the performance; it performs data augmentation on the pool dataset (such as replicating frame images to form a longer video sequence) to enhance the model's adaptability to low-frame-rate inputs, thereby improving the accuracy of drowning recognition.

[0051] Figure 1 It is a schematic flowchart of the drowning recognition method based on the low-frame-rate SlowFast model with knowledge distillation provided by an embodiment of the present invention. As Figure 1 shown, the method includes the following steps S110 to S130.

[0052] S110. Obtain the image to be recognized.

[0053] In this embodiment, the image to be recognized refers to the image data used for drowning behavior recognition. These images usually come from devices such as surveillance cameras, drones, and mobile phones, and represent real-time video frames or static images that need to detect whether a drowning behavior has occurred.

[0054] S120. Input the image to be recognized into the drowning recognition model for drowning behavior recognition to obtain a recognition result; wherein, the drowning recognition model includes a SlowFast model, and the drowning recognition model is a model obtained by training the Fast branch through knowledge distillation technology and combining the parameters of the trained Fast branch with the Slow branch of an open-source SlowFast pre-trained model for transfer learning using the pool environment behavior dataset.

[0055] In this embodiment, the recognition result refers to whether there is a drowning behavior in the image to be recognized.

[0056] The method of this embodiment uses the SlowFast model to identify the behavior of swimming pool visitors. This model processes video data through two parallel paths: the Slow path with a low frame rate and the Fast path with a high frame rate. This design enables the model to capture both slow and fast motion features in the video simultaneously, thus improving the performance of action recognition and video analysis. However, in practical applications, the frame rate of the swimming pool visitor drowning alarm system cannot meet the frame rate requirements of the Fast path. Therefore, the method of this embodiment is based on the knowledge distillation method to train a low-frame-rate SlowFast model with performance equivalent to that of the high-frame-rate SlowFast model.

[0057] The SlowFast model is a deep learning model for video understanding tasks, specifically designed to process spatio-temporal information and capture motion details in videos. This model performs well in tasks such as video classification and action recognition, and is particularly suitable for processing videos containing fast and slow actions.

[0058] The core idea of the SlowFast model is to use two different network branches to process information at different frame rates:

[0059] Slow branch: Responsible for capturing low-frame-rate video information, usually using a lower sampling frequency (e.g., 3 frames per second). This branch focuses on the global scene information of the video, processes dynamic changes over a longer time scale, and helps understand the overall structure of the video.

[0060] Fast branch: Responsible for capturing high-frame-rate video information, usually using a higher sampling frequency (e.g., 30 frames per second or higher). This branch focuses on fast motions and details in the video, capturing local short-time-scale dynamic changes.

[0061] Through the low-frame-rate Slow branch and the high-frame-rate Fast branch, the model can efficiently capture spatio-temporal features in the video. Combining the features of the two time scales can better handle complex actions and video scenes, improving the robustness of the model in various tasks. It can achieve high accuracy in multiple video understanding tasks, especially in action recognition tasks that require quick responses.

[0062] In one embodiment, please refer to Figure 2 , the above-mentioned drowning recognition model is a model obtained by training the Fast branch through knowledge distillation technology, combining the parameters of the trained Fast branch with the Slow branch of an open-source SlowFast pre-trained model to form a network, and performing transfer learning with a swimming pool environment behavior dataset, including steps S121~S124.

[0063] S121. Train the Fast branch through knowledge distillation technology.

[0064] In one embodiment, please refer to Figure 3 , the above step S121 may include steps S1211 to S1213.

[0065] S1211. Using the open-source SlowFast model trained with the Kinetics-700 dataset as a basis, perform high-frame-rate pre-training on the Fast branch of the SlowFast model, and use the existing pre-trained weights of the Fast branch of the SlowFast model as a starting point to obtain a high-frame-rate version of the Fast branch.

[0066] Among them, the high-frame-rate pre-training uses a video frame sequence composed of 32 frames as input for the pre-training of the Fast branch.

[0067] In this embodiment, an open-source SlowFast model that has been trained on the large-scale video dataset Kinetics-700 is selected as the starting point. This model is widely adopted due to its superior performance in various action recognition tasks. Since the Fast branch of the SlowFast model is designed to capture fast movements, it has high requirements for the frame rate, such as the set 32 frames per second. This makes it very suitable as a pre-training model under high-frame-rate conditions.

[0068] Through further high-frame-rate pre-training, it aims to enhance the Fast branch's ability to capture fast-changing actions (such as drowning), thereby obtaining a high-frame-rate Fast branch with more excellent performance.

[0069] Specifically, using the pre-trained weights of the existing Fast branch of the SlowFast model as initialization parameters, a new network structure FastOnly is constructed, and this structure only retains part of the original Fast branch. Then, a continuous video frame sequence composed of 32 frames is used as input for training to ensure that the FastOnly model can fully learn the representation method of video features at high frame rates.

[0070] S1212. Construct a low-frame-rate version of the Fast branch by replicating some frame images to simulate the input format of the full frame.

[0071] Specifically, by selecting 4 frame images and replicating them 8 times respectively, a video frame sequence composed of 32 frame images is used as input for the training of the Fast branch.

[0072] In this embodiment, considering that in the actual application environment, it may not be possible to provide a high enough frame rate. For example, some surveillance cameras can only output video streams with a low frame rate. In order to make the model adapt to these scenarios, a new Fast branch that can run under low-frame-rate conditions but still maintain good performance must be developed.

[0073] To address the above challenges, this embodiment proposes a method for simulating high - frame - rate input. Specifically, 4 frames of images are selected from the original video, and then each frame is copied 8 times to form a hypothetical video frame sequence containing 32 frames as the input. This approach not only meets the requirements of the Fast branch for the input format but also simplifies the processing flow of low - frame - rate videos.

[0074] In this way, the newly constructed low - frame - rate version of the Fast branch can learn how to effectively process low - frame - rate videos while receiving the same input format as the high - frame - rate pre - training, laying a solid foundation for the subsequent knowledge distillation step.

[0075] S1213: Use the high - frame - rate version of the Fast branch as the teacher model and the low - frame - rate version of the Fast branch as the student model, and transfer the information between the teacher model and the student model through knowledge distillation to obtain the trained Fast branch.

[0076] During the training process, the mean - squared error loss function is used for processing to make the low - frame - rate version of the Fast branch approximate the high - frame - rate version of the Fast branch.

[0077] In this embodiment, as Figure 5 shown, the knowledge distillation method is used to optimize the Fast branch so that it can still maintain high - precision behavior recognition ability under low - frame - rate conditions; specifically, knowledge distillation is a model compression technique that allows a more complex "teacher" model to guide the learning process of a simpler or more efficient "student" model. In this case, the Fast branch pre - trained with high frame rate plays the role of the "teacher", while the newly constructed low - frame - rate version of the Fast branch is the "student".

[0078] To quantify the difference between the teacher model and the student model and guide the latter to better imitate the behavior of the former, the mean - squared error (MSE) is introduced as the loss function. MSE is used to evaluate the similarity between the output feature maps of the teacher model and the student model. The formula is as follows: ; , where L represents the MSE loss between the two features, N represents the number of pixels in the feature map, and correspond to the feature values of the teacher model and the student model at the same position respectively; is the output feature map of the teacher model; is the output feature map of the student model.

[0079] Based on the calculated MSE loss, the parameters of the student model can be automatically adjusted through the backpropagation algorithm to gradually approximate the performance of the teacher model. In this process, the key lies in finding a suitable optimization strategy to minimize the gap between the two, while ensuring that the student model does not overfit the data distribution of the teacher model. The ultimate goal is to enable the low-frame-rate version of the Fast branch to accurately identify emergency situations such as drowning in practical applications.

[0080] In summary, through the carefully designed high-frame-rate pre-training, low-frame-rate version construction, and knowledge distillation process, the method of this patent embodiment has successfully reduced the frame-rate requirement of the Fast branch and achieved efficient drowning behavior recognition in resource-constrained environments.

[0081] S122. Combine the parameters of the trained Fast branch with the Slow branch of the open-source SlowFast pre-trained model to form a network, so as to obtain a new SlowFast model.

[0082] In this embodiment, the new SlowFast model refers to the network formed by combining the parameters of the trained Fast branch with the Slow branch of the open-source SlowFast pre-trained model.

[0083] S123. Obtain a pool environment behavior dataset;

[0084] S124. Use the pool environment behavior dataset to apply label smoothing technology to train the new SlowFast model to obtain a drowning recognition model.

[0085] In one embodiment, referring to Figure 4 , the above step S124 may include steps S1241 to S1242.

[0086] S1241. Duplicate each 4-frame image in the pool environment behavior dataset 8 times to form a 32-frame video input to obtain a sample set;

[0087] S1242. Use the sample set to apply label smoothing technology to train the new SlowFast model to obtain a drowning recognition model.

[0088] Specifically, during the training process, the following transformation is adopted for the labels in the sample set , where represents the total number of categories, represents the smoothing parameter, , represents the current actual behavior category; is the label after smoothing.

[0089] In this embodiment, after completing the pre-training of the Fast branch optimized for high and low frame rate conditions, the next crucial step is to apply these results to a specific scenario - namely, human behavior recognition in a pool environment. To achieve this goal, the method of transfer learning is adopted. Specifically, the open-source SlowFast model is customized and adjusted to adapt to the data characteristics of the pool scenario. The following is the detailed implementation process:

[0090] First, use the behavior dataset designed specifically for the pool environment as shown in Figure 6 . This dataset contains various behavior samples that may occur in the pool and is crucial for training a model that can accurately identify emergency situations such as drowning.

[0091] Slow branch: Given that the Slow branch is mainly responsible for capturing the long-term motion trends in the video, its pre-trained parameters are directly sourced from the open-source SlowFast model trained on the Kinetics-700 dataset. This step ensures that the model can fully utilize the advantages of the large-scale dataset while maintaining the ability to understand slower action changes.

[0092] Fast branch: For the Fast branch, the parameters of the low-frame-rate FastOnly model obtained through high-frame-rate pre-training and knowledge distillation as mentioned above are introduced. The purpose of this is to enable the Fast branch to effectively capture fast action features under low frame rate conditions, especially those behaviors that require timely response in the pool environment.

[0093] Considering the particularity of the Fast branch, at the input level, the method described above is continued: 4 frames of images are selected from the original video, and each frame is copied 8 times to form a 32-frame sequence as the input. This method not only simplifies the data preprocessing process but also ensures consistency with the previous pre-training stage, which helps improve the learning efficiency of the model.

[0094] Finally, after the above preparations are completed, the entire SlowFast model is trained on the pool behavior dataset. During this process, the Slow branch and the Fast branch respectively use the pre-trained weights in their respective domains as the starting point, and in combination with the unique behavior patterns in the pool, gradually adjust the model parameters to finally obtain a customized SlowFast model that can operate under low frame rate conditions and efficiently identify key behaviors in the pool.

[0095] In summary, through a carefully designed transfer learning strategy, the general SlowFast model is transformed into a version specifically suitable for the pool environment, significantly enhancing its practicality and accuracy in specific application scenarios. This method not only demonstrates how to effectively utilize existing resources to solve specific problems but also provides a valuable reference case for other similar fields.

[0096] In addition, during the training process, to enhance the model's ability to recognize behaviors (especially drowning situations) in the pool environment and avoid overfitting problems that may occur during training, the LabelSmoothing technique is introduced in the training phase. This technique adjusts the original one-hot encoded labels to make them "smooth", thereby guiding the model to learn more robust and generalized feature representations.

[0097] In this way, the original clear binary classification labels (i.e., 0 or 1) are transformed into a series of values between the two: for the correct class, the label is no longer an absolute 1 but slightly reduced to , while the incorrect class is assigned a small positive value . This change helps reduce the model's sensitivity to noise or outliers in the training dataset and encourages the model to generate a more uniform and better-generalized probability distribution.

[0098] Furthermore, during training, the cross-entropy loss is calculated based on to evaluate the difference between the model's prediction results and the smoothed actual labels, where, is the target probability obtained after applying label smoothing; is the probability distribution output by the model, representing the model's estimation of the attribution to the i-th class, represents the total number of classes.

[0099] The design of this loss function ensures that the model can not only accurately distinguish each class but also maintain a good sense of uncertainty and diversity, which is particularly important for complex and variable application scenarios such as the pool. Through the above measures, a SlowFast model that can adapt to low-frame-rate video inputs and effectively identify potential dangerous behaviors (such as drowning) is finally constructed, greatly improving the practicality and reliability of the system.

[0100] S130. Output the recognition result.

[0101] Output the above recognition result to the terminal for display.

[0102] In this embodiment, the method of this embodiment conducts a detailed analysis of possible drowning situations in the swimming pool environment and specially designs a SlowFast model that can adapt to low-frame-rate video input. Through knowledge distillation technology, an efficient model dedicated to the swimming pool environment is trained, which can more accurately judge three specific behavior categories - struggling while holding onto the dividing line, struggling in the middle of the lane, and the safe state. This improvement not only enhances the accuracy of the system but also effectively reduces false alarms.

[0103] Equipped with advanced camera stream processing capabilities, it can capture and analyze the behavior patterns of swimmers in real time. When a potential drowning behavior is detected, the system will automatically start a warning countdown and send an alarm message to the lifeguard after the countdown ends. This feature enables pool management personnel to take actions in the first place, greatly shortening the response time and improving the rescue efficiency.

[0104] Special waterproof and anti-fog cameras are installed at key positions in the swimming pool to ensure stable operation even in a humid environment. It utilizes advanced technologies such as server-side behavior analysis and AI pattern recognition to achieve intelligent alarm and accident video recording and storage. The hardware system relied on by the entire method consists of multiple components such as a high-performance computer, a waterproof spherical camera, an on-site monitoring screen, and an alarm device, forming a complete automated monitoring network. All devices are designed with full waterproofing, which can provide accurate location guidance at the moment of danger, helping lifeguards quickly locate and rescue drowning victims. In addition, the system also integrates various sensors and audible and visual alarms, further enhancing its safety and reliability.

[0105] Without changing the existing facilities, the system can be seamlessly integrated only by installing a high-resolution underwater camera on the top of the swimming pool. This method avoids the complexity and cost increase problems brought by traditional construction wiring and at the same time ensures a high accuracy of the algorithm in identifying drowning behaviors.

[0106] It can be seen that the method of this embodiment not only has significant advantages in terms of technology and function, but also its design and deployment methods fully consider the convenience and economy in practical applications.

[0107] The above drowning recognition method based on the knowledge distillation-based low-frame-rate SlowFast model inputs the image to be recognized into the SlowFast model and uses its dual-branch structure to process information at different frame rates. The Fast branch is trained through knowledge distillation technology to effectively obtain temporal information from the Slow branch, thereby improving the performance under low-frame-rate conditions; the trained Fast branch is fused with the Slow branch parameters of the open-source SlowFast pre-trained model to form an efficient network structure adapted to low frame rates; transfer learning is performed using the pool environment behavior dataset to enable the model to accurately recognize drowning behavior in a specific pool scenario and improve the adaptability of the model to low-frame-rate images; finally, the drowning behavior is recognized through this model and the recognition result is output, so as to maintain a high-accuracy drowning detection ability under low-frame-rate conditions.

[0108] Figure 7 FIG. 4 is a schematic block diagram of a drowning recognition system 300 based on the knowledge distillation-based low-frame-rate SlowFast model provided by an embodiment of the present invention. As Figure 7 shown, corresponding to the above drowning recognition method based on the knowledge distillation-based low-frame-rate SlowFast model, the present invention also provides a drowning recognition system 300 based on the knowledge distillation-based low-frame-rate SlowFast model. The drowning recognition system 300 based on the knowledge distillation-based low-frame-rate SlowFast model includes units for executing the above drowning recognition method based on the knowledge distillation-based low-frame-rate SlowFast model, and the system can be configured in a server. Specifically, please refer to Figure 7 FIG. 4, the drowning recognition system 300 based on the knowledge distillation-based low-frame-rate SlowFast model includes an image acquisition unit 301, a recognition unit 302, and an output unit 303.

[0109] The image acquisition unit 301 is used to acquire the image to be recognized; the recognition unit 302 is used to input the image to be recognized into the drowning recognition model to perform drowning behavior recognition to obtain a recognition result; wherein, the drowning recognition model includes a SlowFast model, and the drowning recognition model is a model obtained by training the Fast branch through knowledge distillation technology and combining the trained Fast branch with the Slow branch parameters of the open-source SlowFast pre-trained model and performing transfer learning using the pool environment behavior dataset; the output unit 303 is used to output the recognition result.

[0110] In one embodiment, the drowning recognition system 300 based on the knowledge distillation-based low-frame-rate SlowFast model further includes a training unit for:

[0111] Train the Fast branch using knowledge distillation technology; combine the parameters of the trained Fast branch with the Slow branch of the open-source SlowFast pre-trained model to form a network, so as to obtain a new SlowFast model; obtain a pool environment behavior dataset; use the pool environment behavior dataset to train the new SlowFast model using label smoothing technology to obtain a drowning recognition model.

[0112] In one embodiment, the training unit is further configured to: use the open-source SlowFast model trained with the Kinetics-700 dataset as a basis, perform high-frame-rate pre-training on the Fast branch of the SlowFast model, and use the existing pre-trained weights of the Fast branch of the SlowFast model as a starting point to obtain a high-frame-rate version of the Fast branch; construct a low-frame-rate version of the Fast branch by replicating partial frame images to simulate the input format of full frames; use the high-frame-rate version of the Fast branch as the teacher model and the low-frame-rate version of the Fast branch as the student model, and transmit information between the teacher model and the student model through knowledge distillation to obtain the trained Fast branch.

[0113] In one embodiment, the training unit is further configured to:

[0114] Copy each 4-frame image in the pool environment behavior dataset 8 times to form a 32-frame video input to obtain a sample set; use the sample set to train the new SlowFast model using label smoothing technology to obtain a drowning recognition model.

[0115] In one embodiment, the training unit is further configured to:

[0116] During the training process, perform conversion on the labels in the sample set, where represents the total number of categories, represents the smoothing parameter, , represents the current actual behavior category; is the label after smoothing.

[0117] It should be noted that those skilled in the art can clearly understand that the specific implementation processes of the above-mentioned drowning recognition system 300 of the low-frame-rate SlowFast model based on knowledge distillation and each unit can refer to the corresponding descriptions in the foregoing method embodiments. For the convenience and conciseness of description, they will not be elaborated here.

[0118] The above-mentioned drowning recognition system 300 of the low-frame-rate SlowFast model based on knowledge distillation can be implemented in the form of a computer program, and this computer program can be in such asFigure 8 runs on the computer device shown below.

[0119] Please refer to Figure 8 , Figure 8 which is a schematic block diagram of a computer device provided by an embodiment of the present application. The computer device 500 may be a server. Among them, the server may be an independent server or a server cluster composed of multiple servers.

[0120] Refer to Figure 8 , the computer device 500 includes a processor 502, a memory, and a network interface 505 connected through a system bus 501. Among them, the memory may include a non-volatile storage medium 503 and an internal memory 504.

[0121] The non-volatile storage medium 503 can store an operating system 5031 and a computer program 5032. The computer program 5032 includes program instructions. When the program instructions are executed, the processor 502 can be caused to execute a drowning recognition method based on a low-frame-rate SlowFast model using knowledge distillation.

[0122] The processor 502 is used to provide computing and control capabilities to support the operation of the entire computer device 500.

[0123] The internal memory 504 provides an environment for the operation of the computer program 5032 in the non-volatile storage medium 503. When the computer program 5032 is executed by the processor 502, the processor 502 can be caused to execute a drowning recognition method based on a low-frame-rate SlowFast model using knowledge distillation.

[0124] The network interface 505 is used for network communication with other devices. Those skilled in the art can understand that Figure 8 the structure shown in

[0125] is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device 500 to which the solution of the present application is applied. The specific computer device 500 may include more or fewer components than those shown in the figure, or combine some components, or have a different component layout.

[0126] Obtain the image to be recognized; input the image to be recognized into the drowning recognition model for drowning behavior recognition to obtain a recognition result; wherein, the drowning recognition model includes a SlowFast model, and the drowning recognition model is a model obtained by training the Fast branch through knowledge distillation technology, combining the trained Fast branch with the Slow branch parameters of an open-source SlowFast pre-trained model to form a network, and performing transfer learning using a pool environment behavior dataset; output the recognition result.

[0127] In one embodiment, when the processor 502 implements the step of the drowning recognition model being a model obtained by training the Fast branch through knowledge distillation technology, combining the trained Fast branch with the Slow branch parameters of an open-source SlowFast pre-trained model to form a network, and performing transfer learning using a pool environment behavior dataset, the specific implementation is as follows:

[0128] Train the Fast branch through knowledge distillation technology; combine the trained Fast branch with the Slow branch parameters of an open-source SlowFast pre-trained model to form a network to obtain a new SlowFast model; obtain a pool environment behavior dataset; use the pool environment behavior dataset to apply label smoothing technology to train the new SlowFast model to obtain a drowning recognition model.

[0129] Among them, in the process of training the Fast branch through knowledge distillation technology, a mean squared error loss function is used for processing to make the low-frame-rate version of the Fast branch approximate the high-frame-rate version of the Fast branch.

[0130] In one embodiment, when the processor 502 implements the step of training the Fast branch through knowledge distillation technology, the specific implementation is as follows:

[0131] Use the open-source SlowFast model trained with the Kinetics-700 dataset as the basis to perform high-frame-rate pre-training on the Fast branch of the SlowFast model, and use the existing pre-trained weights of the Fast branch of the SlowFast model as the starting point to obtain a high-frame-rate version of the Fast branch; construct a low-frame-rate version of the Fast branch, and simulate the input format of the full frame by copying partial frame images; use the high-frame-rate version of the Fast branch as the teacher model and the low-frame-rate version of the Fast branch as the student model, and transfer the information between the teacher model and the student model through knowledge distillation to obtain the trained Fast branch.

[0132] Among them, the high-frame-rate pre-training is to perform pre-training on the Fast branch with a video frame sequence composed of 32 frames as the input.

[0133] In one embodiment, when the processor 502 implements the step of simulating the input format of a full frame by copying partial frame images, the following steps are specifically implemented:

[0134] By selecting 4 frame images and copying each 8 times, a video frame sequence containing 32 frame images is formed as the input for training the Fast branch.

[0135] In one embodiment, when the processor 502 implements the step of training the new SlowFast model by applying label smoothing technology using the pool environment behavior dataset to obtain a drowning recognition model, the following steps are specifically implemented:

[0136] For every 4 frame images in the pool environment behavior dataset, copy each 8 times to form 32-frame video input to obtain a sample set; apply label smoothing technology using the sample set to train the new SlowFast model to obtain a drowning recognition model.

[0137] In one embodiment, when the processor 502 implements the step of training the new SlowFast model by applying label smoothing technology using the sample set to obtain a drowning recognition model, the following steps are specifically implemented:

[0138] During the training process, apply to the labels in the sample set for conversion, where represents the total number of categories, represents the smoothing parameter, , represents the current actual behavior category; is the label after smoothing processing.

[0139] It should be understood that in the embodiments of the present application, the processor 502 may be a central processing unit (CPU), and this processor 502 may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Among them, the general-purpose processor may be a microprocessor or this processor may also be any conventional processor, etc.

[0140] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program includes program instructions, and the computer program can be stored in a storage medium, which is a computer-readable storage medium. The program instructions are executed by at least one processor in the computer system to implement the process steps of the embodiments of the above methods.

[0141] Therefore, the present invention also provides a storage medium. The storage medium can be a computer-readable storage medium. The storage medium stores a computer program, wherein when the computer program is executed by a processor, the processor executes the following steps:

[0142] Obtain an image to be recognized; input the image to be recognized into a drowning recognition model for drowning behavior recognition to obtain a recognition result; wherein, the drowning recognition model includes a SlowFast model, and the drowning recognition model is a model obtained by training the Fast branch through knowledge distillation technology, combining the parameters of the trained Fast branch with the Slow branch of an open-source SlowFast pre-trained model to form a network, and performing transfer learning with a pool environment behavior dataset; output the recognition result.

[0143] In one embodiment, when the processor executes the computer program to implement the step that the drowning recognition model is a model obtained by training the Fast branch through knowledge distillation technology, combining the parameters of the trained Fast branch with the Slow branch of an open-source SlowFast pre-trained model to form a network, and performing transfer learning with a pool environment behavior dataset, the following steps are specifically implemented:

[0144] Train the Fast branch through knowledge distillation technology; combine the parameters of the trained Fast branch with the Slow branch of an open-source SlowFast pre-trained model to form a network to obtain a new SlowFast model; obtain a pool environment behavior dataset; use the pool environment behavior dataset to train the new SlowFast model by applying label smoothing technology to obtain a drowning recognition model.

[0145] Among them, in the process of training the Fast branch through knowledge distillation technology, a mean square error loss function is used for processing to make the low-frame-rate version of the Fast branch approximate the high-frame-rate version of the Fast branch.

[0146] In one embodiment, when the processor executes the computer program to implement the step of training the Fast branch through knowledge distillation technology, the following steps are specifically implemented:

[0147] Using the open-source SlowFast model trained on the Kinetics-700 dataset as a basis, pre-train the Fast branch of the SlowFast model at a high frame rate. Use the existing pre-trained weights of the Fast branch of the SlowFast model as a starting point to obtain a high-frame-rate version of the Fast branch; construct a low-frame-rate version of the Fast branch by replicating some frame images to simulate the input format of the full frame; use the high-frame-rate version of the Fast branch as the teacher model and the low-frame-rate version of the Fast branch as the student model, and transfer information between the teacher model and the student model through knowledge distillation to obtain the trained Fast branch.

[0148] Among them, the pre-training at the high frame rate uses a video frame sequence composed of 32 frames as the input for pre-training the Fast branch.

[0149] In one embodiment, when the processor executes the computer program to implement the step of simulating the input format of the full frame by replicating some frame images, the following steps are specifically implemented:

[0150] Select 4 frame images and replicate each 8 times to form a video frame sequence containing 32 frame images as the input for training the Fast branch.

[0151] In one embodiment, when the processor executes the computer program to implement the step of training the new SlowFast model using the label smoothing technique with the pool environment behavior dataset to obtain a drowning recognition model, the following steps are specifically implemented:

[0152] Replicate every 4 frame images in the pool environment behavior dataset 8 times to form a 32-frame video input to obtain a sample set; use the sample set to apply the label smoothing technique to train the new SlowFast model to obtain a drowning recognition model.

[0153] In one embodiment, when the processor executes the computer program to implement the step of training the new SlowFast model using the label smoothing technique with the sample set to obtain a drowning recognition model, the following steps are specifically implemented:

[0154] During the training process, for the labels in the sample set, is used for conversion, where represents the total number of categories, represents the smoothing parameter, , represents the current actual behavior category; is the label after smoothing.

[0155] The storage medium may be a variety of computer-readable storage media such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a magnetic disk, or an optical disc, etc., which can store program codes.

[0156] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present invention.

[0157] In several embodiments provided by the present invention, it should be understood that the disclosed systems and methods can be implemented in other ways. For example, the system embodiments described above are merely illustrative. For example, the division of each unit is only a logical function division, and there may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed.

[0158] The steps in the method embodiments of the present invention can be adjusted, combined, and deleted according to actual needs. The units in the system embodiments of the present invention can be combined, divided, and deleted according to actual needs. In addition, the functional units in each embodiment of the present invention can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit.

[0159] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions to enable a computer device (which can be a personal computer, a terminal, or a network device, etc.) to execute all or part of the steps of the methods described in each embodiment of the present invention.

[0160] As described above, the above are only specific embodiments of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of various equivalent modifications or substitutions, and these modifications or substitutions should all be covered within the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.

Claims

1. A drowning recognition method for a low-frame-rate SlowFast model based on knowledge distillation, characterized in that Including: Obtain the image to be recognized; Input the image to be recognized into the drowning recognition model for drowning behavior recognition to obtain a recognition result; Output the recognition result; Among them, obtaining the drowning recognition model includes: Train the Fast branch through knowledge distillation technology; Combine the parameters of the trained Fast branch with the Slow branch of the open-source SlowFast pre-trained model to form a network to obtain a new SlowFast model; Obtain the pool environment behavior dataset; Use the pool environment behavior dataset to apply label smoothing technology to train the new SlowFast model to obtain a drowning recognition model; The training of the Fast branch through knowledge distillation technology includes: Using the open-source SlowFast model trained with the Kinetics-700 dataset as the basis, pre-train the Fast branch of the SlowFast model at high frame rate, and use the existing pre-trained weights of the Fast branch of the SlowFast model as the starting point to obtain a high-frame-rate version of the Fast branch; Construct a low-frame-rate version of the Fast branch, and simulate the input format of the full frame by copying some frame images; Use the high-frame-rate version of the Fast branch as the teacher model and the low-frame-rate version of the Fast branch as the student model, and transfer the information between the teacher model and the student model through knowledge distillation to obtain the trained Fast branch.

2. The drowning recognition method of the low-frame-rate SlowFast model based on knowledge distillation according to claim 1, characterized in that, The high-frame-rate pre-training is to pre-train the Fast branch with a video frame sequence composed of 32 frames as the input.

3. The drowning recognition method of the low-frame-rate SlowFast model based on knowledge distillation according to claim 2, wherein The simulation of the input format of the full frame by copying some frame images includes: Select 4 frame images and copy them 8 times respectively to form a video frame sequence containing 32 frame images as the input for training the Fast branch.

4. The drowning recognition method of the low-frame-rate SlowFast model based on knowledge distillation according to claim 3, characterized in that, During the process of training the Fast branch through knowledge distillation technology, the mean squared error loss function is used for processing to make the low-frame-rate version of the Fast branch approximate the high-frame-rate version of the Fast branch.

5. The drowning recognition method of the low-frame-rate SlowFast model based on knowledge distillation according to claim 1, characterized in that, The use of the pool environment behavior dataset to apply label smoothing technology to train the new SlowFast model to obtain a drowning recognition model includes: Copy each 4-frame image in the pool environment behavior dataset 8 times to form a 32-frame video input to obtain a sample set; Use the sample set to apply label smoothing technology to train the new SlowFast model to obtain a drowning recognition model.

6. The drowning recognition method of the low-frame-rate SlowFast model based on knowledge distillation according to claim 5, wherein, The use of the sample set to apply label smoothing technology to train the new SlowFast model to obtain a drowning recognition model includes: During the training process, the labels in the sample set are processed using , where K represents the total number of categories, represents the smoothing parameter, = 0.1, y represents the current actual behavior category, is the label after smoothing processing.

7. A drowning recognition system for a low frame rate SlowFast model based on knowledge distillation, characterized in that Including: An image acquisition unit for obtaining the image to be recognized; A recognition unit for inputting the image to be recognized into the drowning recognition model for drowning behavior recognition to obtain a recognition result; An output unit for outputting the recognition result; Among them, a training unit is also included for: Train the Fast branch using the knowledge distillation technique; combine the parameters of the trained Fast branch with the Slow branch of the open-source SlowFast pre-trained model to form a network, so as to obtain a new SlowFast model; obtain a swimming pool environment behavior dataset; use the swimming pool environment behavior dataset to train the new SlowFast model by applying label smoothing technique to obtain a drowning recognition model; The training of the Fast branch using the knowledge distillation technique includes: Using the open-source SlowFast model trained with the Kinetics-700 dataset as a basis, pre-train the Fast branch of the SlowFast model at high frame rate, and use the existing pre-trained weights of the Fast branch of the SlowFast model as a starting point to obtain a high frame rate version of the Fast branch; Construct a low frame rate version of the Fast branch, and simulate the input format of the full frame by copying partial frame images; Use the high frame rate version of the Fast branch as the teacher model and the low frame rate version of the Fast branch as the student model, and transfer the information between the teacher model and the student model through knowledge distillation to obtain the trained Fast branch.

Citation Information

Patent Citations

  • Multi-modal human body action recognition method based on knowledge distillation and adversarial learning

    CN112364708A

  • Lightweight online detection method for human body actions

    CN114613004A