Infant safety analysis method and system based on monitoring data

By dividing and optimizing the monitoring data, the problem of low pose recognition accuracy of pose detection model is solved, and efficient and accurate behavior recognition of young children is achieved.

CN120279602AActive Publication Date: 2025-07-08SICHUAN UNIV
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202510770597.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-10
Publication Date
2025-07-08
Estimated Expiration
2045-06-10

AI Technical Summary

Technical Problem

In the prior art, the pose pose detection model has low accuracy in the pose recognition process based on child safety monitoring data.

Method used

The formed feature data set is divided by skeleton monitoring data to determine whether to recognize dangerous postures. If so, the recognition effectiveness analysis is performed, otherwise the windowing operation parameters are optimized. Finally, whether a behavior detection image is generated based on the recognition effectiveness analysis results. If so, visualization is performed, otherwise the model learning rate is optimized.

Benefits of technology

Real-time and efficient processing of children's safety monitoring data is realized, the efficiency of children's behavior recognition is improved, and the accuracy and efficiency of the model recognition is enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120279602A_ABST
    Figure CN120279602A_ABST
Patent Text Reader

Abstract

The invention discloses an infant safety analysis method and system based on monitoring data, and relates to the technical field of electric digital data processing. The infant safety analysis method based on the monitoring data comprises the following steps: data division; identifying validity analysis; and generating a detection image. According to the method, data division is performed on a formed feature data set through skeleton monitoring data, then whether dangerous posture recognition is performed or not is judged based on a data division result, if yes, recognition validity analysis is performed, if not, windowing operation parameter optimization is performed, finally, whether a behavior detection image is generated or not is judged based on a recognition validity analysis result, and if yes, windowing operation parameter optimization is performed. If yes, the generated behavior detection image is visualized, otherwise, the model learning rate is optimized, and the effect of improving the child behavior recognition efficiency based on the child safety monitoring data is achieved. The problem that in the prior art, in the child behavior recognition process based on child safety monitoring data, the posture recognition accuracy of a pose detection model is low is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of electronic digital data processing, and particularly to a method and system for analyzing the safety of young children based on monitoring data. Background Art

[0002] Implementing skeleton-based pose recognition in practical applications is a difficult task because it involves multiple modules, such as human detection and pose estimation. The skeleton-based method has great robustness advantages in understanding actual human behaviors. The increasing maturity of moving object detection and moving object tracking technologies in computer vision technology and the development of artificial intelligence have laid a solid foundation for the detection of abnormal behaviors of moving objects and provided more possibilities for the detection of abnormal behaviors of moving objects. However, due to the particularity and uncertainty of young children's behaviors and the physical, age, and psychological characteristics of young children themselves, general abnormal behavior detection algorithms often cannot achieve good results, and the recognition efficiency or accuracy of existing recognition algorithms is relatively low, unable to achieve good results.

[0003] For example, a method, device, and storage medium for analyzing process behaviors disclosed in the invention patent announcement with the publication number of CN111143182B include: obtaining monitoring data for monitoring multiple application programming interfaces (APIs) called by a preset process; encoding the monitoring data corresponding to the current API called by the preset process according to the relative change of the corresponding value of the attribute associated with the current API and the previous API in the multiple APIs and the storage type of the parameters corresponding to the current API to obtain an API record of the preset process calling the current API, and writing the API record corresponding to the current API into a behavior record file; when analyzing the behavior of the preset process, reading and parsing each API record from the behavior record file, and statistically analyzing the process behavior of the preset process according to the parsed API record to obtain the process behavior analysis result of the preset process.

[0004] For example, a method and system for monitoring patient behaviors based on big data analysis disclosed in the patent application with the publication number of CN119397445A include: collecting physiological characteristic data and behavior characteristic data of patients; preprocessing the collected physiological characteristic data and behavior characteristic data; calculating physiological characteristic abnormal values according to the preprocessed physiological characteristic data; and performing behavior pattern analysis according to the preprocessed behavior characteristic data and physiological characteristic abnormal values.

[0005] In the prior art, based on deep learning object detection and action recognition algorithms, the behavior patterns of children in a video stream are analyzed to automatically identify abnormalities such as crying, pushing, and staying stationary for a long time. At the same time, by comparing and analyzing the monitoring data in different time periods, high-incidence accident periods (such as lunch time and before school dismissal) are identified, and supervision is deployed in advance.

[0006] However, in the process of implementing the inventive technical solution in the embodiments of the present application, it is found that the above technologies have at least the following technical problems: In the prior art, for most real-world videos in the standard dataset, human poses are not easy to detect (i.e., only partially visible or occluded by other objects), and most of the existing pose detection models cannot detect the actions of a person during a falling motion. There is a problem of low accuracy in pose detection of the pose detection model during the process of identifying the behavior of young children based on young children's safety monitoring data. Summary of the Invention

[0007] The present invention provides a method and system for analyzing the safety of young children based on monitoring data, which solves the problem of low accuracy in pose detection of the pose detection model during the process of identifying the behavior of young children based on young children's safety monitoring data, and realizes the improvement of the efficiency of identifying the behavior of young children based on young children's safety monitoring data.

[0008] The present invention provides a method for analyzing the safety of young children based on monitoring data, including the following steps: Step 1, dividing the formed feature dataset through the skeleton monitoring data to obtain a data division result. The skeleton monitoring data is used to visualize the skeleton information of young children in the target monitoring area, and the data division is used to generate a multi-dimensional data list from the result of segmenting the skeleton monitoring data; Step 2, judging whether to perform dangerous pose recognition based on the data division result at the end of the preset division period. If so, performing recognition effectiveness analysis based on the dangerous pose recognition result at the end of the preset recognition period to obtain a recognition effectiveness analysis result. Otherwise, optimizing the windowing operation parameters. The recognition effectiveness analysis is used to quantify the effectiveness of the constructed pose detection model in performing dangerous pose recognition on the skeleton monitoring data, and the optimization of the windowing operation parameters means adjusting the windowing operation parameters to improve the prediction performance of the pose detection model; Step 3, judging whether to generate a behavior detection image based on the recognition effectiveness analysis result. If so, visualizing the generated behavior detection image. Otherwise, optimizing the model learning rate. The optimization of the model learning rate means adjusting the initial learning rate to improve the response rate of the pose detection model.

[0009] The present invention provides a child safety analysis system based on monitoring data, including: a data partitioning module, an identification effectiveness analysis module, and a detection image generation module; wherein, the data partitioning module is used to partition the formed feature data set through the skeleton monitoring data to obtain a data partitioning result, and the skeleton monitoring data is used to visualize the skeleton information of children in the target monitoring area. The data partitioning is used to generate a multi-dimensional data list from the result of segmenting the skeleton monitoring data; the identification effectiveness analysis module is used to determine whether to perform dangerous posture identification based on the data partitioning result at the end of the preset partitioning period. If so, it performs identification effectiveness analysis based on the dangerous posture identification result at the end of the preset identification period to obtain an identification effectiveness analysis result. Otherwise, it optimizes the windowing operation parameters. The identification effectiveness analysis is used to quantify the effectiveness of the constructed pose detection model in performing dangerous posture identification on the skeleton monitoring data. The windowing operation parameter optimization means adjusting the windowing operation parameters to improve the prediction performance of the pose detection model; the detection image generation module is used to determine whether to generate a behavior detection image based on the identification effectiveness analysis result. If so, it visualizes the generated behavior detection image. Otherwise, it optimizes the model learning rate. The model learning rate optimization means adjusting the initial learning rate to improve the response rate of the pose detection model.

[0010] One or more technical solutions provided in the present invention have at least the following technical effects or advantages: 1. The formed feature data set is partitioned through the skeleton monitoring data, and then it is determined whether to perform dangerous posture identification based on the data partitioning result at the end of the preset partitioning period. If so, identification effectiveness analysis is performed. Otherwise, the windowing operation parameters are optimized. Finally, it is determined whether to generate a behavior detection image based on the identification effectiveness analysis result. If so, the generated behavior detection image is visualized. Otherwise, the model learning rate is optimized, thus realizing the real-time and efficient processing of child safety monitoring data, and further realizing the improvement of the efficiency of child behavior recognition based on child safety monitoring data, effectively solving the problem of low accuracy of posture recognition of the pose detection model in the process of child behavior recognition based on child safety monitoring data in the prior art.

[0011] 2. By dynamically adjusting the sliding window length and the sliding window step size, the preset pose prediction model can better adapt to the changes in the skeleton monitoring data, improve the prediction performance of the preset pose prediction model for the postures of children in the target monitoring area, and further improve the real-time performance of child safety monitoring. At the same time, by sending an alarm instruction and prompting preset personnel to intervene, it is ensured that problems can be handled in a timely manner when problems occur during the optimization process, and further improve the accuracy and efficiency of child behavior recognition in the target monitoring area.

[0012] 3. When the total input-output duration obtained is within the allowable range of the total input-output duration in the database, the difference between the obtained input response duration and the reference input response duration in the database is compensated by the input response duration compensation value to obtain the input response duration score. At the same time, the obtained input response duration score, output response duration score, and input-output response duration score are coupled to obtain the recognition effectiveness interference value, thereby improving the accuracy of obtaining the recognition effectiveness interference value, and further realizing a more accurate evaluation of the recognition effectiveness of the pose detection model within the preset recognition period. BRIEF DESCRIPTION OF THE DRAWINGS

[0013] Figure 1 It is a flowchart of the method for analyzing the safety of young children based on monitoring data provided by an embodiment of the present application; Figure 2 It is a framework diagram of the skeleton monitoring and pose recognition of young children provided by an embodiment of the present application; Figure 3 It is a flowchart of the skeleton monitoring and pose recognition of young children provided by an embodiment of the present application; Figure 4 It is a schematic structural diagram of the system for analyzing the safety of young children based on monitoring data provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0014] In the embodiments of the present application, by providing a method and system for analyzing the safety of young children based on monitoring data, the problem of low accuracy of pose detection model pose recognition in the process of young children behavior recognition based on young children safety monitoring data in the prior art is solved. The feature data set formed by the skeleton monitoring data is divided to obtain a data division result, and then based on the data division result at the end of the preset division period, the model prediction performance evaluation value is obtained. At the same time, based on the obtained model prediction evaluation value, it is judged whether to perform dangerous pose recognition through the constructed pose detection model. If so, based on the dangerous pose recognition result at the end of the preset recognition period, the recognition effectiveness analysis is carried out to obtain the recognition effectiveness analysis result. Otherwise, the windowing operation parameters are optimized. Finally, based on the recognition effectiveness analysis result, it is judged whether to generate a behavior detection image. If so, the generated behavior detection image is visualized. Otherwise, the model learning rate is optimized, realizing an improvement in the efficiency of young children behavior recognition based on young children safety monitoring data.

[0015] The technical solution in the embodiments of the present application is to solve the problem of low accuracy of pose detection model pose recognition in the process of young children behavior recognition based on young children safety monitoring data. The general idea is as follows: The feature dataset formed by the skeleton monitoring data is partitioned, and then based on the data partitioning result at the end of the preset partitioning period, it is judged whether to perform dangerous posture recognition. If so, the recognition effectiveness analysis is carried out; otherwise, the windowing operation parameters are optimized. Finally, based on the recognition effectiveness analysis result, it is judged whether to generate a behavior detection image. If so, the generated behavior detection image is visualized; otherwise, the model learning rate is optimized, achieving the effect of improving the efficiency of infant behavior recognition based on infant safety monitoring data.

[0016] To better understand the above technical solution, the above technical solution will be described in detail below in conjunction with the accompanying drawings of the specification and specific embodiments.

[0017] As Figure 1 shown, it is a flowchart of the infant safety analysis method based on monitoring data provided by an embodiment of the present application. The infant safety analysis method based on monitoring data provided by an embodiment of the present application includes the following steps: Step 1: Partition the feature dataset formed by the skeleton monitoring data to obtain a data partitioning result. The skeleton monitoring data is used to visualize the infant skeleton information in the target monitoring area. The infant skeleton information includes the infant's bone position, movement trajectory, and posture characteristics. The skeleton monitoring data represents the result of denoising the data points in the kindergarten monitoring images collected in the target monitoring area and serves as the input of a preset posture prediction model. The preset posture prediction model represents a prediction model that performs posture estimation and movement trajectory analysis based on the obtained infant skeleton information after pre-training. The preset posture prediction model in the present application refers to the YOLOv8 model, that is, the latest version of the YOLO (You Only Look Once) series, which can more quickly and accurately detect multiple data points in the kindergarten monitoring images. The feature dataset represents the result of feature fusion and forms a dataset with spatio-temporal characteristics after combining time series information. The result of feature fusion represents the coordinates corresponding to the data feature points output by the preset posture prediction model.

[0018] Among them, data partitioning is used to generate a multi-dimensional data list of the segmented results of the skeleton monitoring data. The multi-dimensional data list stores the dimensional information corresponding to the infant skeleton information in a structured form. For example, the infant bone position coordinates, posture categories (such as standing, sitting), and movement speed; the following five behavior situations are defined as dangerous behaviors: fighting, choking while eating, on the edge of the stairs, crying, and falling. The training data and test data after data partitioning are both multi-dimensional data lists, and each segment of data needs to be segmented, that is, windowed. The five dangerous behaviors are usually presented in binary in the multi-dimensional data list. For example, if fighting is detected, the position corresponding to fighting in the multi-dimensional data list is recorded as 1, otherwise it is recorded as 0.

[0019] Step 2: Determine whether to perform dangerous pose recognition based on the data division result at the end of the preset division period. If so, perform recognition effectiveness analysis based on the dangerous pose recognition result at the end of the preset recognition period to obtain the recognition effectiveness analysis result. Otherwise, optimize the windowing operation parameters. Recognition effectiveness analysis is used to quantify the effectiveness of the constructed pose detection model in performing dangerous pose recognition on the skeleton monitoring data. Optimizing the windowing operation parameters means adjusting the windowing operation parameters to improve the prediction performance of the pose detection model. The pose detection (human key point detection) model is usually applied in the field of computer vision. The dangerous pose recognition result represents the recognition effectiveness data obtained by the pose detection model at the end of the preset recognition period. The recognition effectiveness data includes the input response duration, the output response duration, and the total input-output duration. The input response duration is used to quantify the time length for the pose detection model to receive and successfully input the training data. The output response duration is used to quantify the time length from when the pose detection model starts to output the pose estimation result to the successful end of the output. The total input-output duration is used to quantify the actual working time length of the pose detection model within the preset recognition period.

[0020] Step 3: Determine whether to generate a behavior detection image based on the recognition effectiveness analysis result. If so, visualize the generated behavior detection image. Otherwise, optimize the model learning rate. Optimizing the model learning rate means adjusting the initial learning rate to improve the response rate of the pose detection model.

[0021] In this embodiment, as Figure 2 shown, it is the framework diagram of the infant skeleton monitoring and pose recognition provided by the embodiment of the present application. As Figure 3 shown, it is the flow chart of the infant skeleton monitoring and pose recognition provided by the embodiment of the present application. By obtaining the total input-output duration, determining whether the convergence is qualified, and taking corresponding adjustment measures, the recognition effectiveness of the pose detection model in the infant safety analysis method based on the monitoring data can be ensured. This method not only improves the recognition efficiency of the model but also enhances the generalization ability of the model, providing a more reliable and efficient monitoring means for infant safety.

[0022] Furthermore, it is determined whether to perform dangerous posture recognition based on the data division result at the end of the preset division period. The specific process includes: when the obtained model prediction performance evaluation value is not greater than the preset model prediction performance evaluation value in the database, the data division result at the end of the preset division period is recorded as qualified data division and dangerous posture recognition is performed. The dangerous posture recognition is executed by the pose detection model trained with the training data in the classification data. The classification data includes the training data and test data obtained from the data division result corresponding to the qualified data division. The pose detection model refers to a classification model with behavior recognition function obtained by inputting the skeleton image of children in the kindergarten surveillance image into the ST-GCN network. During the training of the ST-GCN network: the training data of the classifier mainly comes from the surveillance video data provided by a certain kindergarten institution, and self-shot custom dangerous behavior videos are used as abnormal data to enhance data diversity. The data source is legal and meets the project requirements. For the collected data, the useless parts with messy, blurred, and occluded content are deleted, and the clear valid data is selected.

[0023] When the obtained model prediction performance evaluation value is greater than the preset model prediction performance evaluation value in the database, the data division result at the end of the preset division period is recorded as unqualified data division and the windowing operation parameter optimization is performed. The model prediction performance evaluation value represents the harmonic average result of the generation duration of the multi-dimensional data list and the segmentation duration of the skeleton surveillance data within the preset division period. The harmonic average is used to eliminate the interference of extreme values in the generation duration of the multi-dimensional data list and the segmentation duration of the skeleton surveillance data on the model prediction performance evaluation process of the preset posture prediction model. The windowing operation parameters include the sliding window length and the sliding window step size.

[0024] The aforementioned database is established before the design of the method for analyzing children's safety based on surveillance data and is used to store various types of set data. The database includes, but is not limited to, the preset model prediction performance evaluation value, the preset recognition effectiveness interference value, the preset division period, and the preset recognition period. The various values therein are directly set by technical personnel. Among them, the setting basis of the preset recognition effectiveness interference value can be determined according to the actual application scenario of the pose detection model. For example, the preset recognition effectiveness interference value is represented by the sum average of the historical recognition effectiveness interference values of the pose detection model in the database during the historical recognition period. In addition, the various values in the database can be set and fine-tuned by technical personnel according to actual debugging.

[0025] In this embodiment, when the data division is qualified, the recognition of dangerous postures can ensure the reliability of the recognition results. When the data division is unqualified, optimization operations are carried out, which can continuously improve the prediction performance of the model, laying a solid foundation for subsequent monitoring and recognition work. At the same time, the use of the harmonic mean helps to reduce the interference of extreme values on model evaluation, making the evaluation results more objective and accurate, thereby effectively improving the accuracy and stability of the pose detection model.

[0026] Furthermore, the specific steps for optimizing the windowing operation parameters include: inputting the deviation of the obtained model prediction performance evaluation value into the offline reinforcement learning algorithm of the preset pose prediction model to output the actual increase amplitude of the sliding window length. The deviation of the model prediction performance evaluation value is used to quantify the difference degree between the obtained model prediction performance evaluation value and the preset model prediction performance evaluation value, that is, the difference between the obtained model prediction performance evaluation value and the preset model prediction performance evaluation value. The preset model prediction performance evaluation value represents the difference between the average value of the generation duration of the multi-dimensional data list within the preset division period and the segmentation duration of the skeleton monitoring data and the reference average value in the corresponding database. The reference average value represents the historical average value of the generation duration of the historical multi-dimensional data list and the historical segmentation duration of the historical skeleton monitoring data within the historical division period in the database.

[0027] After increasing the sliding window length once, the obtained first model prediction performance evaluation value deviation is re-input into the offline reinforcement learning algorithm of the preset pose prediction model to output the actual decrease amplitude of the sliding window step. The first model prediction performance evaluation value deviation represents the difference between the re-obtained model prediction performance evaluation value after increasing the sliding window length once and the preset model prediction performance evaluation value.

[0028] If the decrease amplitude corresponding to the second model prediction performance evaluation value deviation obtained after one windowing operation parameter optimization is not greater than the reference decrease amplitude in the corresponding database, an alarm instruction is sent. Otherwise, continue with the windowing operation parameter optimization. If the second model prediction performance evaluation value deviation obtained within the preset number of optimization times of the windowing operation parameter optimization is not greater than 0, the windowing operation parameter optimization is completed. Otherwise, prompt the preset personnel to intervene. The second model prediction performance evaluation value deviation represents the difference between the re-obtained model prediction performance evaluation value after one windowing operation parameter optimization and the preset model prediction performance evaluation value. One windowing operation parameter optimization includes one increase in the sliding window length and one decrease in the sliding window step.

[0029] In this embodiment, the length of the sliding window determines the prediction accuracy of the preset pose prediction model for the skeleton monitoring data. However, an increase in the length of the sliding window may increase the computational amount. Therefore, when the length of the sliding window is increased once, it must be accompanied by a decrease in the sliding window step size because the sliding window step size determines the update frequency of the preset pose prediction model for the skeleton monitoring data. A decrease in the sliding window step size means that the preset pose prediction model will update the prediction result more frequently, thereby improving the real-time performance.

[0030] Therefore, by adjusting the sliding window step size, a balance can be found between the prediction accuracy and the real-time performance, enabling the pose detection model to better adapt to different monitoring scenarios and data characteristics, thereby improving the prediction performance. The optimized model can more accurately identify the dangerous poses of toddlers, reduce false alarms and missed detections, improve the reliability of the monitoring system, and further enhance the operating efficiency of the entire monitoring system.

[0031] Furthermore, before performing the recognition effectiveness analysis, it also includes determining whether to perform the recognition effectiveness analysis based on the obtained total input-output duration. The specific process is as follows: When the obtained total input-output duration is within the allowable range of the total input-output duration in the database, it is determined that the pose detection model converges qualified and the recognition effectiveness analysis is performed; when the obtained total input-output duration is not within the allowable range of the total input-output duration in the database, it is determined that the pose detection model converges unqualified and the regularization term is adjusted. The allowable range of the total input-output duration represents the range corresponding to the maximum and minimum values of the historical input-output response duration of the pose detection model in the database during the historical recognition period.

[0032] Among them, the specific process of adjusting the regularization term is as follows: When the obtained total input-output duration is greater than the maximum allowable input-output duration in the database (i.e., the maximum value of the historical input-output response duration), it is determined that the pose detection model is overfitting and the regularization term of the pose detection model is increased based on the obtained increase amplitude of the regularization term until the newly obtained total input-output duration is within the allowable range of the total input-output duration in the database. The increase amplitude of the regularization term represents the result mapped in the database from the difference between the obtained total input-output duration and the maximum allowable input-output duration.

[0033] When the total input-output duration obtained is less than the minimum allowable input-output duration in the database (i.e., the minimum value of the historical input-output response duration), it is determined that the pose detection model is underfitted, and the regularization term of the pose detection model is reduced based on the reduction amplitude of the obtained regularization term until the total input-output duration obtained again is within the allowable range of the input-output duration in the database. The reduction amplitude of the regularization term represents the result mapped in the database of the difference between the minimum allowable input-output duration and the total input-output duration obtained.

[0034] In this embodiment, by introducing a judgment process based on the total input-output duration and a regularization term adjustment mechanism, the convergence of the pose detection model can be evaluated more effectively, and targeted adjustments can be made according to the fitting state of the pose detection model. This can not only improve the recognition accuracy of the pose detection model but also optimize the response speed of the pose detection model, thereby further enhancing the overall performance of infant skeleton monitoring and dangerous pose recognition.

[0035] Furthermore, an analysis of the recognition effectiveness is performed based on the dangerous pose recognition results at the end of the preset recognition period. The specific process includes: when the total input-output duration obtained is within the allowable range of the input-output duration in the database, first, the difference between the obtained input response duration and the reference input response duration in the database is compensated by the input response duration compensation value to obtain the input response duration score. The specific limiting expression of the input response duration score SYX1 is: SYX1 = y1×Y1 / Y10, where SYX1 represents the input response duration score of the pose detection model within the preset recognition period, y1 represents the input response duration compensation value, Y1 represents the input response duration of the pose detection model within the preset recognition period, and Y10 represents the reference input response duration, and the reference input response duration is represented by the result of summing and averaging the historical input response durations of the pose detection model in the database within the historical recognition period.

[0036] Then, the difference between the obtained output response duration and the reference output response duration in the database is compensated by the output response duration compensation value to obtain the output response duration score. The specific limiting expression of the output response duration score SYX2 is: SYX2 = y2×Y2 / Y20, where SYX2 represents the output response duration score of the pose detection model within the preset recognition period, y2 represents the output response duration compensation value, Y2 represents the output response duration of the pose detection model within the preset recognition period, and Y20 represents the reference output response duration, and the reference output response duration is represented by the result of summing and averaging the historical output response durations of the pose detection model in the database within the historical recognition period.

[0037] Next, the difference between the obtained input-output response duration and the reference input-output response duration in the database is compensated by the input-output response duration compensation value to obtain the input-output response duration score. The specific limiting expression of the input-output response duration score SYX3 is: SYX3 = y3 × Y3 / Y30, where SYX3 represents the input-output response duration score of the pose detection model within the preset recognition period, y3 represents the input-output response duration compensation value, Y3 represents the input-output response duration of the pose detection model within the preset recognition period, and Y30 represents the reference input-output response duration. The reference input-output response duration is represented by the result of summing and averaging the historical input-output response durations of the pose detection model in the database during the historical recognition period.

[0038] Finally, the obtained input response duration score, output response duration score, and input-output response duration score are coupled to obtain the recognition effectiveness interference value. The recognition effectiveness interference value represents the quantification data of the combined influence degree of the input response duration, output response duration, and input-output response duration on the recognition effectiveness of the pose detection model. The specific limiting expression of the recognition effectiveness interference value SYX is: SYX = SYX1 + SYX2 + SYX3, where SYX represents the recognition effectiveness interference value of the pose detection model within the preset recognition period.

[0039] Among them, the input response duration, output response duration, and input-output response duration are all recorded in real time by the time sensor, and have the same unit as the reference input response duration, reference output response duration, and reference input-output response duration, which is milliseconds (ms).

[0040] The input response duration compensation value, output response duration compensation value, and input-output response duration compensation value are respectively the influence degrees of the preset input response duration, output response duration, and input-output response duration in the database on the recognition process of the pose detection model. Specifically, the database stores preset compensation values corresponding to the input response duration, output response duration, and input-output response duration. There is a preset mapping relationship between these compensation values and the input response duration, output response duration, and input-output response duration. This mapping relationship can be one-to-one or many-to-one. For example, in practical applications, the real-time input response duration, output response duration, and input-output response duration can be input into this mapping relationship to quickly obtain the corresponding compensation values. The value ranges of the input response duration compensation value, output response duration compensation value, and input-output response duration compensation value in this example are usually from 0 to 1, and the sum of the three is 1.

[0041] In this embodiment, the recognition effectiveness interference value increases as the input response duration, output response duration, and input-output response duration increase. Among them, the increase in the input response duration may cause the system to accumulate more delays when processing initial data, thereby indirectly affecting the timeliness of the output response. The extension of the output response duration may act on the efficiency of the input response. For example, the accumulation of input data may be exacerbated due to the lag of the feedback mechanism. The synchronous increase in the input-output response duration may form a compound delay effect, causing the overall response ability of the system to show a non-linear attenuation trend.

[0042] By considering the above mutual influence mechanism, it is helpful to establish a multi-dimensional time series response model, quantify the coupling strength of each response duration parameter, and optimize the data flow scheduling strategy through a dynamic weight allocation algorithm. Furthermore, the efficiency of infant behavior recognition based on infant safety monitoring data is improved, effectively solving the problem of low pose recognition accuracy of the pose detection model in the process of infant behavior recognition based on infant safety monitoring data in the prior art.

[0043] Furthermore, the specific process of optimizing the model learning rate is as follows: Map the obtained deviation of the recognition effectiveness interference value in the database to obtain the initial learning rate reduction amplitude of the pose detection model. The deviation of the recognition effectiveness interference value is used to quantify the difference between the obtained recognition effectiveness interference value and the preset recognition effectiveness interference value in the database, that is, the difference between the obtained recognition effectiveness interference value and the preset recognition effectiveness interference value in the database. Input the obtained initial learning rate reduction amplitude and the preset initial learning rate reduction amplitude into the linear regression algorithm of the pose detection model to output the actual initial learning rate reduction amplitude. The actual initial learning rate reduction amplitude is used to ensure that the pose detection model improves the response rate of the pose detection model to the feature data set under the condition of qualified convergence. If the deviation of the recognition effectiveness interference value obtained after the preset number of initial learning rate reductions is not greater than 0, the model learning rate optimization is completed; otherwise, prompt the preset personnel to intervene.

[0044] In this embodiment, through the fusion of dynamic mapping and linear regression, the learning rate adjustment strategy can adapt to the change of interference values in different scenarios, avoiding over-adjustment or under-adjustment problems caused by fixed step size adjustment. On the premise of ensuring model convergence, the convergence efficiency is improved by optimizing the learning rate reduction amplitude, reducing unnecessary iteration times, and significantly shortening the model training time to meet the requirements of different hardware environments and task complexities.

[0045] Such as Figure 4As shown in the figure, it is a schematic structural diagram of the infant safety analysis system based on monitoring data provided by the embodiments of the present application. The infant safety analysis system based on monitoring data provided by the embodiments of the present application includes: a data partitioning module, an identification effectiveness analysis module, and a detection image generation module; among them, the data partitioning module is used to partition the formed feature data set through the skeleton monitoring data to obtain a data partitioning result. The skeleton monitoring data is used to visualize the infant skeleton information in the target monitoring area. The data partitioning is used to generate a multi-dimensional data list of the result of segmenting the skeleton monitoring data; the identification effectiveness analysis module is used to judge whether to perform dangerous posture identification based on the data partitioning result at the end of the preset partitioning period. If so, based on the dangerous posture identification result at the end of the preset identification period, perform identification effectiveness analysis to obtain an identification effectiveness analysis result. Otherwise, perform windowing operation parameter optimization. The identification effectiveness analysis is used to quantify the effectiveness of the constructed pose detection model in performing dangerous posture identification on the skeleton monitoring data. The windowing operation parameter optimization means to improve the prediction performance of the pose detection model by adjusting the windowing operation parameters; the detection image generation module is used to judge whether to generate a behavior detection image based on the identification effectiveness analysis result. If so, visualize the generated behavior detection image. Otherwise, perform model learning rate optimization. The model learning rate optimization means to improve the response rate of the pose detection model by adjusting the initial learning rate.

[0046] In this embodiment, through the collaborative work among the data partitioning module, the identification effectiveness analysis module, and the detection image generation module, in pose detection, a pose detection model of YOLOv8 trained with the Microsoft Common Object in Context (COCO) dataset is used to generate the motion skeleton corresponding to the child; in behavior classification, a spatio-temporal graph convolutional network (ST-GCN) is used for real-time behavior classification; at the same time, multi-modal recognition is introduced, mainly through video and audio for infant safety monitoring.

[0047] Among them, multi-modal recognition: not only uses the image data of the data set as the training basis, but also realizes multi-modal fusion and interaction between image labels and voice labels through the method of voice labeling. Through the cooperation of voice and video, the reliability of detection and the adaptability of the scenario are significantly improved, more natural human-computer interaction is realized, and the situation of single-modal failure can be effectively dealt with.

[0048] At the same time, the "Smart Eye - Campus Guardian" video analysis system is developed in HTML language, providing a simple and convenient interaction platform for users, and visualizing the generated behavior detection images, which helps to overcome the inaccurate detection caused by the limited body shape of children in the current detection model, realizing intelligent monitoring and early warning of the safety status of infants in the target monitoring area, and providing more effective protection for infant safety.

[0049] In summary, in the embodiment of the present application, the formed feature data set is divided by the skeleton monitoring data, and then it is determined whether to perform dangerous posture recognition based on the data division result at the end of the preset division period. If so, the recognition effectiveness analysis is carried out. Otherwise, the windowing operation parameters are optimized. Finally, it is determined whether to generate a behavior detection image based on the recognition effectiveness analysis result. If so, the generated behavior detection image is visualized. Otherwise, the model learning rate is optimized, thereby realizing the real-time and efficient processing of the child safety monitoring data, and further realizing the improvement of the efficiency of child behavior recognition based on the child safety monitoring data, effectively solving the problem of low accuracy of the pose detection model in the process of child behavior recognition based on the child safety monitoring data in the prior art.

[0050] Although the preferred embodiments of the present invention have been described, those skilled in the art can make additional changes and modifications once they learn the basic creative concept. Therefore, the appended claims are intended to be construed to include the preferred embodiments as well as all changes and modifications falling within the scope of the present invention.

[0051] Obviously, those skilled in the art can make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalent technologies, the present invention also intends to include these modifications and variations.

Claims

1. A method for analyzing the safety of young children based on monitoring data, characterized in that, It includes the following steps: Step 1: Divide the formed feature dataset through the skeleton monitoring data to obtain a data division result. The skeleton monitoring data is used to visualize the toddler skeleton information in the target monitoring area, and the data division is used to generate a multi-dimensional data list from the result of segmenting the skeleton monitoring data; Step 2: Based on the data division result at the end of the preset division period, determine whether to perform dangerous pose recognition. If so, perform recognition effectiveness analysis based on the dangerous pose recognition result at the end of the preset recognition period to obtain a recognition effectiveness analysis result. Otherwise, optimize the windowing operation parameters. The recognition effectiveness analysis is used to quantify the effectiveness of the constructed pose detection model in performing dangerous pose recognition on the skeleton monitoring data, and the windowing operation parameter optimization means adjusting the windowing operation parameters to improve the prediction performance of the pose detection model; Step 3: Based on the recognition effectiveness analysis result, determine whether to generate a behavior detection image. If so, visualize the generated behavior detection image. Otherwise, optimize the model learning rate. The model learning rate optimization means adjusting the initial learning rate to improve the response rate of the pose detection model.

2. The method for analyzing the safety of young children based on monitoring data according to claim 1, wherein The toddler skeleton information includes the toddler's bone position, movement trajectory, and pose characteristics; The skeleton monitoring data represents the result after denoising the data points in the kindergarten monitoring images collected in the target monitoring area and serves as the input to the preset pose prediction model; The preset pose prediction model represents a prediction model that performs pose estimation and movement trajectory analysis based on the obtained toddler skeleton information after pre-training; The feature dataset represents the result of feature fusion and forms a dataset with spatio-temporal features after combining time series information; The result of the feature fusion represents the coordinates corresponding to the data feature points output by the preset pose prediction model; The multi-dimensional data list stores the dimension information corresponding to the toddler skeleton information in a structured form.

3. The method for analyzing the safety of young children based on monitoring data according to claim 1, wherein The process of determining whether to perform dangerous pose recognition based on the data division result at the end of the preset division period specifically includes: When the obtained model prediction performance evaluation value is not greater than the preset model prediction performance evaluation value in the database, record the data division result at the end of the preset division period as qualified data division and perform dangerous pose recognition; When the obtained model prediction performance evaluation value is greater than the preset model prediction performance evaluation value in the database, record the data division result at the end of the preset division period as unqualified data division and perform windowing operation parameter optimization; The model prediction performance evaluation value represents the harmonic mean result of the generation duration of the multi-dimensional data list and the segmentation duration of the skeleton monitoring data within the preset division period; The harmonic mean is used to eliminate the interference of extreme values in the generation duration of the multi-dimensional data list and the segmentation duration of the skeleton monitoring data on the process of evaluating the prediction performance of the preset pose prediction model; The windowing operation parameters include the sliding window length and the sliding window step size.

4. The method for analyzing the safety of young children based on monitoring data according to claim 3, wherein, The dangerous pose recognition is performed by the pose detection model trained with the training data in the classification data. The classification data includes training data and test data obtained from the data division results corresponding to qualified data division. The pose detection model represents a classification model with behavior recognition function obtained by inputting the skeleton image of children in the kindergarten surveillance image into the ST-DCN network.

5. The method for analyzing the safety of young children based on monitoring data according to claim 3, wherein The specific steps for optimizing the windowing operation parameters include: Inputting the deviation of the obtained model prediction performance evaluation value into the off-line reinforcement learning algorithm of the preset pose prediction model to output the actual increase amplitude of the sliding window length. The deviation of the model prediction performance evaluation value is used to quantify the difference degree between the obtained model prediction performance evaluation value and the preset model prediction performance evaluation value. After increasing the sliding window length once, inputting the obtained first model prediction performance evaluation value deviation into the off-line reinforcement learning algorithm of the preset pose prediction model to output the actual decrease amplitude of the sliding window step size. If the decrease amplitude corresponding to the second model prediction performance evaluation value deviation obtained after one windowing operation parameter optimization is not greater than the reference decrease amplitude in the corresponding database, an alarm instruction is sent; otherwise, continue to optimize the windowing operation parameters. If the second model prediction performance evaluation value deviation obtained within the preset number of optimization times of windowing operation parameter optimization is not greater than 0, the windowing operation parameter optimization is completed; otherwise, prompt the preset personnel to intervene.

6. The method for analyzing the safety of young children based on monitoring data according to claim 1, wherein, Before performing the recognition effectiveness analysis, it also includes judging whether to perform the recognition effectiveness analysis based on the obtained total input-output duration. The specific process is as follows: When the obtained total input-output duration is within the allowable range of the total input-output duration in the database, it is determined that the pose detection model converges qualified and the recognition effectiveness analysis is performed. When the obtained total input-output duration is not within the allowable range of the total input-output duration in the database, it is determined that the pose detection model converges unqualified and the regularization term is adjusted. The total input-output duration is used to quantify the actual working time length of the pose detection model within the preset recognition period. The dangerous pose recognition result represents the recognition effectiveness data obtained by the pose detection model at the end of the preset recognition period. The recognition effectiveness data includes input response duration, output response duration, and total input-output duration. The input response duration is used to quantify the time length for the pose detection model to receive training data and successfully input it. The output response duration is used to quantify the time length from when the pose detection model starts to output the pose estimation result to the successful output end.

7. The method for analyzing the safety of young children based on monitoring data according to claim 6, wherein The specific process for adjusting the regularization term is as follows: When the obtained total input-output duration is greater than the maximum allowable total input-output duration in the database, it is determined that the pose detection model is overfitting and the regularization term of the pose detection model is increased based on the obtained regularization term increase amplitude. The regularization term increase amplitude represents the result mapped in the database of the difference between the obtained total input-output duration and the maximum allowable total input-output duration. When the total input-output duration obtained is less than the minimum allowable input-output duration in the database, it is determined that the pose detection model is underfitted, and the regularization term of the pose detection model is reduced based on the obtained regularization term reduction amplitude. The regularization term reduction amplitude represents the result mapped in the database for the difference between the minimum allowable input-output duration and the obtained input-output duration.

8. The method for analyzing the safety of young children based on monitoring data according to claim 6, wherein, The specific process of performing recognition effectiveness analysis based on the dangerous pose recognition result at the end of the preset recognition period includes: When the obtained total input-output duration is within the allowable range of the input-output duration in the database, the difference between the obtained input response duration and the reference input response duration in the database is compensated by the input response duration compensation value to obtain the input response duration score. The difference between the obtained output response duration and the reference output response duration in the database is compensated by the output response duration compensation value to obtain the output response duration score. The difference between the obtained input-output response duration and the reference input-output response duration in the database is compensated by the input-output response duration compensation value to obtain the input-output response duration score. The obtained input response duration score, output response duration score, and input-output response duration score are coupled to obtain the recognition effectiveness interference value. The recognition effectiveness interference value represents the quantitative data of the influence degree of the input response duration, output response duration, and input-output response duration on the recognition effectiveness of the pose detection model.

9. The method for analyzing the safety of young children based on monitoring data according to claim 8, wherein, The specific process of model learning rate optimization is as follows: The obtained recognition effectiveness interference value deviation is mapped in the database to obtain the initial learning rate reduction amplitude of the pose detection model. The recognition effectiveness interference value deviation is used to quantify the difference degree between the obtained recognition effectiveness interference value and the preset recognition effectiveness interference value in the database. The obtained initial learning rate reduction amplitude and the preset initial learning rate reduction amplitude are jointly input into the linear regression algorithm of the pose detection model to output the actual initial learning rate reduction amplitude. The actual initial learning rate reduction amplitude is used to ensure that the pose detection model improves the response rate to the feature data set under the condition of qualified convergence. If the recognition effectiveness interference value deviation obtained after the initial learning rate is reduced a preset number of times is not greater than 0, the model learning rate optimization is completed; otherwise, the preset personnel are prompted to intervene.

10. The child safety analysis system based on monitoring data is characterized in that, Including: A data division module, a recognition effectiveness analysis module, and a detection image generation module; Among them, the data division module is used to divide the formed feature data set through the skeleton monitoring data to obtain a data division result. The skeleton monitoring data is used to visualize the toddler skeleton information in the target monitoring area. The data division is used to generate a multi-dimensional data list from the result of segmenting the skeleton monitoring data. The recognition effectiveness analysis module is used to determine whether to perform dangerous pose recognition based on the data partitioning result at the end of the preset partitioning period. If so, it performs recognition effectiveness analysis based on the dangerous pose recognition result at the end of the preset recognition period to obtain the recognition effectiveness analysis result. Otherwise, it optimizes the windowing operation parameters. The recognition effectiveness analysis is used to quantify the effectiveness of the constructed pose detection model in performing dangerous pose recognition on the skeleton monitoring data. The windowing operation parameter optimization means adjusting the windowing operation parameters to improve the prediction performance of the pose detection model; The detection image generation module is used to determine whether to generate a behavior detection image based on the recognition effectiveness analysis result. If so, it visualizes the generated behavior detection image. Otherwise, it optimizes the model learning rate. The model learning rate optimization means adjusting the initial learning rate to improve the response rate of the pose detection model.

Citation Information

Patent Citations

  • A method, apparatus, and storage medium for analyzing process behavior.

    CN111143182B

  • Patient behavior monitoring method and system based on big data analysis

    CN119397445A

  • Sitting posture recognition method and system based on deep learning

    CN116645721A

  • Dangerous human body behavior recognition analysis early warning monitoring system and method

    CN118262410A

  • Old people safety monitoring method and system

    CN118887734A