Equipment security monitoring method based on artificial intelligence
Through deep learning and object detection algorithms, we can identify climbable objects and human postures on campus, monitor and send early warnings in real time, solving the problem of inability to prevent student climbing accidents in the existing technology and improving campus safety.
Patent Information
- Application Number
- CN202510366974.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-26
- Publication Date
- 2025-07-11
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing technology cannot effectively prevent accidental fall accidents caused by climbing objects by students on campus by relying solely on face monitoring, and poses safety hazards.
The convolutional neural network and object detection algorithm based on deep learning are used to identify climbable objects on campus and judge the human posture. The camera and server system monitor in real time and send early warning information when climbing behavior occurs, including high-resolution image data.
Accurately identify climbing behaviors and send early warnings in a timely manner to reduce accidents, provide clear image evidence to support follow-up processing, and ensure campus safety.
Smart Images

Figure CN120298968A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of security monitoring, and particularly to a device security monitoring method based on artificial intelligence. Background Art
[0002] With the continuous development of society and the increasing attention to education, campus security has become a focus issue of concern to all sectors of society. As the main place for students to study and live, the security situation of the campus is directly related to the physical and mental health and academic development of students.
[0003] Publication No. CN108922114B discloses a security monitoring method and system. A camera sends the data of a preset area collected in real time to a server; the server analyzes the data of the preset area to obtain the faces included in the preset area and the number of times and / or the appearance time length of each face. When the number of times and / or the appearance time length of the face exceeds a first monitoring threshold, a warning message is sent to a user terminal.
[0004] However, the above application still has the following problems: The method of Publication No. CN108922114B mainly focuses on monitoring the appearance of faces in a specific area. However, some climbable objects on campus may trigger students' climbing behaviors, which may lead to accidental falling accidents. Students are naturally curious and active, and may climb climbable objects such as buildings, walls, and trees on campus due to exploring the surrounding environment. Without effective monitoring and prevention measures, once a fall occurs, it will cause serious physical harm to students and even endanger their lives. This will not only cause great pain and losses to students and their families, but also have a negative impact on the teaching order and reputation of the school. Therefore, relying solely on face monitoring is far from enough. Summary of the Invention
[0005] In view of the deficiencies of the prior art, the present invention proposes a device security monitoring method based on artificial intelligence.
[0006] A device security monitoring method based on artificial intelligence proposed by the present invention is applied to a system including a camera and a server. The method includes the following steps:
[0007] The camera is installed in the campus coverage area, and the camera sends the image data collected in real time to the server;
[0008] The server analyzes the image data, identifies the climbable objects in the campus coverage area, and performs face recognition and human body posture recognition;
[0009] The server determines whether the human body posture appearing on the climbable object exceeds the anti-climbing threshold;
[0010] If the human body posture appearing on the climbable object exceeds the anti-climbing threshold, the server sends a warning message to the user's user terminal, where the warning message includes the human body posture recognition image data and the face recognition image data.
[0011] Preferably, the server determines whether the human body posture appearing on the climbable object exceeds the anti-climbing threshold as follows:
[0012] Use the convolutional neural network in deep learning and train it with a large number of images marked with climbable objects. The trained model is used to identify the climbable objects on campus;
[0013] For each identified climbable object, confirm the position of the climbable object in the image through the object detection algorithm.
[0014] Use the human body posture estimation model based on the convolutional neural network to obtain the position of the human foot nodes in the image;
[0015] If the position of the human foot nodes in the image is on the surface of the climbable object and the height is higher than the set threshold, the system determines it as a climbing behavior.
[0016] Preferably, when sending the warning message, perform high-resolution picture generation processing on the original human body posture recognition image data and the face recognition image data, and then send the generated high-resolution pictures along with the warning message.
[0017] Preferably, for the original human body posture recognition image data and the face recognition image data:
[0018] Divide the video data in the original human body posture recognition image data and the face recognition image data into several sub-video data, and generate a high-resolution picture for each sub-video data to be sent along with the warning message.
[0019] Preferably, perform high-resolution picture generation processing as follows:
[0020] Suppose there are N frames of video data in the subset. For the i-th frame and the j-th frame, use I i (x, y) and I j (x, y) to represent the intensity values of these two frames at the pixel coordinates (x, y) respectively; where, i≠j;
[0021] The obtained high-resolution frame, that is, the high-resolution picture, has an intensity value of I HR (x, y) at the pixel coordinates (x, y):
[0022]
[0023] Among them, G k (x, y) represents the gradient amplitude of the k-th frame at the pixel coordinates (x, y);
[0024] It represents the intensity value of the k-th frame after alignment at the pixel coordinates (x, y).
[0025] Preferably, within a preset time period, the server calculates the number of people in the camera monitoring area, and the server determines whether the number of people exceeds the set threshold of the number of people.
[0026] If the number of people exceeds the set threshold of the number of people, the server sends a warning message to the user's user terminal, where the warning message includes the information of the number of people data.
[0027] Preferably, the steps for counting the number of people are as follows:
[0028] Within a preset time period, the server processes the received image data frame by frame. In each frame, the number of human targets is determined through a target detection algorithm. Then, over time, the total number of human targets that appear within the preset time period is counted to obtain the information of the number of people data.
[0029] Preferably, the server stores the image data in a distributed file system, and the image data cannot be manually deleted in the distributed file system.
[0030] A terminal includes a processor and a storage medium; the storage medium is used to store instructions;
[0031] The processor is used to operate according to the instructions to execute the steps of the above-mentioned device security monitoring method based on artificial intelligence.
[0032] A computer-readable storage medium stores a computer program, and when the program is executed by a processor, it implements the steps of the above-mentioned device security monitoring method based on artificial intelligence.
[0033] In the present invention, the proposed device security monitoring method based on artificial intelligence has the following beneficial technical effects:
[0034] 1. Through the convolutional neural network and target detection algorithm of deep learning, it can accurately identify the climbable objects on campus and the postures of the human body on the climbable objects. When the foot node of the human body is on the surface of the climbable object and the height is higher than the set threshold, it is determined as a climbing behavior and a warning is issued, which is convenient for relevant personnel to respond and take measures in a timely manner to intervene, thereby effectively preventing dangerous climbing behaviors and reducing the occurrence of accidents.
[0035] 2. The camera captures image data in real time and sends it to the server, which analyzes and processes it in a timely manner. Once a climbing behavior or abnormal pedestrian flow is detected, a warning message containing relevant image data or pedestrian flow data information will be immediately sent to the user terminal, facilitating relevant personnel to respond and take timely measures for intervention to ensure campus safety.
[0036] 3. When sending the warning message, high-resolution picture generation processing is performed on the human pose recognition image data and face recognition image data. This provides clearer and more accurate image evidence for subsequent security processing. This helps relevant personnel to more clearly understand the on-site situation, such as accurately identifying the facial features and pose details of the climber, improving the efficiency and accuracy of subsequent processing.
[0037] Additional aspects and advantages of the present invention will be given in part in the following description, become apparent in part from the following description, or be learned through the practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] Figure 1 It is a flowchart of the method of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0039] Embodiments of the present invention will be described in detail below. Examples of the embodiments are shown in the drawings, where the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below by referring to the drawings are exemplary only for explaining the present invention and should not be construed as limiting the present invention.
[0040] As Figure 1 shown, a device security monitoring method based on artificial intelligence, characterized in that the method is applied to a system including a camera and a server, and the method includes the following steps:
[0041] The camera is installed in the campus coverage area, and the camera sends the image data collected in real time to the server;
[0042] The campus coverage area includes, for example, roads, playgrounds, teaching building corridor areas, and stairwells;
[0043] The server analyzes the image data, identifies climbable objects in the campus coverage area, and performs face recognition and human pose recognition;
[0044] The image data includes picture data and video data;
[0045] In an alternative embodiment, face recognition can directly adopt face recognition technology in the prior art;
[0046] The server determines whether the human pose appearing on the climbable object exceeds the anti-climbing threshold;
[0047] If the human body posture appearing on a climbable object exceeds the anti-climbing threshold, the server sends a warning message to the user's user terminal, where the warning message includes the human body posture recognition image data and the face recognition image data;
[0048] The server determines whether the human body posture appearing on a climbable object exceeds the anti-climbing threshold as follows:
[0049] Use the convolutional neural network in deep learning and train it with a large number of images labeled with climbable objects. The trained model is used to identify climbable objects on campus;
[0050] Climbable objects include trees, escalators, guardrails, parapets, tables, window sills, chairs, and rockeries;
[0051] For each identified climbable object, confirm the position of the climbable object in the image through the object detection algorithm,
[0052] The object detection algorithm is Faster-RCNN or SSD;
[0053] The core idea of the Faster-RCNN algorithm is to introduce Region Proposal Networks to generate candidate regions, thus avoiding the time-consuming and inefficient sliding window and Selective Search algorithms in traditional methods;
[0054] SSD, namely Single Shot MultiBox Detector, is an object detection algorithm based on a deep convolutional neural network;
[0055] Use a human body posture estimation model based on a convolutional neural network to obtain the position of the human foot nodes in the image;
[0056] If the position of the human foot nodes in the image is on the surface of a climbable object and the height is higher than the set threshold, the system determines it as a climbing behavior;
[0057] Through the convolutional neural network of deep learning and the object detection algorithm, accurately identify the climbable objects on campus and the postures of the human body on the climbable objects. When the human foot nodes are on the surface of a climbable object and the height is higher than the set threshold, it is determined as a climbing behavior and a warning is issued, facilitating relevant personnel to make a response and take measures in a timely manner for intervention, thereby effectively preventing dangerous climbing behaviors and reducing the occurrence of accidents.
[0058] In an optional embodiment, the set threshold for the climbing behavior is:
[0059] If the height of the human foot nodes off the ground exceeds 0.3 meters, the system determines it as a climbing behavior.
[0060] The human pose estimation model based on convolutional neural network is OpenPose or AlphaPose or PoseNet;
[0061] OpenPose is a real-time multi-person pose estimation model that can detect the poses of multiple humans simultaneously; it is constructed based on convolutional neural network;
[0062] AlphaPose is a human pose estimation model based on convolutional neural network technology, focusing on achieving high-precision single-person or multi-person pose estimation;
[0063] PoseNet is a human pose estimation model mainly used to estimate the pose of a human from a single image. It is a lightweight model that can run in resource-constrained environments such as mobile devices;
[0064] In an optional embodiment, when sending a warning message, high-resolution picture generation processing is performed on the original human pose recognition image data and face recognition image data, and then the generated high-resolution pictures are sent along with the warning message;
[0065] The camera collects image data in real time and sends it to the server, and the server analyzes and processes it in a timely manner. Once a climbing behavior or abnormal human flow is detected, a warning message can be immediately sent to the user terminal, including relevant image data or human flow data information, facilitating relevant personnel to respond and take measures to intervene in a timely manner to ensure campus safety.
[0066] In an optional embodiment, for the original human pose recognition image data and face recognition image data:
[0067] The video data in the original human pose recognition image data and face recognition image data are evenly divided into several sub-video data, and a high-resolution picture is generated for each sub-video data and sent along with the warning message.
[0068] In an optional embodiment, for high-resolution picture generation processing, SRCNN, that is, super-resolution convolutional neural network, can be used. It takes a low-resolution image as input, extracts features through convolutional layers, enhances feature representation through a non-linear mapping layer, and finally outputs a high-resolution image through a reconstruction layer.
[0069] In an optional embodiment, the high-resolution picture generation processing is as follows:
[0070] Suppose there are N frames of video data in the subset. For the i-th frame and the j-th frame, use I i (x, y) and I j (x, y) to represent the intensity values of these two frames at the pixel coordinates (x, y) respectively; where, i ≠ j;
[0071] In an alternative embodiment, i = 1, j = N, i.e., the first frame and the last frame are selected;
[0072] The obtained high-resolution frame, high-resolution picture, has an intensity value I at pixel coordinates (x, y) HR (x, y):
[0073]
[0074] where,
[0075] To minimize the difference between two frames, a motion vector field M is found ij (x, y):
[0076] M ij (x,y) = (u ij (x,y), v ij (x,y));
[0077] Using the optical flow equation based on the brightness constancy assumption, then:
[0078] I j (x + u ij (x,y), y + v ij (x,y)) - I i (x,y) ≈ 0;
[0079] Solving the optical flow equation based on the brightness constancy assumption, using the gradient descent method to minimize the energy function
[0080]
[0081] where, α is a regularization parameter used to prevent overfitting of the motion vector field, represents the gradient operator;
[0082] By iteratively updating the motion vectors until the energy function E motion (M ij ) converges, the motion vector field M ij (x, y) is obtained;
[0083] For the obtained M ij (x, y), the frames are aligned. Let be the image obtained by aligning the j-th frame to the i-th frame according to the motion vector field M ij (x, y). Then the alignment equation is:
[0084]
[0085] Suppose K frames including the original frames have been aligned. Use Denote the intensity value of the k-th frame after alignment at pixel coordinates (x, y);
[0086] Let the high-resolution frame be I HR (x, y), and the weight function be w k (x, y), then:
[0087]
[0088] Let G k (x, y) represent the gradient magnitude of the k-th frame at pixel coordinates (x, y), then the weight function w k (x, y) can be set as:
[0089]
[0090] Denote starting from m = 1 and accumulating until m = K;
[0091] Specifically,
[0092] In the weight function, plays a role in normalization. By dividing the gradient magnitude of the k-th frame by one can obtain the relative importance of the k-th frame among all frames, that is, the weight;
[0093] where G k (x, y) can be obtained by calculating the gradients in the horizontal and vertical directions and taking the square root of the sum of squares;
[0094] The horizontal gradient G x,k (x, y) is calculated using the central difference formula as follows:
[0095]
[0096] This means that in the horizontal direction, the gradient change in the horizontal direction is approximately obtained by subtracting the pixel value on the left side from the pixel value on the right side of the current pixel point;
[0097] The vertical gradient G x,k (x, y) is calculated using the central difference formula as follows:
[0098]
[0099] That is, the gradient change in the vertical direction is obtained by subtracting the pixel value above from the pixel value below the current pixel point;
[0100] Based on the calculated gradients in the horizontal and vertical directions, G k (x, y) is obtained by taking the square root of the sum of squares, and the formula is as follows:
[0101]
[0102] In this way, the gradient magnitude at each pixel of each frame is obtained, which comprehensively reflects the gradient change intensity of the pixel in the horizontal and vertical directions, that is, a quantitative representation of the richness of details of the image at this point;
[0103] Now, when we process video data, for each pixel of each frame, we operate according to the above steps of calculating the gradients in the horizontal and vertical directions, taking the square root of the sum of squares to obtain the gradient magnitude, and then substitute the obtained gradient magnitude into the high-resolution frame reconstruction equation, and we can reconstruct a high-resolution frame with more details according to the gradient situation at each pixel of each frame.
[0104] Through such a process, in the multi-frame super-resolution algorithm, the detail information of each frame image can be more effectively utilized, thereby improving the quality of the reconstructed high-resolution frame.
[0105] Substitute the weight function w k (x, y) into the high-resolution frame I HR (x, y), and we can get:
[0106]
[0107] It means starting from k = 1 and accumulating until k = K;
[0108] For example It means adding all terms from k = 1 to K;
[0109] In this way, the operation of substituting the weight function into the high-resolution frame reconstruction equation is completed.
[0110] In this way, when reconstructing the high-resolution frame, corresponding weights will be assigned according to the gradient information of each frame. The frame with a larger gradient will have a higher weight in the reconstruction process. The frame with a larger gradient is the frame with more details, so it is more conducive to improving the richness of details of the reconstructed high-resolution frame.
[0111] When sending early warning information, high-resolution image generation processing is performed on the human body pose recognition image data and the face recognition image data. Provide clearer and more accurate image evidence for subsequent security processing. This helps relevant personnel to more clearly understand the on-site situation, such as accurately identifying the facial features and pose details of the climber, and improving the efficiency and accuracy of subsequent processing.
[0112] In an optional embodiment, within a preset time period, the server calculates the number of people flow in the camera monitoring area, and the server determines whether the number of people flow exceeds the set threshold of the number of people flow;
[0113] If the number of people flow exceeds the set threshold of the number of people flow, the server sends a warning message to the user's user terminal, where the warning message includes the information of the number of people flow data;
[0114] The steps for counting the number of people flow are as follows: within a preset time period, the server processes the received image data frame by frame. In each frame, the number of human targets is determined through a target detection algorithm. Then, as time goes by, the total number of human targets that appear within the preset time period is counted to obtain the information of the number of people flow data;
[0115] Sending a warning message in a timely manner when the number of people flow exceeds the set threshold is convenient for school administrators to reasonably arrange campus resources, such as adjusting teaching activity arrangements, diverting the flow of people, etc., to prevent safety problems caused by overcrowding of people. At the same time, it also helps to optimize the use and management of campus facilities subsequently.
[0116] In an alternative embodiment, the server stores the image data in a distributed file system, and the image data cannot be manually deleted in the distributed file system:
[0117] The distributed file system is Ceph. Ceph is an open-source distributed file system. In Ceph, the image data can be protected from being manually deleted by setting the read-only attribute of the image data, and the image data that exceeds the set time threshold can be automatically deleted by setting an automatic deletion program.
[0118] This ensures the integrity and security of the security monitoring data, preventing the data from being maliciously tampered with or lost. When a retrospective investigation is needed, it can provide reliable historical data support, provide long-term data guarantee for campus security management, help analyze and summarize campus security incidents, and further improve security measures.
[0119] A terminal includes a processor and a storage medium; the storage medium is used to store instructions;
[0120] The processor is used to operate according to the instructions to execute the steps of the above-mentioned device security monitoring method based on artificial intelligence.
[0121] A computer-readable storage medium stores a computer program, and when the program is executed by a processor, it implements the steps of the above-mentioned device security monitoring method based on artificial intelligence.
[0122] This application can be applied to schools to protect the safety of students.
[0123] Meanwhile, the content not described in detail in this specification belongs to the prior art well-known to those skilled in the art.
[0124] In the embodiments provided by the present invention, it should be understood that the disclosed system or method can be implemented in other ways. For example, the above-described invention embodiments are merely illustrative. For example, the division of modules is only a logical function division, and there may be other division methods in actual implementation.
[0125] The modules described as separate components may or may not be physically separated. The components shown as modules may or may not be physical modules. They can be located in one place or distributed to multiple network modules. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0126] In addition, the functional modules in various embodiments of the present invention can be integrated into a processing module, or each module can exist physically alone, or two or more modules can be integrated into one module. The above-integrated modules can be implemented in the form of hardware or in the form of a combination of hardware and software functional modules.
[0127] For those skilled in the operation and maintenance field, it is obvious that the present invention is not limited to the details of the above exemplary embodiments, and can be implemented in other specific forms without departing from the basic characteristics of the present invention.
[0128] The above is only a preferred specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention, according to the technical solution and inventive concept of the present invention, makes equivalent replacements or changes, and should be covered by the protection scope of the present invention.
Claims
1. An artificial intelligence-based device security monitoring method, characterized in that, The method is applied to a system including a camera and a server. The method includes the following steps: The camera is installed in the campus coverage area, and the camera sends the real-time collected image data to the server; The server analyzes the image data, identifies the climbable objects in the campus coverage area, and performs face recognition and human pose recognition; The server determines whether the human pose appearing on the climbable object exceeds the anti-climbing threshold; If the human pose appearing on the climbable object exceeds the anti-climbing threshold, the server sends a warning message to the user's user terminal. Among them, the warning message includes the human pose recognition image data and the face recognition image data.
2. The method for device security monitoring based on artificial intelligence according to claim 1, wherein, The server determines whether the human pose appearing on the climbable object exceeds the anti-climbing threshold as follows: Use the convolutional neural network in deep learning to train with a large number of images marked with climbable objects. The trained model is used to identify the climbable objects on the campus; For each identified climbable object, confirm the position of the climbable object in the image through the object detection algorithm, Use the human pose estimation model based on the convolutional neural network to obtain the position of the human foot node in the image; If the position of the human foot node in the image is on the surface of the climbable object and the height is higher than the set threshold, the system determines it as a climbing behavior.
3. The method for device security monitoring based on artificial intelligence according to claim 1, characterized in that When sending the warning message, perform high-resolution picture generation processing on the original human pose recognition image data and face recognition image data, and then send the generated high-resolution pictures along with the warning message.
4. The method for device security monitoring based on artificial intelligence according to claim 3, wherein, For the original human pose recognition image data and face recognition image data: Divide the video data in the original human pose recognition image data and face recognition image data into several sub-video data, and generate a high-resolution picture for each sub-video data to be sent along with the warning message.
5. The method for device security monitoring based on artificial intelligence according to claim 3 or 4, characterized in that Perform high-resolution picture generation processing as follows: Suppose there are N frames of video data in the subset. For the i-th frame and the j-th frame, use I i (x, y) and I j (x, y) to represent the intensity values of these two frames at the pixel coordinates (x, y) respectively; where i ≠ j; The resulting high-resolution frame, i.e., the resulting high-resolution image, has an intensity value of I at pixel coordinates (x, y). HR (x, y): Among them, G k (x, y) represents the gradient magnitude at the pixel coordinates (x, y) in the k-th frame; Indicates the intensity value of the k-th frame after alignment at the pixel coordinates (x, y).
6. The method for device security monitoring based on artificial intelligence according to claim 1, wherein, Within a preset time period, the server calculates the number of people in the camera monitoring area, and the server determines whether the number of people exceeds the set threshold of the number of people; If the number of people exceeds the set threshold of the number of people, the server sends a warning message to the user's user terminal. Among them, the warning message includes the number of people data information.
7. The method for device security monitoring based on artificial intelligence according to claim 6, wherein The steps for counting the number of people are: Within a preset time period, the server processes the received image data frame by frame. In each frame, determine the number of human targets through the object detection algorithm, and then, as time goes by, count the total number of human targets that appear within the preset time period to obtain the number of people data information.
8. The method for device security monitoring based on artificial intelligence according to claim 1, wherein The server stores the image data in the distributed file system, and the image data cannot be manually deleted in the distributed file system.
9. A terminal, comprising a processor and a storage medium; characterized in that: The storage medium is used to store instructions; The processor is used to operate according to the instructions to execute the steps of the artificial intelligence-based device security monitoring method according to any one of claims 1-8.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the steps of the artificial intelligence-based device security monitoring method according to any one of claims 1-8.
Citation Information
Patent Citations
Security monitoring methods and systems
CN108922114B