Monitoring Method and Device for Video Images
Create a human posture detection model through an AI processor and analyze video images using support vector machine SVM, which solves the problems of high human resource occupation and cost in the existing video surveillance methods, and realizes high-precision and low-cost attitude monitoring for staff.
Patent Information
- Application Number
- CN202011168835.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-10-28
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2040-10-28
AI Technical Summary
The existing video surveillance methods require manual real-time viewing, occupying a large amount of human resources, and the video analysis method is costly.
An AI processor is used to create a human posture detection model, and video image analysis is performed through the support vector machine SVM, which identifies staff posture abnormalities, and automatically monitors it with the camera, CPU and alarm device.
High-precision and low-cost staff attitude monitoring is achieved, reducing human resource occupation and monitoring costs.
Smart Images

Figure CN112163566B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of video processing, and particularly to a method and device for monitoring video images. Background Art
[0002] Currently, when staff are working in a duty room, video monitoring or video analysis methods are usually adopted to monitor the staff. The video monitoring method requires manual real-time viewing within 24 hours, occupying a large amount of human resources; the video analysis method requires an artificial intelligence server for video analysis, with high costs. Summary of the Invention
[0003] In view of this, the purpose of the present invention is to provide a method and device for monitoring video images, which use an AI processor to monitor video images, determine whether the posture of the staff is abnormal, and have high recognition accuracy and low costs.
[0004] In a first aspect, an embodiment of the present invention provides a method for monitoring video images, the method comprising:
[0005] Collecting video images in a working environment;
[0006] Creating a human pose detection model;
[0007] Inputting the video images into the human pose detection model to obtain a plurality of predicted key points;
[0008] Inputting the plurality of predicted key points into a support vector machine (SVM) to obtain a classification result;
[0009] Determining the posture of the target object according to the classification result;
[0010] Determining whether the target object is normal or abnormal according to the posture of the target object.
[0011] Further, the creating of the human pose detection model includes:
[0012] Annotating the video images to obtain annotated key points;
[0013] Inputting the video images into a deep learning neural network algorithm to obtain a key point heat map;
[0014] Calculating the annotated key points in a downsampling manner to obtain the key points of the training images;
[0015] In the training images, distributing the key points of the training images onto the key point heat map in a downsampling manner through a Gaussian filtering algorithm with the annotated key points;
[0016] The key points of the training image and the predicted key points are corrected through a loss function to obtain a first correction difference value;
[0017] An initialized bias value is set, and the initialized bias value is trained through an L1 loss function to obtain a bias value;
[0018] When the bias value reaches a first preset condition, the predicted key points are corrected through the bias value to obtain corrected predicted key points;
[0019] When the first correction difference value meets a second preset condition, the correction of the key points of the training image and the predicted key points is completed, and a human pose detection model is obtained according to the corrected predicted key points.
[0020] Further, the method further includes:
[0021] Determine the body shape of the target object according to the video image;
[0022] According to the body shape of the target object, a plurality of center points are obtained;
[0023] According to the postures of the plurality of center points, each predicted key point is parameterized to obtain the offset of each predicted key point relative to the center point;
[0024] The offset of each predicted key point relative to the center point is calculated through the L1 loss function to obtain the offset of each predicted key point.
[0025] Further, the method further includes:
[0026] Select a plurality of sample points that meet a third preset condition from the plurality of center points, where each sample point corresponds to the plurality of predicted key points;
[0027] Construct a plurality of combinations according to each sample point and the plurality of predicted key points corresponding to each sample point;
[0028] Calculate the confidence corresponding to each combination;
[0029] If the confidence is equal to 1, the combination is determined to be the target object in the video image;
[0030] If the confidence is equal to 0, the combination is determined to be the background in the video image.
[0031] Further, the selecting a plurality of sample points that meet a third preset condition from the plurality of center points includes repeatedly performing the following processing until each center point is traversed:
[0032] Select any center point from the multiple center points as the current center point;
[0033] If the value of the current center point is greater than or equal to the values of other center points adjacent to the current center point, then use the current center point as the sample point.
[0034] In a second aspect, an embodiment of the present invention provides a monitoring device for video images. The device includes: a camera, a CPU, an AI processor, an alarm light, and a speaker. The camera, the AI processor, the alarm light, and the speaker are respectively connected to the CPU;
[0035] The camera is used to collect video images in the working environment;
[0036] The AI processor is used to create a human pose detection model; input the video images into the human pose detection model to obtain multiple predicted key points; input the multiple predicted key points into a support vector machine (SVM) to obtain a classification result; determine the pose of the target object according to the classification result; determine whether the target object is normal or abnormal according to the pose of the target object; in the case where the target object is abnormal, send an abnormal reminder message to the CPU;
[0037] The CPU is used to control the alarm light to flash and / or the speaker to give a voice prompt according to the abnormal reminder message.
[0038] Further, it further includes a communication module;
[0039] The communication module, which is connected to the CPU, is used to send the abnormal reminder message to a remote control system.
[0040] Further, it further includes an image acquisition module;
[0041] The image acquisition module, which is respectively connected to the camera and the CPU, is used to decode the video images to obtain decoded video images, and send the decoded video images to the CPU.
[0042] In a third aspect, an embodiment of the present invention provides an electronic device, including a memory and a processor. A computer program that can run on the processor is stored on the memory. When the processor executes the computer program, the method described above is implemented.
[0043] In a fourth aspect, an embodiment of the present invention provides a computer-readable medium having non-volatile program code executable by a processor. The program code causes the processor to execute the method described above.
[0044] An embodiment of the present invention provides a method and apparatus for monitoring video images, including: collecting video images in a working environment; creating a human pose detection model; inputting the video images into the human pose detection model to obtain a plurality of predicted key points; inputting the plurality of predicted key points into a support vector machine (SVM) to obtain a classification result; determining the pose of a target object according to the classification result; determining whether the target object is normal or abnormal according to the pose of the target object, and using an AI processor to monitor the video images to determine whether the pose of a staff member is abnormal, with high recognition accuracy and low cost.
[0045] Other features and advantages of the present invention will be described in the following specification, and in part, will be obvious from the specification, or will be understood by implementing the present invention. The objectives and other advantages of the present invention are achieved and obtained by the structures specifically pointed out in the specification, claims, and drawings.
[0046] To make the above objectives, features, and advantages of the present invention more obvious and understandable, the following specifically enumerates preferred embodiments and, in conjunction with the accompanying drawings, provides a detailed description as follows. BRIEF DESCRIPTION OF THE DRAWINGS
[0047] To more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following will briefly introduce the drawings required for the description of the specific embodiments or the prior art. Obviously, the following drawings are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0048] Figure 1 It is a flowchart of the method for monitoring video images provided in Embodiment 1 of the present invention;
[0049] Figure 2 It is a schematic diagram of predicted key points provided in Embodiment 1 of the present invention;
[0050] Figure 3 It is a schematic diagram of the apparatus for monitoring video images provided in Embodiment 2 of the present invention;
[0051] Figure 4 It is a schematic diagram of another apparatus for monitoring video images provided in Embodiment 3 of the present invention.
[0052] ICON:
[0053] 1 - Camera; 2 - CPU; 3 - AI Processor; 4 - Alarm Light; 5 - Speaker; 6 - Communication Module; 7 - Image Acquisition Module; 8 - I / O Module; 9 - Sound Module; 61 - Ethernet Module; 62 - WiFi Module; 63 - 4G Module. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0054] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings. Apparently, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the scope of protection of the present invention.
[0055] For ease of understanding of this embodiment, the embodiments of the present invention will be introduced in detail below.
[0056] Embodiment 1:
[0057] Figure 1 It is a flowchart of the monitoring method for video images provided in Embodiment 1 of the present invention.
[0058] Step S101, collect video images in the working environment;
[0059] Step S102, create a human pose detection model;
[0060] Step S103, input the video images into the human pose detection model to obtain multiple predicted key points;
[0061] Step S104, input the multiple predicted key points into an SVM (support vector machines) to obtain a classification result;
[0062] Step S105, determine the pose of the target object according to the classification result;
[0063] Step S106, determine whether the target object is normal or abnormal according to the pose of the target object.
[0064] In this embodiment, video images in the working environment are collected, a human pose detection model is constructed, the video images are used as input and input into the human pose detection model, and multiple predicted key points are output. Specifically, refer to Figure 2 ; the multiple predicted key points are used as input and input into the SVM for classification to obtain a classification result, and human pose estimation is performed according to the classification result to determine whether the target object is normal or abnormal; if the target object is abnormal, an abnormal reminder message is generated, where the target object includes but is not limited to staff members. Among them, the corresponding relationship between the predicted key points and the human body is shown in Table 1:
[0065] Number Name Number Name 0 Nose 9 Right foot 1 Right shoulder 10 Left hip 2 Right elbow 11 Left knee 3 Right hand 12 Left foot 4 Left shoulder 13 Right eye 5 Left elbow 14 Left eye 6 Left hand 15 Right ear 7 Right hip 16 Left ear 8 Right knee
[0066] Further, step S102 includes the following steps:
[0067] Step S201, annotate the video images to obtain annotated key points;
[0068] Step S202: Input the video image into the deep learning neural network algorithm to obtain the key-point heat map;
[0069] Specifically, set the video image as I, where I ∈ R W×H×3 , W is the width, H is the height, and use the deep learning neural network algorithm to predict the video image I to obtain the key-point heat map where R is the downsampling factor, set to 4; C is the number of types of preset key points, set to 17. The deep learning neural network algorithm includes a DLA (Deep Layer Aggregation) fully convolutional encoder-decoder network.
[0070] When it means that the predicted key points are detected; when it means that the background is detected.
[0071] Step S203: Calculate the key points of the training image by downsampling the labeled key points;
[0072] Specifically, when training the key-point prediction network, label the video image to obtain the labeled key points GT (Ground Truth), and the position of the labeled key points is P ∈ R 2 , and calculate the key points of the training image by downsampling the labeled key points Here, the labeled key points GT on the video image are downsampled (low-resolution processed) to obtain a 128*128 training image.
[0073] Step S204: In the training image, distribute the key points of the training image to the key-point heat map by downsampling the labeled key points through the Gaussian filtering algorithm;
[0074] Specifically, in the training image, distribute the key points of the training image to the heat map by downsampling the labeled key points GT through the Gaussian filtering algorithm. Among them, the Gaussian filtering algorithm can be known from formula (1):
[0075]
[0076] where σ p is the standard deviation adapted to the target scale, is the x-axis coordinate of the key point of the training image, is the y-axis coordinate of the key point of the training image.
[0077] Step S205: Correct the key points of the training image and the predicted key points through the loss function to obtain the first correction difference;
[0078] Specifically, the key points of the training image and the predicted key points are corrected through a loss function to obtain a first correction difference. The second preset condition is: setting the number of training times. If the first correction difference tends to be stable within the number of training times, the correction of the key points of the training image and the predicted key points is completed.
[0079] Among them, the loss function refers to formula (2):
[0080]
[0081] Among them, both α and β are hyperparameters of the loss function, and N is the number of predicted key points of the video image I.
[0082] Step S206, set an initial bias value, and train the initial bias value through the L1 loss function to obtain a bias value;
[0083] Step S207, when the bias value reaches the first preset condition, correct the predicted key points through the bias value to obtain corrected predicted key points;
[0084] Specifically, since the downsampling method is adopted for the video image, there will be a certain error. Therefore, a bias value is set to compensate for the predicted key points through the bias value.
[0085] Set an initial bias value, and train the initial bias value through the L1 loss function to obtain a bias value. The first preset condition is: setting the number of correction times. When the bias value tends to be stable within the number of correction times, it means that the bias value meets the accuracy requirements. At this time, use the bias value to correct the predicted key points to obtain corrected predicted key points. Specifically, refer to formula (3):
[0086]
[0087] Among them, is the predicted value, is the value after downsampling the center point of the video image, P is the position of the center point on the video image, is the position of the key point of the training image.
[0088] Step S208, when the first correction difference meets the second preset condition, the correction of the key points of the training image and the predicted key points is completed, and a human pose detection model is obtained according to the corrected predicted key points.
[0089] Furthermore, the method further includes the following steps:
[0090] Step S401, determine the body shape of the target object according to the video image;
[0091] Step S402: Obtain multiple center points according to the body shape of the target object;
[0092] Step S403: Parameterize each predicted key point according to the poses of the multiple center points to obtain the offset of each predicted key point relative to the center point;
[0093] Step S404: Calculate the offset of each predicted key point through the L1 loss function to obtain the offset of each predicted key point.
[0094] Specifically, for predicting key point detection based on the center point, let the pose of the center point be k×2-dimensional (k is the number of human body predicted key points), and then parameterize each predicted key point (the point corresponding to the joint point) to obtain the offset of each predicted key point relative to the center point, and then directly regress the offset of each predicted key point (in pixel units) through the L1 loss function
[0095] To improve the predicted key points, a bottom-up multi-person pose estimation algorithm is adopted to further estimate the heat maps of k human body key points Find the nearest initial prediction value on the key point heat map, and then use the offset of the predicted key point as a clue to assign the nearest person to each predicted key point.
[0096] Let be the detected center point, and the regression result of the key point is for j∈1...k; obtain the predicted key points through the key point heat map, referring to formula (4):
[0097]
[0098] Then match each regression position l j with the key point heat map, and then assign it to the closest human body.
[0099] Furthermore, the method further includes the following steps:
[0100] Step S501: Select multiple sample points that meet the third preset condition from the multiple center points, where each sample point corresponds to multiple predicted key points;
[0101] Step S502: Construct multiple combinations according to each sample point and the multiple predicted key points corresponding to each sample point;
[0102] Step S503: Calculate the confidence corresponding to each combination;
[0103] Step S504: If the confidence is equal to 1, determine the combination as the target object in the video image;
[0104] Step S505, if the confidence level is equal to 0, determine the combination as the background in the video image.
[0105] Further, step S501 includes the following steps: repeatedly execute the following process until each center point is traversed:
[0106] Step S601, select any center point from multiple center points as the current center point;
[0107] Step S602, if the value of the current center point is greater than or equal to the values of other center points adjacent to the current center point, then use the current center point as a sample point.
[0108] Here, the value of the current center point is the pixel value of the current center point. If the pixel value of the current center point is greater than or equal to the pixel values of other center points adjacent to the current center point, then use the current center point as a sample point. The other center points adjacent to the current center point are the eight adjacent points around the current center point. Multiple sample points can be selected by using the 3×3 MaxPool method, and the number of multiple sample points can be 100.
[0109] Specifically, select multiple sample points that meet the third preset condition from multiple center points, where each sample point corresponds to multiple predicted key points; construct multiple combinations according to each sample point and the multiple predicted key points corresponding to each sample point; calculate the confidence level corresponding to each combination If the confidence level is equal to 1, determine the combination as the target object in the video image; if the confidence level is equal to 0, determine the combination as the background in the video image.
[0110] The embodiment of the present invention provides a monitoring method for video images, including: collecting video images in the working environment; creating a human body pose detection model; inputting the video images into the human body pose detection model to obtain multiple predicted key points; inputting the multiple predicted key points into a support vector machine (SVM) to obtain a classification result; determining the pose of the target object according to the classification result; determining whether the target object is normal or abnormal according to the pose of the target object, and using an AI processor to monitor the video images to determine whether the pose of the staff is abnormal, with high recognition accuracy and low cost.
[0111] Embodiment 2:
[0112] Figure 3 It is a schematic diagram of a monitoring device for video images provided by the second embodiment of the present invention.
[0113] Refer to Figure 3 , the device includes: a camera 1, a CPU 2, an AI processor 3, an alarm lamp 4, and a speaker 5. The camera 1, the AI processor 3, the alarm lamp 4, and the speaker 5 are respectively connected to the CPU 2;
[0114] Camera 1, for collecting video images in the working environment;
[0115] AI processor 3, for creating a human pose detection model; inputting the video image into the human pose detection model to obtain multiple predicted key points; inputting the multiple predicted key points into a support vector machine (SVM) to obtain a classification result; determining the pose of the target object according to the classification result; determining whether the target object is normal or abnormal according to the pose of the target object; in the case where the target object is abnormal, sending an abnormal reminder message to CPU 2;
[0116] CPU 2, for controlling the alarm lamp 4 to flash and / or the speaker 5 to give a voice prompt according to the abnormal reminder message.
[0117] Further, it further includes a communication module 6;
[0118] Communication module 6, connected to CPU 2, for sending the abnormal reminder message to a remote control system.
[0119] Embodiment 3:
[0120] Figure 4 It is a schematic diagram of another video image monitoring device provided by the embodiment 3 of the present invention.
[0121] Refer to Figure 4 As shown in the figure, the device includes a camera 1, a CPU (Central Processing Unit) 2, an AI (Artificial Intelligence) processor 3, an alarm lamp 4, a communication module 6 and a speaker 5. The camera 1, the AI processor 3, the alarm lamp 4 and the speaker 5 are respectively connected to the CPU 2; it further includes an image acquisition module 7, an I / O module 8 and a sound module 9; the communication module 6 includes an Ethernet module 61, a WiFi (Wireless Fidelity) module 62 and a 4G (the 4th Generation) module 63. Among them, the Ethernet module 61 is connected to an Ethernet interface, the WiFi module 62 is connected to a WiFi antenna, and the 4G module 63 is connected to a 4G antenna;
[0122] The image acquisition module 7 is respectively connected to the camera 1 and the CPU 2, the I / O module 8 is respectively connected to the CPU 2 and the alarm lamp 4, and the sound module 9 is respectively connected to the CPU 2 and the speaker 5.
[0123] The image acquisition module 7 is used for decoding the video image to obtain the decoded video image, and sending the decoded video image to the CPU 2;
[0124] The I / O module 8 is used to trigger the flashing of the alarm light 4;
[0125] The sound module 9 is used to drive the speaker 5 for voice prompts.
[0126] In addition, when sending the abnormal reminder information to the remote control system, it can be sent through the Ethernet module 61, the WiFi module 62, and the 4G module 63. Among them, the remote control system includes, but is not limited to, the alarm system and the security system.
[0127] When using the Ethernet module 61 for sending, the abnormal reminder information can be sent to the remote control system through the Ethernet interface;
[0128] When using the WiFi module 62 for sending, the abnormal reminder information can be sent to the remote control system through the WiFi antenna.
[0129] When using the 4G module 63 for sending, the abnormal reminder information can be sent to the remote control system through the 4G antenna.
[0130] A monitoring system for video images includes the monitoring device for video images as described above.
[0131] The embodiments of the present invention provide a monitoring device and system for video images, including: a camera, a CPU, an AI processor, an alarm light, and a speaker. The camera, the AI processor, the alarm light, and the speaker are respectively connected to the CPU; the camera is used to collect video images of the working environment; the AI processor is used to create a human pose detection model; input the video images into the human pose detection model to obtain multiple predicted key points; input the multiple predicted key points into the SVM to obtain a classification result; determine the pose of the target object according to the classification result; determine whether the target object is normal or abnormal according to the pose of the target object; in the case where the target object is abnormal, send an abnormal reminder information to the CPU; the CPU is used to control the flashing of the alarm light and / or the speaker for voice prompts according to the abnormal reminder information. By using the AI processor to monitor the video images and determine whether the staff's pose is abnormal, the recognition accuracy is high and the cost is low.
[0132] The embodiments of the present invention also provide an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, it implements the steps of the video image monitoring method provided in the above embodiments.
[0133] The embodiments of the present invention also provide a computer-readable medium having non-volatile program code executable by a processor. A computer program is stored on the computer-readable medium. When the computer program is run by the processor, it executes the steps of the video image monitoring method in the above embodiments.
[0134] The computer program product provided by the embodiments of the present invention includes a computer-readable storage medium storing program code. The instructions included in the program code can be used to execute the methods described in the foregoing method embodiments. For specific implementation, reference can be made to the method embodiments and will not be elaborated herein.
[0135] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the systems and devices described above can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.
[0136] In addition, in the description of the embodiments of the present invention, unless otherwise clearly specified and limited, the terms "installation", "connection", and "connection" should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be directly connected or indirectly connected through an intermediate medium, and it can be the communication inside two components. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific situations.
[0137] If the above functions are implemented in the form of software function units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the embodiments of the present invention. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs that can store program code.
[0138] In the description of the present invention, it should be noted that the orientation or positional relationship indicated by the terms "center", "upper", "lower", "left", "right", "vertical", "horizontal", "inner", "outer", etc. is based on the orientation or positional relationship shown in the drawings. It is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be understood as a limitation of the present invention. In addition, the terms "first", "second", and "third" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance.
[0139] Finally, it should be noted that the above-described embodiments are only specific embodiments of the present invention, used to illustrate the technical solutions of the present invention, rather than limiting it. The protection scope of the present invention is not limited thereto. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that any technician familiar with the technical field of the present invention can still modify the technical solutions recorded in the foregoing embodiments or can easily think of changes, or perform equivalent replacements for some of the technical features; and these modifications, changes or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be covered within the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.
Claims
1. A monitoring method for video images, characterized in that, The method includes: Collecting video images in the working environment; Creating a human pose detection model; Inputting the video images into the human pose detection model to obtain multiple predicted key points; Inputting the multiple predicted key points into a support vector machine (SVM) to obtain a classification result; Determining the pose of the target object according to the classification result; Determining whether the target object is normal or abnormal according to the pose of the target object; The creating of the human pose detection model includes: Annotating the video images to obtain annotated key points; Inputting the video images with the annotated key points into a deep learning neural network algorithm to obtain key point heat maps; Calculating the key points of the training images by downsampling the annotated key points; In the training images, distributing the key points of the training images to the key point heat maps by means of Gaussian filtering algorithm with the annotated key points downsampled; Correcting the key points of the training images and the predicted key points through a loss function to obtain a first correction difference; Setting an initialized bias value, training the initialized bias value through an L1 loss function to obtain a bias value; When the bias value reaches a first preset condition, correcting the predicted key points through the bias value to obtain corrected predicted key points; When the first correction difference meets a second preset condition, the correction of the key points of the training images and the predicted key points is completed, and the human pose detection model is obtained according to the corrected predicted key points; The method further includes: Determining the body shape of the target object according to the video images; Obtaining multiple center points according to the body shape of the target object; Parameterizing each predicted key point according to the poses of the multiple center points to obtain the offset of each predicted key point relative to the center point; Calculating the offset of each predicted key point through the L1 loss function with the offset of each predicted key point relative to the center point; The method further includes: Selecting multiple sample points that meet a third preset condition from the multiple center points, where each sample point corresponds to the multiple predicted key points; Constructing multiple combinations according to each sample point and the multiple predicted key points corresponding to each sample point; Calculating the confidence level corresponding to each combination; If the confidence level is equal to 1, determining the combination as the target object in the video images; If the confidence level is equal to 0, determining the combination as the background in the video images.
2. The monitoring method of video images according to claim 1, characterized in that, The selecting of multiple sample points that meet a third preset condition from the multiple center points includes repeatedly performing the following process until each center point is traversed: Selecting any center point from the multiple center points as the current center point; If the value of the current center point is greater than or equal to the values of other center points adjacent to the current center point, taking the current center point as the sample point.
3. A monitoring device for video images, characterized in that, When the device runs, it executes the monitoring method of the video image described in claim 1 or 2. The device includes: a camera, a CPU, an AI processor, an alarm light, and a speaker. The camera, the AI processor, the alarm light, and the speaker are respectively connected to the CPU; The camera is used to collect video images in the working environment; The AI processor is used to create a human pose detection model; input the video image into the human pose detection model to obtain a plurality of predicted key points; input the plurality of predicted key points into a support vector machine (SVM) to obtain a classification result; determine the pose of the target object according to the classification result; determine whether the target object is normal or abnormal according to the pose of the target object; in the case where the target object is abnormal, send an abnormal reminder message to the CPU; The CPU is used to control the alarm light to flash and / or the speaker to give a voice prompt according to the abnormal reminder message.
4. The monitoring device for video images according to claim 3, characterized in that, It further includes a communication module; The communication module, which is connected to the CPU, is used to send the abnormal reminder message to a remote control system.
5. The monitoring device for video images according to claim 3, characterized in that, It further includes an image acquisition module; The image acquisition module, which is respectively connected to the camera and the CPU, is used to decode the video image to obtain a decoded video image, and send the decoded video image to the CPU.
6. An electronic device, comprising a memory and a processor, wherein a computer program capable of running on the processor is stored on the memory, characterized in that When the processor executes the computer program, it implements the method described in claim 1 or 2 above.
7. A computer-readable medium having non-volatile program code executable by a processor, characterized in that, The program code causes the processor to execute the method described in claim 1 or 2.
Citation Information
Patent Citations
Artificial intelligence early warning system
CN109447048A
Method and device for identifying figure position in image, computer equipment and storage medium
CN110502986A