Video Surveillance Method and System for Visual Operation
By analyzing the time series and timing differences in facial information acquisition efficiency in monitoring videos, the Turbopixels algorithm is used for clustering and labeling, which solves the problem that static or small amplitude actions are difficult to identify in the video, and improves the accuracy of abnormal behavior detection and the practicality of visual operations.
Patent Information
- Application Number
- CN202510112334.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-24
- Publication Date
- 2025-08-01
- Estimated Expiration
- 2045-01-24
AI Technical Summary
The prior art is difficult to effectively identify and mark static or small amplitude abnormal behaviors in surveillance videos, such as napping, dazedness, etc., resulting in a low recognition accuracy rate.
By obtaining the efficiency-time series of facial information acquisition of target objects, the Turbopixels algorithm is used for clustering, the behavioral state similarity and timing differences are analyzed, the possibility of abnormal behavior is calculated, and the abnormal behavior is marked based on the marking weight.
It improves the accuracy of abnormal behavior detection of static or small amplitude actions, avoids misjudgment of normal behavior and abnormal behavior, and enhances the practicality of the visual operation monitoring system.
Smart Images

Figure CN119600547B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of digital information transmission, and particularly relates to a video monitoring method and system for visual operation. Background Art
[0002] Visual operation refers to the process of enabling users to intuitively interact with computer systems, applications, or data through a GUI (Graphical User Interface) or other visual representation forms. Video data is collected through a high-definition camera, and based on the actual monitoring requirement direction, a corresponding GUI interface is designed to provide tools such as thumbnails and timelines to facilitate users to quickly locate and retrieve videos. For the current visual operations added during the processing of surveillance videos, it includes the ability to reverse-play the video using the timeline to obtain relevant monitoring information at specific moments.
[0003] For example, in a classroom, the movement amplitude of students is relatively small, and the target users of the video surveillance using visual operations are usually parents of students. Most of them want to understand the historical listening status of students in the classroom. Existing image detection technologies can analyze and capture the actions of people or objects in video images. However, for static or small-amplitude actions, the changes in video images over a continuous period of time are small, so the similarity is high, resulting in a low recognition accuracy, that is, abnormal behaviors that may exist, such as dozing and being in a daze, cannot be marked. Summary of the Invention
[0004] In order to solve the technical problem that it is difficult to identify and mark static and small-amplitude abnormal behavior actions in surveillance videos, the purpose of the present invention is to provide a video monitoring method and system for visual operation, and the specific technical solutions adopted are as follows:
[0005] The present invention provides a video monitoring method for visual operation, and the method includes:
[0006] Obtain the target video information of the target object, and use the target video information to determine the current acquisition efficiency-time series of the facial information of the target object;
[0007] Determine the behavior state similarity between each adjacent time point in the current acquisition efficiency-time series, and determine each target cluster of the current acquisition efficiency-time series according to the behavior state similarity;
[0008] Determine the timing difference between the target time period corresponding to any target cluster and the other time periods corresponding to other objects, and determine the possibility of continuous abnormal behavior of the target object in the target time period according to the timing difference;
[0009] Determine and mark the abnormal behavior of the target object according to the possibility of continuous abnormal behavior.
[0010] Further, the step of determining the similarity of behavior states corresponding to each adjacent time point in the current acquisition efficiency - time series includes:
[0011] Determine the efficiency difference of the facial information acquisition efficiency between any two adjacent time points in the current acquisition efficiency - time series;
[0012] Using the efficiency difference, calculate the similarity of behavior states between any two adjacent time points.
[0013] Further, the step of calculating the similarity of behavior states between any two adjacent time points using the efficiency difference includes:
[0014] Determine the instantaneous change rate of the facial information acquisition efficiency corresponding to each of any two adjacent time points;
[0015] Determine the difference in instantaneous change rates between any two adjacent time points;
[0016] Using the efficiency difference and the difference in instantaneous change rates, calculate the similarity of behavior states between any two adjacent time points.
[0017] Further, the step of determining each target cluster of the current acquisition efficiency - time series according to the similarity of behavior states includes:
[0018] Cluster the current acquisition efficiency - time series using the Turbopixels algorithm and the similarity of behavior states to obtain each target cluster; where the target clustering represents the behavior state of the target object.
[0019] Further, the step of determining the temporal difference between the target time period corresponding to any target cluster and the other time periods corresponding to other objects includes:
[0020] Determine the other time periods corresponding to other objects that are closest to the target time period corresponding to any target cluster;
[0021] Determine the start time interval and the end time interval between the target time period and the other time periods;
[0022] Using the start time interval and / or the end time interval, determine the temporal difference between the target time period and the other time periods.
[0023] Further, the step of determining the possibility of continuous abnormal behavior of the target object in the target time period according to the temporal difference includes:
[0024] Determine the average value of the target facial information acquisition efficiency of the target object in the target time period;
[0025] Determine the average value of the acquisition efficiency of other facial information of other objects in other time periods closest to the target time period;
[0026] Utilize the time series difference, the average value of the target facial information acquisition efficiency, and the average value of the acquisition efficiency of other facial information to calculate the likelihood of continuous abnormal behavior of the target object in the target time period.
[0027] Furthermore, the steps of determining and marking the abnormal behavior of the target object according to the likelihood of continuous abnormal behavior include:
[0028] Determine the historical acquisition efficiency - time series of the target object;
[0029] Determine the time period difference between the current acquisition efficiency - time series and the historical acquisition efficiency - time series;
[0030] Determine and mark the abnormal behavior of the target object according to the time period difference and the likelihood of continuous abnormal behavior.
[0031] Furthermore, the time period difference includes: the total number of time period differences;
[0032] The steps of determining and marking the abnormal behavior of the target object according to the time period difference and the likelihood of continuous abnormal behavior include:
[0033] Utilize the total number of time period differences to calculate the overall behavioral abnormality of the target object;
[0034] Determine and mark the abnormal behavior of the target object according to the overall behavioral abnormality and the likelihood of continuous abnormal behavior.
[0035] Furthermore, the steps of determining and marking the abnormal behavior of the target object according to the overall behavioral abnormality and the likelihood of continuous abnormal behavior include:
[0036] Calculate the abnormal behavior marking weight obtained by normalizing the product of the overall behavioral abnormality and the likelihood of continuous abnormal behavior;
[0037] Compare the abnormal behavior marking weight with the preset abnormal marking threshold to determine and mark the abnormal behavior of the target object.
[0038] The present invention also provides a video surveillance system for visualization operations, and the system is used to implement the video surveillance method for visualization operations as described in any one of the above; the system includes:
[0039] A video acquisition module, configured to acquire the target video information of the target object, and utilize the target video information to determine the current acquisition efficiency - time series of the facial information of the target object;
[0040] An information clustering module, configured to determine the similarity of behavior states corresponding to each adjacent time point in the current acquisition efficiency-time series, and determine each target cluster of the current acquisition efficiency-time series according to the similarity of behavior states;
[0041] A behavior comparison module, configured to determine the temporal difference between the target time period corresponding to any target cluster and other time periods corresponding to other objects, and determine the possibility of continuous abnormal behavior of the target object in the target time period according to the temporal difference;
[0042] A behavior marking module, configured to determine and mark the abnormal behavior of the target object according to the possibility of continuous abnormal behavior.
[0043] The present invention has the following beneficial effects:
[0044] By using image processing technology to obtain the facial information acquisition efficiency of the human object at each time node, and based on the change of the facial information acquisition efficiency in the time series, segmenting and clustering the time periods to obtain several time periods in the same state, this operation avoids the problem that the changes generated by static or small-amplitude actions in the video images for a continuous period of time are small and difficult to identify, and improves the accuracy of abnormal behavior detection;
[0045] Furthermore, by analyzing the temporal difference between the time period of the target object and the time periods of other objects, as well as the difference in the facial information acquisition efficiency within the time period, the possibility of abnormal behavior of the target object in any time period is judged; and the marking weight can be further determined in combination with the historical acquisition efficiency-time series of the student, and based on the marking weight, the rapid positioning and search of the abnormal behavior of the target object are realized. This operation solves the problem that small-amplitude actions are not easily recognized in the prior art, and avoids misjudgment between normal behavior and abnormal behavior, improving the accuracy of abnormal behavior detection and the practicality of the visual operation monitoring system. Description of the Drawings
[0046] In order to more clearly illustrate the technical solutions and advantages in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required to be used in the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained according to these drawings without creative efforts.
[0047] Figure 1 It is a step flow chart of a video monitoring method for visual operation provided by an embodiment of the present invention;
[0048] Figure 2 It is a detailed flow chart of step S3 in a video monitoring method for visual operation provided by an embodiment of the present invention;
[0049] Figure 3 The refined flowchart of step S3 in a video surveillance method for visual operation provided by another embodiment of the present invention;
[0050] Figure 4 The refined flowchart of step S4 in a video surveillance method for visual operation provided by an embodiment of the present invention;
[0051] Figure 5 The structural schematic diagram of the hardware operating environment of a video surveillance device for visual operation involved in the solution of the embodiment of the present invention;
[0052] Figure 6 The framework structural schematic diagram of a video surveillance system for visual operation involved in the solution of the embodiment of the present invention. Detailed implementation manners
[0053] In order to further elaborate on the technical means and effects adopted by the present invention to achieve the predetermined invention purpose, the following, in combination with the accompanying drawings and preferred embodiments, details the specific implementation manners, structures, features, and effects of a video surveillance method for visual operation proposed according to the present invention. In the following description, different "one embodiment" or "another embodiment" do not necessarily refer to the same embodiment. In addition, the specific features, structures, or characteristics in one or more embodiments can be combined in any suitable form.
[0054] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present invention belongs.
[0055] The following specifically describes the specific solution of a video surveillance method for visual operation provided by the present invention with reference to the accompanying drawings.
[0056] Embodiment 1:
[0057] For a video surveillance method for visual operation provided by the present invention, please refer to Figure 1 , which shows the flowchart of the steps of a video surveillance method for visual operation provided by an embodiment of the present invention.
[0058] The method includes:
[0059] Step S1, obtaining the target video information of the target object, and using the target video information to determine the current acquisition efficiency - time series of the facial information of the target object;
[0060] The target object here can be a student, an employee, or even any object whose behavior can be recognized as abnormal due to a certain degree of behavior. Here, students are taken as exemplary research objects to carry out the following various embodiments.
[0061] Install multiple cameras in the space to be monitored to collect video images (video information) within a certain range of the space. In this embodiment, the classroom monitoring scenario is taken as an example; the collected video images are transmitted to the monitoring terminal based on the Internet of Things for processing and analysis.
[0062] The behaviors involved in the scenario include those of students during class and during breaks. Usually, parents are the main users of video monitoring for visual operations and can query and locate the behaviors of a specified student in historical videos. To achieve this query and location function, it is first necessary to locate and mark different students appearing in the video images, and then identify relevant actions.
[0063] An example of the specific process of location and marking is as follows:
[0064] 1. Preprocess the video images based on existing image processing techniques, including mean value processing, Gaussian denoising, etc.
[0065] 2. Analyze the currently collected and historical video information based on semantic segmentation technology to obtain important components in the video. In this case, each component in each frame of the video image is marked separately or classified. For example, for a certain component existing in the video image at a certain time point in the historical video information, it is marked as the j component, also known as the j object. In this embodiment, it refers to any jth student representative target object in the space.
[0066] 3. Evaluate the efficiency of obtaining facial features (facial information collection efficiency) for each student. Specifically, pre-enter the front ID photos of all students in the current space (class) into the monitoring system, and compare the facial features of each target component with the template photos through existing image matching techniques. The similarity and facial information collection efficiency will change in real time with the movement of the target. Therefore, it is necessary to obtain the facial information collection efficiency of any target object in real time.
[0067] Based on the monitoring camera, the target video information of the target object in the current classroom is obtained. Based on the above positioning and recognition process, the current acquisition efficiency-time series of the facial information of the target object can be analyzed and read from the target video information. This series refers to the distribution of the facial information acquisition (obtaining) efficiency at each time point. When reflected on the coordinate axis, each time point can be the horizontal axis, and the facial information acquisition efficiency is the vertical axis. Due to the changes in the actions of the target object, the corresponding facial information acquisition efficiency at each time point is often different. Here, the facial information acquisition efficiency at each time point can refer to the record of the facial information acquisition efficiency within the time interval from the adjacent previous time point to the current time point. For example, if there are 5 localizations and recognitions of the target object's face from time point 1 second to time point 2 seconds, then the corresponding facial information acquisition efficiency is 5 times / s. This is just an example for easy understanding.
[0068] It is also possible to obtain the video information of the historical D = 10 classes, and then obtain multiple historical acquisition efficiency-time series, which is convenient for subsequent analysis and processing.
[0069] Step S2: Determine the similarity of the behavior states corresponding to each pair of adjacent time points in the current acquisition efficiency-time series, and determine each target cluster of the current acquisition efficiency-time series according to the similarity of the behavior states.
[0070] Specifically, the steps to determine the similarity of the behavior states corresponding to each pair of adjacent time points in the current acquisition efficiency-time series include:
[0071] Determine the efficiency difference of the facial information acquisition efficiency between any two adjacent time points in the current acquisition efficiency-time series.
[0072] Using the efficiency difference, calculate the similarity of the behavior states between any two adjacent time points.
[0073] More specifically, the steps to calculate the similarity of the behavior states between any two adjacent time points using the efficiency difference include:
[0074] Determine the instantaneous change rate of the facial information acquisition efficiency corresponding to each of any two adjacent time points.
[0075] Determine the difference in the instantaneous change rate between any two adjacent time points.
[0076] Using the efficiency difference and the difference in the instantaneous change rate, calculate the similarity of the behavior states between any two adjacent time points.
[0077] In the current classroom, the movement amplitude of students is relatively small. The target users of video monitoring using visual operations are usually parents of students, and most of them want to understand the listening status of students in previous classes. Existing image detection technologies can analyze and capture the movements of people or objects in video images. However, for static or small-amplitude movements, the changes in video images over a continuous period of time are small, resulting in high similarity and low recognition accuracy; that is, abnormal behaviors that may exist, such as dozing off and being in a daze, cannot be marked.
[0078] According to the analysis, the general behaviors of students in the classroom include raising the head, lowering the head, getting up, walking, etc. For large-amplitude movements such as getting up and walking, they can be marked by existing technologies; for the movements of raising and lowering the head, they have the characteristics of small change amplitude and long duration under normal circumstances. For example, during the teacher's lecture, students continuously raise their heads to look or continuously lower their heads to think and take notes, etc.; due to the particularity of classroom teaching, usually students will look in the same direction, so the efficiency of facial information collection obtained by installing a camera in the direction where the students' eyes are looking will change over time as the students' movements change. Therefore, the behaviors of students can be marked according to similar efficiency time periods.
[0079] For each object, that is, each student, obtain the facial information collection efficiency that changes over time during the historical classroom period; arrange these data on the coordinate axis in chronological order, and the vertical axis of the coordinate is the facial information collection efficiency.
[0080] Among them, the specific method for measuring the similarity of behavior states is based on the efficiency difference between the facial information collection efficiencies between two time points. For discrimination, where 、 respectively represent two adjacent time points;
[0081] It should be noted that since the process of students changing their postures is a continuous change process, the collection efficiency of facial information increases or decreases slowly during the process of movement change. To avoid dividing the time points in the gradual change process into several small intervals during the time period marking process, the efficiency difference between any moment and the previous adjacent moment is used to represent the instantaneous change rate at the corresponding moment, denoted as 、 It should be noted that b - 1 here is different from a, but is a video acquisition moment of a smaller unit between 、 The method for obtaining this instantaneous quantity is the basic method for calculating instantaneous quantities in physics; in addition, the smaller the instantaneous change rate, the more stable the current moment is in the change state.
[0082] Thus, the discriminant formula of behavior state similarity is obtained:
[0083]
[0084] In the formula, according to the above logic, the instantaneous rate of change difference is obtained respectively ( ), and the efficiency difference between it and the efficiency of facial information collection Using the Euclidean norm, we get , the smaller the value, the higher the similarity of the behavior state.
[0085] The step of determining each target cluster of the current acquisition efficiency-time series according to the similarity of the behavior state specifically includes:
[0086] The Turbopixels algorithm and behavioral state similarity are used to cluster the current acquisition efficiency-time series to obtain target clusters. The target clusters represent the behavioral states of the target objects.
[0087] According to the above-mentioned behavior state similarity discrimination method, the current student's current collection efficiency-time series is clustered. That is, a certain behavior state similarity distinction rule can be set, such as setting a behavior state similarity threshold. Data with behavior state similarity greater than or equal to the threshold, which is considered to have a high behavior state similarity, are grouped into the same cluster, and data with behavior state similarity less than the threshold, which is considered to have a low behavior state similarity, are grouped into another cluster, thereby obtaining N target clusters. Each cluster represents any behavior state that the student is in. It should be noted that in order to make the final cluster a continuous time period, the existing Turbopixels algorithm can be used to cluster the facial information collection efficiency of each marked student. This algorithm can ensure the connectivity of the final clustering result, so that the data points within each cluster are continuous in time series.
[0088] Step S3: determining the timing difference between the target time period corresponding to any target cluster and other time periods corresponding to other objects, and determining the possibility of the target object's continued abnormal behavior in the target time period based on the timing difference;
[0089] For details, please refer to Figure 2 The step of determining the time series difference between a target time period corresponding to any target cluster and other time periods corresponding to other objects includes:
[0090] Step S30, determining other time periods corresponding to other objects closest to the target time period corresponding to any target cluster;
[0091] Step S31, determining the starting time interval and the ending time interval between the target time period and other time periods;
[0092] Step S32: Determine the timing difference between the target time period and other time periods by using the start time interval and / or the end time interval.
[0093] Since some abnormal behaviors are prone to being confused with each other. For example, dozing and daydreaming are likely to be confused with the head-up state and the head-down state. According to the analysis, students' behaviors should be consistent within the same time period. For example, when the teacher is giving a lecture, students generally raise their heads uniformly. There may be individual students with their heads down during the process, but the time is short. According to the analysis, the abnormal behavior of the current student (target object) is manifested in that the duration of a certain behavior state within the corresponding time period has a relatively obvious difference from the behavior performance of other students.
[0094] Therefore, for any time period k in the N clusters obtained according to the above embodiments, obtain the start and end time points of the time period k (i.e., the start time and the end time).
[0095] Select the other time period of other students that is closest to the start time or the end time of the k-th time period of the current student j; obtain the start time interval between the other time period of other student r and the time period k of the current j-th target student, denoted as ; and obtain the end time interval between the end time point of the other time period and the end time point of the current time period k, denoted as ; the greater the above two time intervals, the higher the possibility that the current student j is in an abnormal behavior state.
[0096] Thus, obtain the timing difference of any time period k, which is also called behavioral slowness:
[0097]
[0098] In the formula, according to the above logic, the time intervals corresponding to the start and end time points can be added and averaged respectively, and thus the behavioral slowness is obtained , or the start time interval and the end time interval can be used as the timing difference and the behavioral slowness without taking the average value. The larger this value is, the higher the possibility that the current student j is in an abnormal behavior state.
[0099] Traverse all students, and M - 1 behavioral slownesses respectively corresponding to the current student j can be obtained , where M represents the total number of all students in the current space.
[0100] Specifically, please refer to Figure 3 , the steps of determining the possibility of continuous abnormal behavior of the target object in the target time period according to the timing difference include:
[0101] Step S300, determine the average value of the target facial information acquisition efficiency of the target object over the target time period;
[0102] Step S310, determine the average value of the other facial information acquisition efficiency of other objects over other time periods closest to the target time period;
[0103] Step S320, calculate the likelihood of persistent abnormal behavior of the target object over the target time period by using the time series difference, the average value of the target facial information acquisition efficiency, and the average value of the other facial information acquisition efficiency.
[0104] Based on the above embodiments, if there are obvious differences in the facial acquisition efficiency obtained by the current student j and other students within a similar time period, it indicates that the likelihood of the current student j having persistent abnormal behavior in the current time period k is higher;
[0105] Thus, obtain the average value of the other facial information acquisition efficiency of other students in the corresponding time period and the average value of the target facial information acquisition efficiency of student j in time period k ;
[0106] Combined with the behavioral slowness (time series difference) in the above embodiments, the likelihood of the current student j having persistent abnormal behavior in time period k is obtained as follows:
[0107]
[0108] In the formula, according to the above logic, obtain the difference in the average efficiency between the kth time period of the current student j and a similar time period r of another student, and represent it with an absolute value, thus obtaining , the larger this value, the greater the likelihood that the current student has abnormal behavior; at the same time, weight it using the above calculated behavioral slowness and traverse all students, thus obtaining , the larger this value, the greater the likelihood that the current student j has persistent abnormal behavior in time period k.
[0109] Step S4, determine and mark the abnormal behavior of the target object according to the likelihood of persistent abnormal behavior.
[0110] Specifically, please refer to Figure 4 , the said step S4 includes:
[0111] Step S40, determine the historical acquisition efficiency - time series of the target object;
[0112] Step S41, determine the time period difference between the current acquisition efficiency - time series and the historical acquisition efficiency - time series;
[0113] Step S42: Determine and mark the abnormal behavior of the target object according to the time period difference and the possibility of continuous abnormal behavior.
[0114] Furthermore, the time period difference includes: the difference in the total number of time periods;
[0115] The step of determining and marking the abnormal behavior of the target object according to the time period difference and the possibility of continuous abnormal behavior includes:
[0116] Utilize the difference in the total number of time periods to calculate the overall behavioral abnormality of the target object;
[0117] Determine and mark the abnormal behavior of the target object according to the overall behavioral abnormality and the possibility of continuous abnormal behavior.
[0118] Even more specifically, the step of determining and marking the abnormal behavior of the target object according to the overall behavioral abnormality and the possibility of continuous abnormal behavior includes:
[0119] Calculate the abnormal behavior marking weight obtained by normalizing the product of the overall behavioral abnormality and the possibility of continuous abnormal behavior;
[0120] Compare the abnormal behavior marking weight with the preset abnormal marking threshold to determine and mark the abnormal behavior of the target object.
[0121] Based on the operations of the above embodiments, the possibility of continuous abnormal behavior existing in any time period k is obtained. This possibility reflects the difference in the behavioral performance of any student in a local time period from that of other students. However, according to the analysis, for different students in an abnormal state, the frequency of switching behavioral states is significantly different from the frequency of historical behavioral states. To avoid misjudgment of the calculated possibility of continuous abnormal behavior and abnormal behavior, the possibility of continuous abnormal behavior should be corrected in combination with the historical behavioral state of the student;
[0122] Therefore, based on what was mentioned in the foregoing embodiments: It is also possible to obtain the video information of the historical D = 10 classes, and then obtain multiple historical acquisition efficiency - time series, which is convenient for subsequent analysis and processing.
[0123] Obtain the total number of time periods of the current student j in the current target class , that is, the total number of time periods obtained by clustering the current acquisition efficiency - time series;
[0124] Thus, obtain the overall behavioral abnormality of the j - th student in the current class:
[0125]
[0126] In the formula, That is, the total number of time periods corresponding to the historical collection efficiency-time series in other classes of the current student's history. represents the normalized range of the total number of current and historical time periods; D represents the number of history classes, that is, the historical collection efficiency minus the number of time series.
[0127] According to the above logic, the difference between the total number of time periods in the current class and the historical classes is obtained, that is, the time period difference between the current collection efficiency-time series and each historical collection efficiency-time series. , and normalize its range, thus obtaining the overall behavioral abnormality ,The larger the value is, the more abnormal the student’s behavior is in the current class;
[0128] Combined with the possibility of persistent abnormal behavior , thus the abnormal behavior marking weight of the current k-th time period is obtained:
[0129]
[0130] In the formula, the above two characteristics and Combined and normalized by the norm function, the abnormal behavior marking weight is obtained ,The larger the value is, the greater the possibility that the current student has abnormal behavior in the current time period, and the higher the weight of being marked;
[0131] A preset abnormality marking threshold can be set as needed, for example, 0.7, and all time periods greater than the threshold will be marked as abnormal;
[0132] According to the visual operation, the starting time points of these marked time periods are assigned pointers or other prominent interactive buttons. When users use the visual operation video surveillance system to monitor the history classroom behavior, they can click these interactive buttons and view the abnormal behavior of the monitored object.
[0133] The present invention uses image processing technology to obtain the facial information collection efficiency of human objects at each time node, and based on the changes in the facial information collection efficiency in the time series, it divides and clusters the time periods to obtain several time periods with the same state. This operation avoids the problem that static or small-amplitude movements have small changes in the video images over a continuous period of time and are difficult to identify, thereby improving the accuracy of abnormal behavior detection.
[0134] Further, by analyzing the temporal differences between the time periods of the target object and those of other objects, as well as the differences in the facial information collection efficiency within the time periods, the possibility of abnormal behavior of the target object in any time period is judged; and the marking weight can be further determined in combination with the historical collection efficiency-time series of the student. Based on this marking weight, the rapid positioning and search of the abnormal behavior of the target object are realized. This operation solves the problem that small movements are not easily recognized in the prior art and avoids misjudgments between normal behavior and abnormal behavior, improving the accuracy of abnormal behavior detection and the practicality of the visual operation monitoring system.
[0135] Embodiment 2:
[0136] The embodiment of the present invention also proposes a video monitoring device for visual operation. The video monitoring device for visual operation can be a data calculation and processing device such as a monitoring camera, a computer, a server, or a combination of multiple devices.
[0137] As Figure 5 shown, Figure 5 is a schematic structural diagram of the hardware operating environment of the video monitoring device for visual operation involved in the embodiment of the present invention.
[0138] As Figure 5 shown, the video monitoring device for visual operation may include: a processor 1001, such as a CPU, a network interface 1004, a user interface 1003, a memory 1005, and a communication bus 1002. Among them, the communication bus 1002 is used to realize the connection and communication between these components. The user interface 1003 may include a display (Display) and an input unit such as a control panel. Optionally, the user interface 1003 may further include a standard wired interface and a wireless interface. The network interface 1004 may optionally include a standard wired interface and a wireless interface (such as a WIFI interface). The memory 1005 may be a high-speed RAM memory or a stable memory (non-volatile memory), such as a disk memory. Optionally, the memory 1005 may also be a storage device independent of the aforementioned processor 1001. The memory 1005, as a computer storage medium, may include a video monitoring program for visual operation.
[0139] Those skilled in the art can understand that Figure 5 the hardware structure shown in
[0140] Continue to refer to Figure 5 , Figure 5The memory 1005 as a computer-readable storage medium may include an operating system, a user interface module, a network communication module, and a video surveillance program for visual operation.
[0141] In Figure 5 it, the network communication module is mainly used to connect to the server and can communicate with the server for data; while the processor 1001 can call the video surveillance program stored in the memory 1005 for visual operation and execute the steps in each of the above embodiments.
[0142] Based on the above hardware structure of the video surveillance device for visual operation, each embodiment for implementing the video surveillance method for visual operation of the present invention is realized.
[0143] In addition, the present invention also provides a video surveillance system for visual operation. Please refer to Figure 6 and the video surveillance system for visual operation includes:
[0144] A video acquisition module A10, configured to obtain target video information of a target object and determine a current acquisition efficiency-time series of the facial information of the target object using the target video information;
[0145] An information clustering module A20, configured to determine the behavioral state similarity corresponding to each adjacent time point in the current acquisition efficiency-time series and determine each target cluster of the current acquisition efficiency-time series according to the behavioral state similarity;
[0146] A behavior comparison module A30, configured to determine the timing difference between the target time period corresponding to any target cluster and other time periods corresponding to other objects, and determine the possibility of continuous abnormal behavior of the target object in the target time period according to the timing difference;
[0147] A behavior marking module A40, configured to determine and mark the abnormal behavior of the target object according to the possibility of continuous abnormal behavior.
[0148] Further, the information clustering module A20 is further configured to:
[0149] Determine the efficiency difference of the facial information acquisition efficiency between any two adjacent time points in the current acquisition efficiency-time series;
[0150] Use the efficiency difference to calculate the behavioral state similarity between any two adjacent time points.
[0151] Further, the information clustering module A20 is further configured to:
[0152] Determine the instantaneous change rate of the facial information acquisition efficiency corresponding to any two adjacent time points;
[0153] Determine the difference in the instantaneous change rate between any two adjacent time points;
[0154] Using the efficiency difference and the difference in the instantaneous change rate, calculate the similarity of the behavior states between any two adjacent time points.
[0155] Furthermore, the information clustering module A20 is further configured to:
[0156] Cluster the current acquisition efficiency-time series using the Turbopixels algorithm and the similarity of the behavior states to obtain each target cluster; wherein, the target clustering characterizes the behavior state of the target object.
[0157] Furthermore, the behavior comparison module A30 is further configured to:
[0158] Determine the other time period corresponding to the other object that is closest to the target time period corresponding to any target cluster;
[0159] Determine the start time interval and the end time interval between the target time period and the other time period;
[0160] Using the start time interval and / or the end time interval, determine the timing difference between the target time period and the other time period.
[0161] Furthermore, the behavior comparison module A30 is further configured to:
[0162] Determine the average value of the target facial information acquisition efficiency of the target object during the target time period;
[0163] Determine the average value of the other facial information acquisition efficiency of the other object during the other time period that is closest to the target time period;
[0164] Using the timing difference, the average value of the target facial information acquisition efficiency, and the average value of the other facial information acquisition efficiency, calculate the likelihood of continuous abnormal behavior of the target object during the target time period.
[0165] Furthermore, the behavior marking module A40 is further configured to:
[0166] Determine the historical acquisition efficiency-time series of the target object;
[0167] Determine the time period difference between the current acquisition efficiency-time series and the historical acquisition efficiency-time series;
[0168] Determine and mark the abnormal behavior of the target object according to the time period difference and the likelihood of continuous abnormal behavior.
[0169] Furthermore, the behavior marking module A40 is further configured to:
[0170] Steps for determining and marking abnormal behaviors of a target object based on time period differences and the possibility of continuous abnormal behaviors, including:
[0171] Using the difference in the total number of time periods, calculate the overall behavioral abnormality of the target object;
[0172] Determine and mark the abnormal behaviors of the target object based on the overall behavioral abnormality and the possibility of continuous abnormal behaviors.
[0173] Furthermore, the behavior marking module A40 is also used for:
[0174] Calculate the abnormal behavior marking weight obtained by normalizing the product of the overall behavioral abnormality and the possibility of continuous abnormal behaviors;
[0175] Compare the abnormal behavior marking weight with a preset abnormal marking threshold to determine and mark the abnormal behaviors of the target object.
[0176] The specific implementation manner of the video surveillance system for visual operation in the present invention is basically the same as each embodiment of the above-mentioned video surveillance method for visual operation, and will not be elaborated here.
[0177] In addition, the present invention also provides a computer-readable storage medium. A video surveillance program for visual operation is stored on the computer-readable storage medium of the present invention. When the video surveillance program for visual operation is executed by a processor, the steps of the above-mentioned video surveillance method for visual operation are implemented.
[0178] Among them, the method implemented when the video surveillance program for visual operation is executed can refer to each embodiment of the video surveillance method for visual operation of the present invention, and will not be elaborated here.
[0179] It should be noted that: the above sequence of embodiments of the present invention is only for description and does not represent the superiority or inferiority of the embodiments. The processes depicted in the drawings do not necessarily require the specific order or continuous order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0180] Each embodiment in this specification is described in a progressive manner. The same or similar parts among the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments.
[0181] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0182] The above are only the preferred embodiments of the present invention, and do not limit the protection scope of the present invention. Any equivalent structure / method transformation made using the content of the specification and drawings of the present invention under the inventive concept of the present invention, or any direct / indirect application in other related technical fields is included in the protection scope of the present invention.
Claims
1. A video surveillance method for visualization operations, characterized in that, The method includes: Obtaining target video information of a target object, and determining a current acquisition efficiency-time series of the facial information of the target object by using the target video information; Determining the behavioral state similarity corresponding to each adjacent time point in the current acquisition efficiency-time series, and determining each target cluster of the current acquisition efficiency-time series according to the behavioral state similarity; Determining the temporal difference between the target time period corresponding to any target cluster and the other time periods corresponding to other objects, and determining the possibility of continuous abnormal behavior of the target object in the target time period according to the temporal difference; Determining and marking the abnormal behavior of the target object according to the possibility of continuous abnormal behavior; Among them, the step of determining the behavioral state similarity corresponding to each adjacent time point in the current acquisition efficiency-time series includes: Determining the efficiency difference of the facial information acquisition efficiency between any two adjacent time points in the current acquisition efficiency-time series; Calculating the behavioral state similarity between any two adjacent time points by using the efficiency difference; Among them, the step of calculating the behavioral state similarity between any two adjacent time points by using the efficiency difference includes: Determining the instantaneous change rate of the facial information acquisition efficiency corresponding to any two adjacent time points respectively; Determining the instantaneous change rate difference between any two adjacent time points; Calculating the behavioral state similarity between any two adjacent time points by using the efficiency difference and the instantaneous change rate difference; Among them, the step of determining each target cluster of the current acquisition efficiency-time series according to the behavioral state similarity includes: Clustering the current acquisition efficiency-time series by using the Turbopixels algorithm and the behavioral state similarity to obtain each target cluster; wherein, the target cluster represents the behavioral state of the target object.
2. The video surveillance method for visualization operations according to claim 1, characterized in that, The step of determining the temporal difference between the target time period corresponding to any target cluster and the other time periods corresponding to other objects includes: Determining the other time periods corresponding to other objects that are closest to the target time period corresponding to any target cluster; Determining the start time interval and the end time interval between the target time period and the other time periods; Determining the temporal difference between the target time period and the other time periods by using the start time interval and / or the end time interval.
3. The video surveillance method for visualization operations according to claim 1, wherein, The step of determining the possibility of continuous abnormal behavior of the target object in the target time period according to the temporal difference includes: Determining the mean value of the target facial information acquisition efficiency of the target object in the target time period; Determining the mean value of the other facial information acquisition efficiency of other objects in the other time periods that are closest to the target time period; Calculating the possibility of continuous abnormal behavior of the target object in the target time period by using the temporal difference, the mean value of the target facial information acquisition efficiency, and the mean value of the other facial information acquisition efficiency.
4. The video surveillance method for visualization operations according to claim 1, characterized in that, The step of determining and marking the abnormal behavior of the target object according to the possibility of continuous abnormal behavior includes: Determining the historical acquisition efficiency-time series of the target object; Determining the time period difference between the current acquisition efficiency-time series and the historical acquisition efficiency-time series; Determining and marking the abnormal behavior of the target object according to the time period difference and the possibility of continuous abnormal behavior.
5. The video surveillance method for visualization operations according to claim 4, wherein The time period differences include: the total number of time period differences; The steps of determining and marking the abnormal behavior of the target object according to the time period difference and the possibility of continuous abnormal behavior include: Using the total number of time period differences, calculate the overall behavior abnormality of the target object; Determine and mark the abnormal behavior of the target object according to the overall behavior abnormality and the possibility of continuous abnormal behavior.
6. The video surveillance method for visualization operations according to claim 5, characterized in that The steps of determining and marking the abnormal behavior of the target object according to the overall behavior abnormality and the possibility of continuous abnormal behavior include: Calculate the abnormal behavior marking weight obtained by normalizing the product of the overall behavior abnormality and the possibility of continuous abnormal behavior; Compare the abnormal behavior marking weight with the preset abnormal marking threshold to determine and mark the abnormal behavior of the target object.
7. A video surveillance system for visualization operations, characterized in that, The system is used to implement the video monitoring method for visualization operations described in any one of claims 1 to 6; the system includes: A video acquisition module, configured to obtain the target video information of the target object, and use the target video information to determine the current acquisition efficiency-time series of the facial information of the target object; An information clustering module, configured to determine the behavior state similarity corresponding to each adjacent time point in the current acquisition efficiency-time series, and determine each target cluster of the current acquisition efficiency-time series according to the behavior state similarity; A behavior comparison module, configured to determine the time series difference between the target time period corresponding to any target cluster and the other time periods corresponding to other objects, and determine the possibility of continuous abnormal behavior of the target object in the target time period according to the time series difference; A behavior marking module, which determines and marks the abnormal behavior of the target object according to the possibility of continuous abnormal behavior; The information clustering module is further configured to determine the efficiency difference of the facial information acquisition efficiency between any two adjacent time points in the current acquisition efficiency-time series; use the efficiency difference to calculate the behavior state similarity between any two adjacent time points; Determine the instantaneous change rate of the facial information acquisition efficiency corresponding to any two adjacent time points; determine the instantaneous change rate difference between any two adjacent time points; use the efficiency difference and the instantaneous change rate difference to calculate the behavior state similarity between any two adjacent time points; Use the Turbopixels algorithm and the behavior state similarity to cluster the current acquisition efficiency-time series to obtain each target cluster; wherein, the target clustering represents the behavior state of the target object.
Citation Information
Patent Citations
Face tracking recognition technique based on video
CN102306290A
User behavior sequence anomaly detection method, terminal and storage medium
CN112491877A
Building intelligent comprehensive security monitoring method and system
CN118397506A