Unsupervised abnormal behavior detection method and device, equipment and storage medium
By performing frame extraction processing on video data, optical flow estimation and interpolation network generation prediction frames, and combining feature extraction network calculation loss function values, the false alarm and missed detection problems of existing abnormal behavior detection methods are solved, and more efficient abnormal behavior recognition is achieved.
Patent Information
- Application Number
- CN202510534824.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-25
- Publication Date
- 2025-08-08
AI Technical Summary
The existing abnormal behavior detection methods lack adaptability to spatial and temporal dynamic characteristics, and are prone to false alarms or missed detection, which is difficult to meet the safety protection needs of the elderly.
By collecting video data and performing frame extraction processing, the previous and next frames of the current frame are input to the optical flow estimation network to generate an optical flow estimation frame, and the interpolated network generates a predicted frame, and the network extracts the loss function value is calculated in combination with the feature extraction network to determine abnormal behavior.
It improves the accuracy and robustness of abnormal behavior detection, reduces the amount of data, avoids false alarms, and enhances the efficiency of identifying abnormal motion patterns.
Smart Images

Figure CN120452062A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of image processing technology and can be applied to fields such as elderly care. In particular, it relates to an unsupervised abnormal behavior detection method, device, equipment and storage medium. Background Art
[0002] Due to factors such as declining physical function and slowed reaction times, the elderly face multiple risks during outdoor activities, including falls, collisions, and injuries. Statistics show that the incidence of health problems such as fractures and brain injuries caused by travel accidents among the elderly is increasing year by year. The long recovery period and high medical costs further increase the burden on families and social insurance programs. Against this backdrop, how to use technology to improve travel safety for the elderly while also addressing their health status has become a key issue in the healthcare and insurance sectors.
[0003] In recent years, the rapid development of augmented reality (AR) technology has provided innovative solutions for the safe travel of the elderly. Safety systems based on AR devices can build active protection barriers for the elderly through real-time environmental perception, dynamic information superposition, and intelligent interaction. For example, AR glasses can identify environmental risks such as road obstacles and step heights in real time, and guide avoidance through visual cues; at the same time, combined with inertial sensors and physiological monitoring modules (such as heart rate and gait analysis), they can synchronously track the physical condition of the elderly, and provide early warnings when unstable gait or abnormal heart rate is detected, preventing falls or sudden illnesses. This technical model that deeply integrates safety protection and health monitoring is gradually becoming the core direction of smart elderly care.
[0004] In AR security systems, abnormal behavior detection is a key technology for protecting seniors from external threats. By integrating computer vision and artificial intelligence algorithms, the system can analyze surrounding behavior in real time to identify dangerous behaviors such as violence, harassment, illegal stalking, or snatching. For example, if it detects someone approaching quickly, following for a long time, or engaging in physical conflict, the system can immediately trigger an audible and visual alarm, send location information to emergency contacts, and even initiate intervention with the community security system.
[0005] However, the practical application of this technology faces multiple challenges. First, abnormal behavior is highly sudden and dynamic, making its timing, location, and manifestation difficult to predict. Second, factors such as changing lighting, crowd occlusion, and motion blur in complex environments can interfere with detection accuracy. Traditional rule-based or fixed pattern matching methods, lacking adaptability to spatiotemporal dynamics, are prone to false alarms or missed detections, making them unable to meet the high-reliability safety protection needs of the elderly. Summary of the Invention
[0006] The purpose of the present invention is to provide an unsupervised abnormal behavior detection method, device, equipment and storage medium, aiming to solve the problems that existing abnormal behavior detection methods lack adaptability to spatiotemporal dynamic characteristics and are prone to false alarms or missed detections.
[0007] In a first aspect, an embodiment of the present invention provides an unsupervised abnormal behavior detection method, comprising:
[0008] Collecting video data and performing frame extraction processing on the video data to obtain a previous frame, a current frame, and a next frame of the current frame;
[0009] Inputting a previous frame of the current frame and a subsequent frame of the current frame into an optical flow estimation network to generate a front optical flow estimation frame and a back optical flow estimation frame respectively;
[0010] Inputting a previous frame of the current frame, a subsequent frame of the current frame, the previous optical flow estimation frame, and the subsequent optical flow estimation frame into an interpolation network to generate a predicted frame;
[0011] Inputting the current frame and the predicted frame into a feature extraction network to obtain current features and predicted features respectively;
[0012] Calculating a loss function value according to the current frame, the predicted frame, the current feature, and the predicted feature;
[0013] Determine whether the loss function value reaches a preset threshold;
[0014] If the loss function value does not reach the preset threshold, it is determined to be abnormal behavior.
[0015] In a second aspect, an embodiment of the present invention provides an unsupervised abnormal behavior detection device, comprising:
[0016] An acquisition unit is used to acquire video data and perform frame extraction processing on the video data to obtain a previous frame, a current frame, and a next frame of the current frame;
[0017] an optical flow estimation unit, configured to input a frame preceding the current frame and a frame following the current frame into an optical flow estimation network to generate a preceding optical flow estimation frame and a succeeding optical flow estimation frame, respectively;
[0018] A prediction frame generation unit, configured to input a previous frame of the current frame, a subsequent frame of the current frame, the previous optical flow estimation frame, and the subsequent optical flow estimation frame into an interpolation network to generate a prediction frame;
[0019] A feature extraction unit, configured to input the current frame and the predicted frame into a feature extraction network to obtain current features and predicted features, respectively;
[0020] a calculation unit, configured to calculate a loss function value based on the current frame, the predicted frame, the current feature, and the predicted feature;
[0021] A judgment unit, configured to judge whether the loss function value reaches a preset threshold;
[0022] A determination unit is configured to determine that the behavior is abnormal if the loss function value does not reach a preset threshold.
[0023] In a third aspect, an embodiment of the present invention further provides a computer device, which includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the unsupervised abnormal behavior detection method described in the first aspect above is implemented.
[0024] In a fourth aspect, an embodiment of the present invention further provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the unsupervised abnormal behavior detection method described in the first aspect above is implemented.
[0025] The present invention discloses an unsupervised abnormal behavior detection method, device, equipment and storage medium, including: collecting video data, and performing frame extraction processing on the video data to obtain the previous frame of the current frame, the current frame and the next frame of the current frame; inputting the previous frame of the current frame and the next frame of the current frame into the optical flow estimation network and the interpolation network to generate a predicted frame; inputting the current frame and the predicted frame into the feature extraction network to obtain the current feature and the predicted feature respectively; calculating the loss function value according to the current frame, the predicted frame, the current feature and the predicted feature; judging whether the loss function value reaches a preset threshold; if the loss function value does not reach the preset threshold, it is determined to be abnormal behavior. The present invention utilizes the spatiotemporal continuity between adjacent frames through the interpolation network to synthesize a more accurate normal behavior trajectory, avoiding the false alarm problem that may occur in traditional methods. At the same time, compared with using a video sequence as input, the bidirectional frame interpolation method uses two frames of images as input, which not only reduces the amount of data, but also improves the efficiency of recognizing abnormal motion patterns. Furthermore, the present invention uses optical flow estimation technology to analyze motion information from adjacent frames, effectively identifying whether others exhibit unusual motion trajectories, such as rapid approach or unnatural movements, thereby improving the accuracy and robustness of abnormal behavior detection. The present invention also provides an unsupervised abnormal behavior detection device, a computer-readable storage medium, and a computer device, all of which exhibit the aforementioned beneficial effects and are not further detailed here. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0027] Figure 1 A schematic diagram of an application environment of an unsupervised abnormal behavior detection method according to an embodiment of the present invention;
[0028] Figure 2 Schematic diagram of the process of unsupervised abnormal behavior detection method;
[0029] Figure 3 Another flowchart of the unsupervised abnormal behavior detection method;
[0030] Figure 4 Schematic diagram of the sub-process of the unsupervised abnormal behavior detection method;
[0031] Figure 5 Schematic diagram of inference for unsupervised abnormal behavior detection method;
[0032] Figure 6 It is a structural diagram of the decoding module of the converter;
[0033] Figure 7 is a schematic block diagram of an unsupervised abnormal behavior detection device;
[0034] Figure 8 is a structural diagram of a computer device in one embodiment of the present invention;
[0035] Figure 9 FIG. 2 is another structural diagram of a computer device according to an embodiment of the present invention. DETAILED DESCRIPTION
[0036] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of them. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.
[0037] It will be understood that when used in this specification and the appended claims, the terms “comprises” and “comprising” indicate the presence of described features, integers, steps, operations, elements and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or collections thereof.
[0038] It should also be understood that the terminology used in this specification is for the purpose of describing particular embodiments only and is not intended to limit the present invention. As used in the specification and appended claims, the singular forms "a," "an," and "the" are intended to include the plural forms unless the context clearly indicates otherwise.
[0039] It should be further understood that the term "and / or" used in the present description and the appended claims refers to and includes any and all possible combinations of one or more of the associated listed items.
[0040] The unsupervised abnormal behavior detection method provided by the embodiment of the present invention can be applied in the following situations: Figure 1 In an application environment, the client communicates with the server through a network. The server can collect video data through the client, and perform frame extraction processing on the video data to obtain the previous frame of the current frame, the current frame, and the next frame of the current frame; then the previous frame of the current frame and the next frame of the current frame are input into the optical flow estimation network to generate a front optical flow estimation frame and a back optical flow estimation frame respectively; then the previous frame of the current frame, the next frame of the current frame, the previous optical flow estimation frame, and the back optical flow estimation frame are input into the interpolation network to generate a predicted frame; then the current frame and the predicted frame are input into the feature extraction network to obtain the current feature and the predicted feature respectively; then the loss function value is calculated according to the current frame, the predicted frame, the current feature, and the predicted feature; it is judged whether the loss function value reaches a preset threshold; if the loss function value does not reach the preset threshold, it is determined to be abnormal behavior. The abnormal behavior is fed back to the client. In the present invention, the temporal and spatial continuity between adjacent frames is used through the interpolation network to synthesize a more accurate normal behavior trajectory, thereby avoiding the false alarm problem that may occur in the traditional method. At the same time, compared to using a video sequence as input, the bidirectional frame interpolation method uses two frames of images as input, which not only reduces the amount of data, but also improves the efficiency of recognizing abnormal motion patterns. In addition, the present invention analyzes the motion information of adjacent frames through optical flow estimation technology, which can effectively identify whether others show abnormal motion trajectories, such as rapid approach or unnatural movements, thereby improving the accuracy and robustness of abnormal behavior detection. Among them, the client can be but is not limited to various personal computers, laptops, smart phones, tablets, AR devices and portable wearable devices. The server can be implemented with an independent server or a server cluster composed of multiple servers. The present invention is described in detail below through specific embodiments.
[0041] See also Figure 2 and Figure 3 This embodiment provides an unsupervised abnormal behavior detection method, including:
[0042] S101: Collect video data and perform frame extraction processing on the video data to obtain a previous frame, a current frame, and a next frame of the current frame;
[0043] In this embodiment, video data can come from a variety of channels. Common acquisition channels include: real-time collection of first-person video streams through the binocular camera of a head-mounted AR device to capture dynamic information of the elderly's surrounding environment, including pedestrian behavior, spatial position, and interactive actions. Alternatively, access to the monitoring system of public places such as communities and nursing homes to obtain global scene videos with a fixed perspective to supplement the detection of potential threats outside the field of view of the AR device (such as abnormal approach to blind spots). Alternatively, integrate multi-angle videos of smart glasses or portable recorders worn by the elderly to improve the accuracy of behavioral analysis through multi-perspective fusion. Alternatively, in applications in the medical and health field, video data comes from cameras deployed in wards and rehabilitation centers, and the cameras are used to collect videos of patient activities.
[0044] After collecting video data, it is necessary to extract the video data and pre-process each frame. For details, please refer to Figure 4 , collecting video data, and performing frame extraction processing on the video data to obtain a previous frame of a current frame, a current frame, and a next frame of the current frame include:
[0045] S201: Collect video data and extract a video sequence of a predetermined length through a sliding window to obtain a previous frame of a temporary current frame, a temporary current frame, and a subsequent frame of the temporary current frame;
[0046] After capturing video data, a sliding window mechanism is used to extract a continuous frame sequence. By setting the window length and step size (e.g., a window size of 3 frames and a step size of 1 frame), video segments are dynamically captured to generate the temporary previous frame, the temporary current frame, and the temporary next frame. This method covers the spatiotemporal characteristics of the video by sliding the window, avoiding the problem of missing action key frames caused by extracting frames at fixed intervals. For example, in a 30fps (30 frames per second) video, the sliding window moves in steps of 1 frame, ensuring the continuity of action between adjacent frames.
[0047] S202: performing wavelet transform decomposition on the previous frame of the temporary current frame, the temporary current frame, and the next frame of the temporary current frame to obtain high-frequency components and low-frequency components;
[0048] Specifically, a discrete wavelet transform (DWT) is performed on the frame before the temporary current frame, the temporary current frame, and the frame after the temporary current frame to decompose them into high-frequency components (detail information) and low-frequency components (approximate information). For example, a Daubechies wavelet basis (such as db1, db1 is the simplest Daubechies wavelet with a support length of 2 and only a first-order vanishing moment) is used to implement a two-layer decomposition:
[0049] High-frequency components are used to capture detailed features such as edges and textures, but are susceptible to noise interference;
[0050] The low-frequency component is used to preserve the main structure and brightness information of the image, but the contrast is low.
[0051] This embodiment separates noise and valid signals through multi-scale decomposition, which can lay the foundation for subsequent differentiation processing.
[0052] In the application of elderly care or medical health, taking the elderly care scene as an example, when someone is detected approaching the elderly, a discrete wavelet transform (DWT) is performed on the relevant video frames. The high-frequency component can capture the edge details of the person's body movements, such as subtle features such as quickly raising his hand (which may be a threatening action), but such details are easily disturbed by noise such as the movement of other objects in the environment (such as fluttering clothes, falling objects); the low-frequency component retains the overall structure of the scene (such as the location of the elderly, the layout of surrounding facilities) and brightness information, but there is a problem of insufficient contrast between the subject of the action and the background. Through the multi-scale decomposition of this embodiment, noise and effective signals can be accurately separated. If the edge of the action in the high-frequency component shows an abnormal sharp mutation (a characteristic of unnatural behavior), combined with the relative position of the person and the elderly in the low-frequency component (such as excessive approach), it can more timely and accurately identify abnormal behaviors such as pushing and snatching that threaten the safety of the elderly, providing a key basis for triggering an early warning for caregivers or monitoring systems.
[0053] S203: performing non-local mean denoising on the high-frequency component to obtain a denoised high-frequency component;
[0054] Specifically, we input the high-frequency component to be processed, select a pixel i to be denoised in the high-frequency component, and define a 7×7 neighborhood block N(i) centered on the pixel i. Its pixel value set is {v(i+k,i+l)|-3≤k,l≤3}. Then, we set a larger search window (e.g., 15×15) centered on i, traverse all remaining pixels j, and intercept the corresponding neighborhood block N(j).
[0055] Calculate the similarity between N(i) and N(j) using the following formula:
[0056]
[0057] Among them, d(i,j) represents the similarity between N(i) and N(j); G αis a Gaussian kernel function, and α controls the weight of the center pixel (default α = 1). k and l represent the coordinate offsets of the pixels in the neighborhood relative to the center pixel. k corresponds to the offset in the row direction (vertical direction), and l corresponds to the offset in the column direction (horizontal direction). For example, when k = 1 and l = 1, it indicates a position shifted one column to the right and one row down relative to the center pixel.
[0058] The similarity is then mapped to a weight value w(i,j) according to the following formula:
[0059]
[0060] Where h is the filtering parameter (the empirical value is usually 10σ, where σ is the noise standard deviation).
[0061] Then the normalization factor of the weight value w(i,j) is calculated according to the following formula:
[0062]
[0063] Then make sure the weights sum to 1 according to the following formula:
[0064]
[0065] Then perform weighted aggregation on all pixel points j in the search window according to the following formula:
[0066]
[0067] Among them, v(j) is the original value of pixel j, and u(i) is the high-frequency component after denoising.
[0068] By globally searching for similar image blocks and weightedly fusing them, rather than relying solely on local neighborhood information, similarity matching effectively identifies high-frequency information such as edges and textures when removing noise, avoiding the edge blurring caused by traditional methods such as Gaussian and median filtering. It is particularly effective at denoising repetitive textures (such as fabrics and skin pores) and is particularly suitable for detail enhancement in medical images (such as X-rays and dermatoscopes). By adjusting the filter parameter h, it can handle mixed scenarios with salt and pepper noise and Gaussian noise.
[0069] S204: performing contrast limited histogram equalization on the low-frequency component to obtain an equalized low-frequency component;
[0070] Specifically, the low-frequency component image is divided into S×S rectangular sub-blocks, with sizes of 8×8 or 16×16 being common. Smaller sub-blocks yield finer local contrast enhancement, but may also increase noise. Larger sub-blocks achieve a near-global equalization effect. If the image size cannot be divided evenly by the sub-block size, the edge regions are mirrored or replicated to ensure that all pixels are covered.
[0071] Then, the pixel value distribution of each sub-block is counted to obtain a histogram Hm,n(k), where k∈[0, L-1] (L is the number of gray levels).
[0072] Then calculate the maximum allowed histogram height Among them, clipLimit is a user-set parameter (usually 1.5 to 3.0).
[0073] The portion of the histogram that exceeds the threshold is then clipped and redistributed according to the following formula:
[0074]
[0075] Then calculate the normalized cumulative distribution function CDF of the clipped histogram m,n (k):
[0076]
[0077] Then, the original pixel value v in the sub-block is mapped to the equalized value v′ according to the normalized cumulative distribution function:
[0078] v ′ =round((L-1)×CDF m,n (v))
[0079] Among them, round means adjusting a value to the nearest integer.
[0080] Then, for any pixel (x, y) in the image, find the center coordinates of its four adjacent sub-blocks (m1, n1), (m1, n2), (m2, n1) and (m2, n2).
[0081] Then calculate the interpolation weight based on the coordinates of any pixel and its four adjacent sub-block centers:
[0082]
[0083] w 11 =(1-d x )(1-d y ),w 12 =(1-d x )d y
[0084] w 21 =d x (1-d y ),w 22 =d x d y
[0085] The weighted fusion mapping value is then calculated based on the interpolation weights of any pixel and its four adjacent sub-block center coordinates:
[0086]
[0087] Among them, v final represents the weighted fusion mapping value, that is, the low-frequency component after equalization in this embodiment; Represents the equalization value of the coordinate (m1, n1); Represents the equalization value of the coordinate (m1, n2); Represents the equalization value of the coordinate (m2, n1); Represents the equalization value of the coordinate (m2, n2).
[0088] Setting the maximum allowable height of each sub-block histogram (such as clipLimit = 2.0) can prevent excessive concentration of single grayscale pixels. At the same time, a small amount of high-frequency noise may remain in the low-frequency component, and cropping can prevent the grayscale corresponding to the noise from being over-enhanced. In AR device video processing, by setting the maximum allowable height of each sub-block histogram and cropping, it is possible to avoid significant amplification of dark area noise, affecting the accuracy of subsequent optical flow estimation. In addition, weighted fusion of pixel values at the sub-block boundaries can eliminate block artifacts.
[0089] In applications such as monitoring the activity areas of nursing homes, if a stranger attempts to approach an elderly person and potentially harass them, the video footage of the elderly person's surroundings (such as the floor and chairs) may contain multiple grayscale levels. Setting the maximum allowable height of each sub-block histogram (e.g., clipLimit = 2.0) prevents excessive concentration of pixels at a specific grayscale level on the ground, thus preventing overly bright or dark backgrounds from interfering with the observation of the person's movements. Furthermore, small amounts of high-frequency noise remaining within the low-frequency components (such as noise from swaying leaves in the distance) are clipped to prevent excessive enhancement of the corresponding grayscale levels, which could lead to misjudgment. Weighted fusion of pixel values at sub-block boundaries eliminates blocking artifacts and creates a more natural image. This processing improves the accuracy of optical flow estimation when analyzing human behavior, accurately capturing unusual movements such as a stranger suddenly accelerating towards an elderly person. Caregivers can detect and intervene promptly to ensure the safety of the elderly.
[0090] S205: Aggregate the denoised high-frequency components and the equalized low-frequency components to obtain a final previous frame of the current frame, a final current frame, and a final subsequent frame of the current frame.
[0091] The denoised high-frequency components and the equalized low-frequency components are subjected to an inverse discrete wavelet transform (IDWT) layer by layer to obtain the final frame before the current frame, the final current frame, and the final frame after the current frame. For example, for a two-layer decomposition result, the low-frequency and high-frequency components of the second layer are first reconstructed and then merged layer by layer to the original resolution, preserving edge details and global contrast.
[0092] Through two-layer decomposition and layer-by-layer reconstruction, contrast enhancement of low-frequency components and detail preservation of high-frequency components are synergistically optimized. This method is particularly suitable for scenarios requiring simultaneous improvement of global visibility and local precision (such as medical image analysis and security monitoring). It enhances information readability while avoiding noise amplification and edge distortion caused by traditional methods.
[0093] S102: Inputting a frame before the current frame and a frame after the current frame into an optical flow estimation network to generate a front optical flow estimation frame and a back optical flow estimation frame respectively;
[0094] Specifically, the previous frame of the current frame and the next frame of the current frame are input into the optical flow estimation network to generate the previous optical flow estimation frame and the next optical flow estimation frame respectively, including:
[0095] According to the following formula, the features of the previous frame and the next frame of the current frame are extracted to obtain the front optical flow estimation frame and the back optical flow estimation frame respectively;
[0096] F t-1→t+1 ,F t+1→t-1 =Φ(I t-1 ,I t+1 )
[0097] Among them, F t-1→t+1 represents the previous optical flow estimation frame; F t+1→t-1 Represents the post-optical flow estimation frame; I t-1 Indicates the previous frame of the current frame; I t+1 represents the next frame after the current frame; Φ represents the optical flow estimation network.
[0098] By analyzing the motion information of adjacent frames, optical flow estimation technology can effectively identify whether others exhibit abnormal motion trajectories, such as rapid approach or unnatural movements, thereby improving the accuracy and robustness of abnormal behavior detection.
[0099] In the application of the elderly care field, when a stranger is detected suddenly accelerating towards an elderly person living alone in the corridor of a nursing home, the optical flow estimation technology can accurately capture the abnormal trajectory by analyzing the motion information of adjacent frames. The previous frame (time t-1) and the next frame (time t+1) captured by the camera in front of the elderly are input into the optical flow estimation network to generate the previous optical flow estimation frame F t-1→t+1 and the post-optical flow estimation frame F t+1→t-1 If F t-1→t+1The displacement vector of the pixel block of the person in the image increases significantly (e.g., horizontal speed exceeds 1.5m / s), and the direction of movement continues to point toward the elderly person's location. Combined with the deformation characteristics of the human silhouette in the optical flow field (e.g., abnormal arm swing amplitude), it can be determined as a threatening and rapidly approaching behavior. This technology can accurately separate the target motion trajectory in complex backgrounds (such as people walking in the corridor and moving wheelchairs), avoiding the false alarms caused by background disturbances in traditional frame difference methods, and providing reliable abnormality warning basis for elderly care monitoring systems.
[0100] In healthcare, optical flow estimation networks can effectively identify sudden violent movements in patients in psychiatric wards, monitoring abnormal behavior. When a patient engages in unusual movements such as punching or kicking in the bedside area, their arm may be naturally drooping in the previous frame, but rapidly swinging to their chest in the next frame. The optical flow estimation frame will reveal an abnormal displacement vector at the elbow joint (the displacement exceeds three standard deviations of the normal motion mean), and the gradient rate of change of the limb edge in the optical flow field will increase significantly. By comparing the optical flow feature library with historical normal behavior (such as the optical flow distribution of eating and walking), the system can detect these unnatural movements in real time. By combining the relative position of the patient and surrounding personnel (such as an aggressive posture facing medical staff), it can assist medical staff in preemptive intervention to prevent injuries. This technology is robust to low-contrast scenes (such as nighttime hospital wards) and improves the accuracy of abnormal behavior detection by analyzing the direction and amplitude distribution of the optical flow vector, rather than simply relying on pixel intensity changes.
[0101] In some embodiments, after the front optical flow estimation frame and the rear optical flow estimation frame are generated, the front optical flow estimation frame and the rear optical flow estimation frame are input into the discriminant network, and the discriminant network performs discriminant analysis on the two optical flow estimation frames after a series of convolutional layers, pooling layers and fully connected layers. The output of the discriminant network is a probability value, which indicates the degree of matching between the front optical flow estimation frame and the rear optical flow estimation frame. If this probability value is greater than a preset threshold, it is considered that the two optical flow estimation frames are matched, that is, the optical flow estimation before and after the current frame is reasonable; on the contrary, if the probability value is less than the threshold, it means that the two optical flow estimation frames do not match, and it is necessary to re-extract features from the frames before and after the current frame or adjust the parameters of the optical flow estimation network until a matching optical flow estimation frame is obtained.
[0102] S103: Inputting a previous frame of the current frame, a subsequent frame of the current frame, the previous optical flow estimation frame, and the subsequent optical flow estimation frame into an interpolation network to generate a predicted frame;
[0103] Specifically, the previous frame of the current frame, the next frame of the current frame, the previous optical flow estimation frame, and the next optical flow estimation frame are input into the interpolation network to generate a predicted frame, including:
[0104] The predicted frame is calculated based on the previous frame of the current frame, the next frame of the current frame, the previous optical flow estimation frame and the next optical flow estimation frame according to the following formula:
[0105]
[0106] in, F represents the predicted frame; t-1→t+1 represents the previous optical flow estimation frame; F t+1→t-1 Represents the post-optical flow estimation frame; I t-1 Indicates the previous frame of the current frame; I t+1 represents the next frame after the current frame; ω represents the parameters of the interpolation network Ψ.
[0107] This embodiment uses bidirectional frame interpolation to exploit the spatiotemporal continuity between adjacent frames to synthesize more accurate normal motion trajectories, avoiding the false positives that can occur with traditional methods. Compared to using a video sequence as input, bidirectional frame interpolation uses two frames as input, which not only reduces the amount of data but also improves the efficiency of identifying abnormal motion patterns.
[0108] In the application of the elderly care field, when a stranger is detected to be abnormally close to the elderly in the activity area of the nursing home, the interpolation network can accurately reconstruct the predicted frame by fusing the previous and next frames and the optical flow information. For example, the previous frame (t-1) of the current frame shows that the person is 1.5 meters in front of the left of the elderly, and the next frame (t+1) shows that he is close to 0.5 meters. The optical flow estimation frame F t-1→t+1 and F t+1→t-1 The displacement vector of the figure's outline shows continuous high-speed motion (speed > 2m / s) toward the elderly. This information is input into the interpolation network Ψ, and the interpolation weight is dynamically adjusted based on the parameter ω to generate a predicted frame. If the person's body posture in the actual current frame shows abnormal contact movements (such as the arm extending forward close to the elderly's shoulder), and the predicted frame only shows a natural walking posture due to normal behavior trajectory learning, the pixel-level difference between the two (such as the grayscale value change in the shoulder area > 30%) and the feature space cosine similarity (<0.4) will trigger an abnormal alarm. This can effectively identify dangerous behaviors such as pushing and pulling, and can also avoid the detection delay caused by motion blur in the traditional frame difference method.
[0109] In the healthcare field, interpolation networks can detect unnatural sudden changes in limb movement in rehabilitation training areas. For example, during gait training, a patient's left leg is in the stance phase in the previous frame (t-1), but suddenly exhibits abnormal knee flexion (a non-trained movement) in the next frame (t+1). The optical flow vector direction at the knee joint in the optical flow estimation frame deviates by more than 2σ from the mean of the normal gait. Based on the joint position information and optical flow field motion trends of the previous and next frames, the interpolation network generates a predicted frame containing only gentle knee flexion movements that meet the training specifications. When the actual knee angle change rate (>150° / s) in the current frame differs significantly from the predicted frame, combined with the abnormal joint spatial position (e.g., center of gravity offset) output by the feature extraction network, the system can detect unbalanced movements or abnormal force application before a fall in real time, providing medical staff with precise intervention opportunities and improving the efficiency of safety monitoring in rehabilitation scenarios. This technology effectively suppresses the interference of complex background (such as reflections from training equipment and ground shadows) on prediction accuracy through the interpolation mechanism guided by optical flow, and enhances the ability to recognize abnormal actions in low-contrast environments.
[0110] It's important to note that both the optical flow estimation network and the interpolation network in this example are trained from scratch on normal video sequences. The optical flow estimation network learns to estimate regular optical flow corresponding to normal motion, but may produce poor optical flow estimation results for unseen abnormal motion patterns. During the inference phase, a large interpolation error will indicate an abnormal frame.
[0111] S104: Inputting the current frame and the predicted frame into a feature extraction network to obtain current features and predicted features respectively;
[0112] Specifically, the current frame and the predicted frame are input into the feature extraction network to obtain the current features and predicted features respectively, including:
[0113] The features of the current frame and the predicted frame are extracted through multiple convolutional layers and pooling layers to obtain the spatial features of the current frame and the spatial features of the predicted frame respectively;
[0114] A 3D convolutional network is used to extract features from the current frame and the predicted frame to obtain the temporal features of the current frame and the predicted frame respectively;
[0115] The spatial features and the corresponding temporal features are fused to obtain the current features and the predicted features respectively.
[0116] Spatial features provide information about the static structure and appearance of objects within a frame, while temporal features reflect the dynamic changes of objects over time. Combining these two features allows features to contain richer information, overcoming the shortcomings of individual features. For example, when recognizing an action, spatial features can identify the body parts involved, while temporal features can determine the order and pattern of movement of these parts, leading to more accurate action recognition.
[0117] In the elderly care sector, feature extraction networks play a key role in monitoring common areas of nursing homes. For example, consider an elderly person walking in a hallway. At a given moment, the current frame depicts the individual's normal walking posture, while the predicted frame is generated based on the model's understanding of normal behavior. Through multiple convolutional and pooling layers, spatial features such as the individual's body contour and clothing texture are extracted from the current frame, and the predicted frame also obtains corresponding spatial features. For example, the convolutional layer captures pattern details on the individual's clothing, while the pooling layer filters and reduces these features to highlight key features.
[0118] Specifically, the features of the current frame and the predicted frame are extracted through multiple convolutional layers and pooling layers based on the improved ResNet-50 to obtain the spatial features of the current frame and the spatial features of the predicted frame respectively;
[0119] Among them, the spatial features of the current frame Mathematical representation of:
[0120]
[0121] Spatial features of predicted frames Mathematical representation of:
[0122]
[0123] Then, 3D-ResNet18 (3D convolutional network) is used to extract features of the current frame and the predicted frame to obtain the temporal features of the current frame and the predicted frame respectively;
[0124] Among them, the temporal characteristics of the current frame Mathematical representation of:
[0125]
[0126] Mathematical representation of the temporal features of the predicted frame:
[0127]
[0128] Then, the gated attention mechanism (GatedAttention) is used to adaptively weight the spatial features and the corresponding temporal features;
[0129] Specifically, the temporal features of the current frame and the spatial features of the current frame are spliced together according to the following formula, and the temporal features of the predicted frame and the spatial features of the predicted frame are spliced together to generate the fusion features of the current frame respectively: and the fusion features of the predicted frame
[0130]
[0131] Then, according to the fusion features of the current frame and the fusion features of the predicted frame, the attention weight α of the current frame and the attention weight of the predicted frame are generated respectively according to the following formula:
[0132]
[0133] Among them, σ is the Sigmoid function, W a and b a are learnable parameters.
[0134] Then generate the current feature according to the attention weight of the current frame, the spatial feature of the current frame and the temporal feature of the current frame At the same time, the prediction features are generated according to the attention weight of the prediction frame, the spatial features of the prediction frame and the temporal features of the prediction frame
[0135]
[0136] The gated attention mechanism enables the system to dynamically adjust the weights of spatial and temporal features based on the scene, allowing for more accurate processing of information in different scenarios. For example, in video analysis tasks, when the action in the scene changes rapidly, the system can increase the weight of temporal features to better capture the coherence of the action; whereas, when the layout and positional relationships of objects in the scene are complex, the weight of spatial features is increased accordingly, helping to accurately identify objects and their relationships. This ability to dynamically adjust weights greatly enhances the system's adaptability and flexibility.
[0137] In the field of elderly care, take the monitoring of public activity areas in nursing homes as an example. During daily activities, the elderly engage in various activities in this area, such as chatting and playing chess. Based on the convolutional and pooling layers of the improved ResNet-50, spatial features are extracted from the current frame and the predicted frame. The captured information such as the elderly person's body posture, the position of surrounding objects (such as tables and chairs), and other details such as an elderly person leaning forward slightly, approaching the chessboard, with both hands on the edge of the chessboard, and the distribution of chess pieces on the chessboard are all recorded in the spatial features. The spatial features of the predicted frame predict the elderly person's possible movements and the state of objects based on normal activity patterns, such as the elderly person's hand may reach out to a certain chess piece.
[0138] 3D-ResNet18 then processes these frames to extract temporal features. The temporal features of the current frame capture the dynamic information of the elderly person's movements, such as the speed and trajectory of their hand as they reach for the chess piece. The temporal features of the predicted frame estimate the temporal trend of subsequent movements, such as the likelihood of the hand continuing to move and pick up the piece, and the speed changes.
[0139] The gated attention mechanism then takes effect, concatenating the temporal and spatial features of the current frame to generate a fused feature for the current frame. The predicted frame is generated similarly. In this scenario, if an elderly person suddenly loses balance, the fused feature of the current frame will include both the elderly person's unusual posture (spatial features) and the rapid changes in movement at the moment of loss of balance (temporal features). Attention weights are then calculated, and the attention weights of the current frame and the predicted frame are adjusted based on learnable parameters and the fused features. Because the movement in this scenario is so rapid, the gated attention mechanism increases the weight of the temporal feature. The resulting current and predicted features clearly reflect this abnormality. By comparing the current and predicted features, the system can quickly detect the elderly person's unusual movement and promptly notify staff for assistance, preventing them from falling and injuring themselves.
[0140] S105: Calculating a loss function value according to the current frame, the predicted frame, the current feature, and the predicted feature;
[0141] Specifically, the loss function value is calculated based on the current frame, the predicted frame, the current feature, and the predicted feature, including:
[0142] The loss function value is calculated based on the current frame, predicted frame, current features and predicted features according to the following formula:
[0143] l=l1+λl2
[0144]
[0145] Where l represents the total loss function value; l1 and l2 represent the loss function values; λ represents the balance hyperparameter, which is used for the balance weight and is generally set to 0.0001; I t Indicates the current frame; represents the predicted frame; represents the prediction feature; f t Indicates the current feature.
[0146] S106: Determine whether the loss function value reaches a preset threshold;
[0147] In this embodiment, determining whether the loss function value reaches a preset threshold includes:
[0148] Obtain different environmental information;
[0149] Then, the environment is divided into multiple categories according to different environmental information to obtain multiple environmental categories;
[0150] Then, for each environment category, we collect loss function value samples of normal behavior under the environment category and calculate the preset thresholds under different environment categories;
[0151] When detecting, obtain the current environment information;
[0152] Then, the environment category is determined according to the current environment information to obtain the current environment category;
[0153] Then adopt the corresponding preset threshold according to the current environment category.
[0154] Specifically, GPS / Beidou positioning is used to obtain geographic location (e.g., longitude and latitude coordinates), which is then combined with timestamps to divide time periods (e.g., morning rush hour / nighttime). Cameras / sensors are used to collect light intensity, crowd density (using YOLOv8 for real-time crowd counting), and noise decibel levels. Map APIs (e.g., Amap / Google Places) are then used to identify scene labels (e.g., "hospital entrance," "bus stop," or "park bench").
[0155] Then, several sets of environmental feature vectors (including geographic location, time period, light intensity, crowd density, noise decibel value and scene labels, etc.) are collected, and the environment is divided into five typical scenes based on several sets of environmental feature vectors: dense traffic areas (such as intersections and subway stations), leisure places (such as parks and squares), commercial areas (such as supermarkets and shopping malls), medical places (such as hospital corridors and pharmacies) and residential communities (such as community roads and elevators).
[0156] Then, for each environment category, FocalLoss is used to calculate the predicted loss value of the normal behavior samples of the environment category, and the 95% quantile of the loss value of each environment category is taken as the abnormal trigger threshold (i.e., the preset threshold).
[0157] In densely populated areas, a higher abnormal trigger threshold needs to be set due to the following characteristics:
[0158] Behavior in densely populated areas is more complex. For example, normal behavior in densely populated areas includes frequent physical contact (such as avoiding and squeezing through the crowd), rapid movement (speed > 1.2m / s) and group synchronous movement (such as queuing and moving forward in waves), which leads to the overall high loss value predicted by the model.
[0159] Secondly, dense occlusion leads to an increase in the pose estimation error, which indirectly increases the calculated value of the loss function.
[0160] In addition, maintaining a social distance of >1m in sparse areas (such as community gardens) is normal, but a distance of <0.5m may still be reasonable in dense areas.
[0161] When detecting, obtain the current environment information;
[0162] Then, the environment category is determined according to the current environment information to obtain the current environment category;
[0163] Then, the corresponding preset threshold is adopted according to the current environment category.
[0164] This embodiment effectively solves the adaptability problem of cross-scenario behavior recognition through environmental perception and dynamic threshold setting, and has higher sensitivity to weak abnormal signals (such as slow falls and slight pushes) during the travel of the elderly.
[0165] After obtaining the corresponding preset threshold, the behavior of others is judged by determining whether the loss function value reaches the preset threshold. Specifically, if the total loss function value l does not reach the preset threshold, it is judged as abnormal behavior. If the total loss function value l reaches the preset threshold, it is judged as normal behavior.
[0166] For example, in a senior care scenario, while an elderly person strolls in the garden of a nursing home, the system continuously analyzes surveillance video. Suppose the loss function value calculated after a series of processing for the current video frame is 0.2, while the preset threshold is 0.3. Since 0.2 is less than 0.3, meaning the loss function value does not reach the preset threshold, the system determines that abnormal behavior has occurred. Further analysis reveals that a stranger is rapidly approaching the elderly person, and their movement trajectory and posture differ significantly from normal behavior patterns, consistent with the abnormality indicated by the lower loss function value. The system then immediately triggers an alarm, notifying staff to investigate, effectively ensuring the elderly person's safety.
[0167] Another example is monitoring patient behavior in a hospital ward in the healthcare sector. When the patient is resting normally, the system calculates a loss function value of 0.4, with a preset threshold of 0.35. At this point, the loss function value is greater than the preset threshold, and the system determines this behavior as normal. This indicates that the patient's movements, vital signs, and other comprehensive characteristics are consistent with normal resting behavior, and no additional intervention is required from medical staff. However, if the patient suddenly stands up and moves violently, causing the loss function value to drop to 0.25, below the preset threshold, the system will determine this behavior as abnormal, alerting medical staff to the patient's condition, as the patient may be feeling unwell or have other emergencies that require attention.
[0168] In some embodiments, the total loss function value l is used to optimize the optical flow estimation network and the interpolation network. The total loss function value l is used to optimize the optical flow estimation network and the interpolation network, and it comprehensively considers the errors of these two networks during the training process. The optical flow estimation network aims to accurately predict the optical flow information in the image, while the interpolation network is responsible for interpolating the optical flow to improve the accuracy and continuity of the optical flow. By minimizing the total loss function value l, the optical flow estimation network can be prompted to continuously adjust its internal parameters, thereby improving the accuracy of the optical flow estimation. At the same time, the interpolation network will also optimize the parameters of the interpolation algorithm based on the feedback of the total loss function value l to better interpolate the optical flow. In each iterative training, the optical flow estimation network generates a preliminary optical flow prediction result based on the input image data, and the interpolation network then interpolates this result. The total loss function value l is then calculated, and the error is backpropagated based on this value to update the parameters of the two networks. As training progresses, the total loss function value l will gradually decrease, indicating that the performance of the optical flow estimation network and the interpolation network is constantly improving, until the pre-set convergence condition is reached. At this time, both networks can complete their respective tasks well and work together to provide high-quality optical flow information.
[0169] Also, see Figure 5 , which determines abnormal behavior by determining whether the loss function value l2 reaches a preset threshold. If the loss function value l2 does not reach the preset threshold, it is determined to be abnormal behavior. If the loss function value l2 reaches the preset threshold, it is determined to be normal behavior.
[0170] The range of l2 is [-1, 1]. The closer the value of l2 is to 1, the more similar the two vectors are; the closer the value of l2 is to -1, the less similar the two vectors are; and the value of l2 equal to 0 indicates that there is no similarity between the two vectors.
[0171] S107: If the loss function value does not reach a preset threshold, it is determined to be abnormal behavior.
[0172] S108: If the loss function value reaches a preset threshold, it is determined to be normal behavior.
[0173] In some embodiments, during the training process and the inference process, the optical flow estimation network and the interpolation network are both part of the decoder-transformer module of the transformer, and the structure is as follows: Figure 6 shown.
[0174] The specific parameters are: a 16-layer network is used, the number of hidden layer units is hidden size 2048, the number of attention heads is 16, and the number of trainable tokens is token trained is 2T.
[0175] In some embodiments, the last two layers or one layer of decoding layer decoder block can be arranged to run on the AR device.
[0176] Because these final two layers or one decoder block run on the AR device, they can process data directly locally. This not only reduces reliance on external devices for data transmission but also improves the speed and efficiency of data processing. When these decoding layers run on an AR device, they can more quickly convert received information into recognizable content, such as images, sounds, or specific commands. This helps enhance the user's sense of real-time interactivity in the AR experience, making the integration of virtual elements and real scenes more smooth and natural.
[0177] This embodiment integrates an optical flow estimation network and an interpolation network through the above method, and learns the feature representation of normal video sequences through joint training on normal video sequences.
[0178] First, during the training phase, the optical flow estimation network and the interpolation network work together to capture the characteristics of normal behavior in the video. This approach differs from traditional deep learning-based anomaly detection techniques in that it introduces bidirectional interpolation and optical flow estimation mechanisms, enhancing the distinction between normal and abnormal behavior.
[0179] Specifically, during the inference phase, the system uses interpolation error to identify anomalies. This process involves interpolating the current frame with its preceding and following frames to generate intermediate frames. The interpolation error between these intermediate frames is then calculated as a key metric for anomaly detection. Simultaneously, the optical flow estimation network analyzes the motion trajectories between adjacent frames, providing additional features representing abnormal behavior. This two-pronged approach effectively improves the accuracy of anomaly detection.
[0180] Furthermore, the system utilizes the latest decoder-transformer module for bidirectional interpolation and optical flow estimation, supporting an end-to-end inference process. To optimize resource utilization and improve response speed, some inference tasks can be performed locally on the AR device, reducing the server burden and enabling faster device response. This not only improves detection efficiency but also ensures real-time performance and accuracy, making it ideal for applications in elderly care settings where high reliability is crucial.
[0181] See also Figure 7 This embodiment provides an unsupervised abnormal behavior detection device 300, comprising:
[0182] The acquisition unit 301 is used to acquire video data and perform frame extraction processing on the video data to obtain a previous frame, a current frame, and a next frame of the current frame;
[0183] An optical flow estimation unit 302 is configured to input a frame preceding the current frame and a frame following the current frame into an optical flow estimation network to generate a preceding optical flow estimation frame and a succeeding optical flow estimation frame, respectively;
[0184] The prediction frame generation unit 303 is configured to input a previous frame of the current frame, a subsequent frame of the current frame, the previous optical flow estimation frame, and the subsequent optical flow estimation frame into an interpolation network to generate a prediction frame;
[0185] A feature extraction unit 304 is configured to input the current frame and the predicted frame into a feature extraction network to obtain current features and predicted features, respectively;
[0186] A calculation unit 305 is configured to calculate a loss function value based on the current frame, the predicted frame, the current feature, and the predicted feature;
[0187] A judging unit 306 is configured to judge whether the loss function value reaches a preset threshold;
[0188] The determination unit 307 is configured to determine that the behavior is abnormal if the loss function value does not reach a preset threshold.
[0189] Furthermore, the acquisition unit 301 includes:
[0190] An extraction subunit is used to collect video data and extract a video sequence of a predetermined length through a sliding window to obtain a previous frame of a temporary current frame, a temporary current frame, and a next frame of the temporary current frame;
[0191] a decomposition subunit, configured to perform wavelet transform decomposition on a frame preceding the temporary current frame, the temporary current frame, and a frame following the temporary current frame to obtain high-frequency components and low-frequency components;
[0192] a denoising subunit, configured to perform non-local mean denoising on the high-frequency component to obtain a denoised high-frequency component;
[0193] an equalization subunit, configured to perform contrast-limited histogram equalization on the low-frequency component to obtain an equalized low-frequency component;
[0194] The aggregation subunit is used to aggregate the denoised high-frequency components and the equalized low-frequency components to obtain a final previous frame of the current frame, a final current frame, and a final subsequent frame of the current frame.
[0195] Furthermore, the optical flow estimation unit 302 includes:
[0196] An extraction subunit is used to perform feature extraction on a frame before the current frame and a frame after the current frame according to the following formula to obtain a front optical flow estimation frame and a back optical flow estimation frame respectively;
[0197] F t-1→t+1 ,F t+1→t-1 =Φ(I t-1 ,I t+1 )
[0198] Among them, F t-1→t+1 represents the previous optical flow estimation frame; F t+1→t-1 Represents the post-optical flow estimation frame; I t-1 Indicates the previous frame of the current frame; I t+1 represents the next frame after the current frame; Φ represents the optical flow estimation network.
[0199] Furthermore, the predicted frame generation unit 303 includes:
[0200] The prediction frame calculation subunit is configured to calculate a prediction frame based on a previous frame of the current frame, a next frame of the current frame, the previous optical flow estimation frame, and the next optical flow estimation frame according to the following formula:
[0201]
[0202] in, F represents the predicted frame; t-1→t+1 represents the previous optical flow estimation frame; F t+1→t-1 Represents the post-optical flow estimation frame; I t-1 Indicates the previous frame of the current frame; I t+1 represents the next frame after the current frame; ω represents the parameters of the interpolation network Ψ.
[0203] Furthermore, the feature extraction unit 304 includes:
[0204] A spatial feature extraction subunit, configured to extract features from the current frame and the predicted frame through a plurality of convolutional layers and pooling layers, to obtain spatial features of the current frame and spatial features of the predicted frame, respectively;
[0205] A temporal feature extraction subunit, configured to extract features from the current frame and the predicted frame using a 3D convolutional network to obtain temporal features of the current frame and temporal features of the predicted frame, respectively;
[0206] The fusion subunit is used to fuse the spatial features and the corresponding temporal features to obtain current features and predicted features respectively.
[0207] Furthermore, the calculation unit 305 includes:
[0208] The loss function value calculation subunit is configured to calculate the loss function value according to the current frame, the predicted frame, the current feature, and the predicted feature according to the following formula:
[0209] l=l1+λl2
[0210]
[0211] Among them, l represents the total loss function value; l1 and l2 represent the loss function value; λ represents the balance hyperparameter; I t Indicates the current frame; represents the predicted frame; represents the prediction feature; f t Indicates the current feature.
[0212] Furthermore, the judging unit 306 includes:
[0213] Different information acquisition subunits are used to obtain different environmental information;
[0214] A division subunit is used to divide the environment into multiple categories according to different environmental information to obtain multiple environmental categories;
[0215] A collecting subunit is used to collect loss function value samples of normal behaviors under each environment category and calculate preset thresholds under different environment categories;
[0216] The current information acquisition subunit is used to obtain the current environment information when detecting;
[0217] The determination subunit is used to determine the environment category according to the current environment information and obtain the current environment category;
[0218] The adopting subunit is used to adopt the corresponding preset threshold value according to the current environment category.
[0219] The present invention provides an unsupervised abnormal behavior detection device, which first collects video data and performs frame extraction processing on the video data to obtain the previous frame, the current frame and the next frame of the current frame; then the previous frame of the current frame and the next frame of the current frame are input into an optical flow estimation network and an interpolation network to generate a predicted frame; then the current frame and the predicted frame are input into a feature extraction network to obtain current features and predicted features, respectively; then a loss function value is calculated based on the current frame, the predicted frame, the current features and the predicted features; then it is determined whether the loss function value reaches a preset threshold; if the loss function value does not reach the preset threshold, it is determined to be abnormal behavior, which can effectively identify whether others show abnormal movement trajectories, thereby improving the accuracy and robustness of abnormal behavior detection.
[0220] For the specific definition of the unsupervised abnormal behavior detection device, please refer to the definition of the unsupervised abnormal behavior detection method above, which will not be repeated here. Each unit in the above-mentioned unsupervised abnormal behavior detection device can be implemented in whole or in part by software, hardware, or a combination thereof. The above-mentioned units can be embedded in or independent of the processor in the computer device in the form of hardware, or can be stored in the memory of the computer device in the form of software, so that the processor can call and execute the operations corresponding to the above-mentioned units.
[0221] In one embodiment, a computer device is provided. The computer device may be a server, and its internal structure diagram may be as follows: Figure 8 As shown. The computer device includes a processor, a memory, a network interface and a database connected via a system bus. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile and / or volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external client via a network connection. When the computer program is executed by the processor, it realizes the functions or steps on the server side of an unsupervised abnormal behavior detection method.
[0222] In one embodiment, a computer device is provided. The computer device may be a client, and its internal structure diagram may be as follows: Figure 9 As shown. The computer device includes a processor, memory, network interface, display screen and input device connected via a system bus. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external server via a network connection. When the computer program is executed by the processor, it realizes the functions or steps of the client side of an unsupervised abnormal behavior detection method.
[0223] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the following steps are performed:
[0224] Collecting video data and performing frame extraction processing on the video data to obtain a previous frame, a current frame, and a next frame of the current frame;
[0225] Inputting a previous frame of the current frame and a subsequent frame of the current frame into an optical flow estimation network to generate a front optical flow estimation frame and a back optical flow estimation frame respectively;
[0226] Inputting a previous frame of the current frame, a subsequent frame of the current frame, the previous optical flow estimation frame, and the subsequent optical flow estimation frame into an interpolation network to generate a predicted frame;
[0227] Inputting the current frame and the predicted frame into a feature extraction network to obtain current features and predicted features respectively;
[0228] Calculating a loss function value according to the current frame, the predicted frame, the current feature, and the predicted feature;
[0229] Determine whether the loss function value reaches a preset threshold;
[0230] If the loss function value does not reach the preset threshold, it is determined to be abnormal behavior.
[0231] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the following steps are implemented:
[0232] Collecting video data and performing frame extraction processing on the video data to obtain a previous frame, a current frame, and a next frame of the current frame;
[0233] Inputting a previous frame of the current frame and a subsequent frame of the current frame into an optical flow estimation network to generate a front optical flow estimation frame and a back optical flow estimation frame respectively;
[0234] Inputting a previous frame of the current frame, a subsequent frame of the current frame, the previous optical flow estimation frame, and the subsequent optical flow estimation frame into an interpolation network to generate a predicted frame;
[0235] Inputting the current frame and the predicted frame into a feature extraction network to obtain current features and predicted features respectively;
[0236] Calculating a loss function value according to the current frame, the predicted frame, the current feature, and the predicted feature;
[0237] Determine whether the loss function value reaches a preset threshold;
[0238] If the loss function value does not reach the preset threshold, it is determined to be abnormal behavior.
[0239] It should be noted that the above functions or steps that can be implemented by the computer-readable storage medium or computer device can be found in the relevant descriptions of the server side and the client side in the aforementioned method embodiment. To avoid repetition, they will not be described one by one here.
[0240] Those skilled in the art will appreciate that all or part of the processes in the above-mentioned embodiments can be implemented by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM).
[0241] Those skilled in the art will clearly understand that for the sake of convenience and brevity of description, only the division of the above-mentioned functional units and modules is used as an example. In actual applications, the above-mentioned functions can be distributed and completed by different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.
[0242] The embodiments described above are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention, and should all be included in the scope of protection of the present invention.
Claims
1. An unsupervised abnormal behavior detection method, characterized in that: include: Collecting video data and performing frame extraction processing on the video data to obtain a previous frame, a current frame, and a next frame of the current frame; Inputting a previous frame of the current frame and a subsequent frame of the current frame into an optical flow estimation network to generate a front optical flow estimation frame and a back optical flow estimation frame respectively; Inputting a previous frame of the current frame, a subsequent frame of the current frame, the previous optical flow estimation frame, and the subsequent optical flow estimation frame into an interpolation network to generate a predicted frame; Inputting the current frame and the predicted frame into a feature extraction network to obtain current features and predicted features respectively; Calculating a loss function value according to the current frame, the predicted frame, the current feature, and the predicted feature; Determine whether the loss function value reaches a preset threshold; If the loss function value does not reach the preset threshold, it is determined to be abnormal behavior.
2. The unsupervised abnormal behavior detection method according to claim 1, characterized in that The collecting video data and performing frame extraction processing on the video data to obtain a frame before a current frame, a current frame, and a frame after the current frame include: Collecting video data, and extracting a video sequence of a predetermined length through a sliding window to obtain a previous frame of a temporary current frame, a temporary current frame, and a next frame of the temporary current frame; Perform wavelet transform decomposition on the previous frame of the temporary current frame, the temporary current frame, and the next frame of the temporary current frame to obtain high-frequency components and low-frequency components; Performing non-local mean denoising on the high-frequency component to obtain a denoised high-frequency component; performing contrast limited histogram equalization on the low-frequency component to obtain an equalized low-frequency component; The denoised high-frequency components and the equalized low-frequency components are aggregated to obtain the previous frame of the final current frame, the final current frame, and the next frame of the final current frame.
3. The unsupervised abnormal behavior detection method according to claim 1, characterized in that Inputting a previous frame of the current frame and a subsequent frame of the current frame into an optical flow estimation network to generate a previous optical flow estimation frame and a subsequent optical flow estimation frame respectively includes: Perform feature extraction on the previous frame of the current frame and the next frame of the current frame according to the following formula to obtain a front optical flow estimation frame and a back optical flow estimation frame respectively; F t-1→t+1 ,F t+1→t-1 =Φ(I t-1 ,I t+1 ) Among them, F t-1→t+1 represents the previous optical flow estimation frame; F t+1→t-1 Represents the post-optical flow estimation frame; I t-1 Indicates the previous frame of the current frame; I t+1 represents the next frame after the current frame; Φ represents the optical flow estimation network.
4. The unsupervised abnormal behavior detection method according to claim 1, characterized in that The step of inputting a previous frame of the current frame, a subsequent frame of the current frame, the previous optical flow estimation frame, and the subsequent optical flow estimation frame into an interpolation network to generate a predicted frame comprises: The predicted frame is obtained by calculating based on the previous frame of the current frame, the next frame of the current frame, the previous optical flow estimation frame and the next optical flow estimation frame according to the following formula: in, represents the predicted frame; F t-1→t+1 represents the previous optical flow estimation frame; F t+1→t-1 Represents the post-optical flow estimation frame; I t-1 Indicates the previous frame of the current frame; I t+1 represents the next frame after the current frame; ω represents the parameters of the interpolation network Ψ.
5. The unsupervised abnormal behavior detection method according to claim 1, characterized in that: Inputting the current frame and the predicted frame into a feature extraction network to obtain current features and predicted features respectively includes: Performing feature extraction on the current frame and the predicted frame through multiple convolutional layers and pooling layers to obtain spatial features of the current frame and spatial features of the predicted frame respectively; Using a 3D convolutional network to extract features from the current frame and the predicted frame, respectively obtaining a temporal feature of the current frame and a temporal feature of the predicted frame; The spatial features and the corresponding temporal features are fused to obtain current features and predicted features respectively.
6. The unsupervised abnormal behavior detection method according to claim 1, characterized in that: Calculating the loss function value according to the current frame, the predicted frame, the current feature, and the predicted feature includes: The loss function value is calculated according to the current frame, the predicted frame, the current feature and the predicted feature according to the following formula: l=l1+λl2 Among them, l represents the total loss function value; l1 and l2 represent the loss function value; λ represents the balance hyperparameter; I t Indicates the current frame; represents the predicted frame; represents the prediction feature; f t Indicates the current feature.
7. The unsupervised abnormal behavior detection method according to claim 1, characterized in that: The step of determining whether the loss function value reaches a preset threshold comprises: Obtain different environmental information; According to different environmental information, the environment is divided into multiple categories to obtain multiple environmental categories; For each environment category, collect loss function value samples of normal behavior under the environment category, and calculate preset thresholds under different environment categories; When detecting, obtain the current environment information; Determine the environment category according to the current environment information and obtain the current environment category; Use the corresponding preset threshold according to the current environment category.
8. An unsupervised abnormal behavior detection device, characterized in that: include: An acquisition unit is used to acquire video data and perform frame extraction processing on the video data to obtain a previous frame, a current frame, and a next frame of the current frame; an optical flow estimation unit, configured to input a frame preceding the current frame and a frame following the current frame into an optical flow estimation network to generate a preceding optical flow estimation frame and a succeeding optical flow estimation frame, respectively; A prediction frame generation unit, configured to input a previous frame of the current frame, a subsequent frame of the current frame, the previous optical flow estimation frame, and the subsequent optical flow estimation frame into an interpolation network to generate a prediction frame; A feature extraction unit, configured to input the current frame and the predicted frame into a feature extraction network to obtain current features and predicted features, respectively; a calculation unit, configured to calculate a loss function value based on the current frame, the predicted frame, the current feature, and the predicted feature; A judgment unit, configured to judge whether the loss function value reaches a preset threshold; A determination unit is configured to determine that the behavior is abnormal if the loss function value does not reach a preset threshold.
9. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the unsupervised abnormal behavior detection method according to any one of claims 1 to 7 is implemented.
10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, which, when executed by a processor, causes the processor to perform the unsupervised abnormal behavior detection method according to any one of claims 1 to 7.
Citation Information
Cited By
Model training method, task processing method based on optical flow estimation and related products
CN121147668A
High-risk operator state intelligent analysis system based on AI large model
CN121438182A