Video monitoring anomaly identification method and system applied to public safety

CN122551053APending Publication Date: 2026-08-11浙江飞至科技有限公司
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-03
Publication Date
2026-08-11

AI Technical Summary

Technical Problem

背景建模方法通过构建背景模型,将当前帧与背景模型进行差分来检测运动目标,该类方法计算效率较高,但对光照变化、阴影干扰等环境因素较为敏感,且难以区分正常运动与异常运动

Benefits of technology

[0006]Based on the above, embodiments of the present invention segment the original monitoring video stream acquired by image acquisition devices deployed in public safety monitoring areas into a set of video segment units with temporal sequence continuity, each carrying an acquisition timestamp and an acquisition location identifier. Visual primitive decomposition is then performed on each video segment unit to obtain static scene primitive components representing the distribution of structured visual elements in the background scene and dynamic target primitive components representing the instantaneous motion of visual elements in the foreground moving target. The static scene primitive components and dynamic target primitive components are then input into a pre-constructed spatiotemporal primitive coupling coding network, and co-coding is performed across primitive components to generate a spatiotemporal coupling feature representation with a spatiotemporal coupling relationship between the two. This spatiotemporal coupling feature representation is then input into a pre-constructed abnormal situation inference network, and multi-directional situation inference is performed through situation differentiation branches to generate a set of multi-directional situation differentiation trajectories describing the possible evolution of visual primitive components in the next video segment unit. Finally, each situation differentiation trajectory in the multi-directional situation differentiation trajectory set is compared with the normal situation trajectory template library constructed from normal monitoring video streams that do not contain abnormal events to measure the situation deviation, generate situation deviation distribution features, and identify abnormal situation types accordingly. This invention achieves a unified spatiotemporal representation of scene structure and moving targets through the decomposition and co-coding of static scene primitive components and dynamic target primitive components, avoiding the contextual information loss problem caused by the separation of background and foreground in traditional methods. Through multi-directional situational deduction based on spatiotemporal coupling feature representation, anomaly identification is no longer limited to static discrimination at the current moment, but possesses the ability to predict the possible situational trajectory at the next moment. This allows for early warning through early signs of situational deviation before anomalies fully manifest. By comparing trajectory morphologies with a normal situational trajectory template library, fine-grained anomaly type identification based on situational evolution patterns is achieved. Compared to traditional threshold judgment or reconstruction error judgment methods, this can more accurately characterize the category attributes of anomalies, significantly improving the accuracy, timeliness, and type distinguishability of anomaly identification in public safety video monitoring.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122551053A_ABST
    Figure CN122551053A_ABST
Patent Text Reader

Abstract

This invention provides a video surveillance anomaly identification method and system for public safety. The method involves acquiring the original surveillance video stream and segmenting it into a set of video segment units carrying timestamps and location identifiers. Each video segment unit undergoes visual primitive decomposition to obtain static scene primitive components and dynamic target primitive components. These are then input into a spatiotemporal primitive coupling coding network for cross-primitive co-coding, generating spatiotemporal coupling feature representations with spatiotemporal coupling relationships. These spatiotemporal coupling feature representations are then input into an anomaly situation inference network for multi-directional situation inference, generating a set of multi-directional situation differentiation trajectories. Finally, the situation deviation of each situation differentiation trajectory is measured against normal trajectories in a normal situation trajectory template library to identify the type of anomaly situation. This invention achieves a unified representation of scene and motion through primitive decomposition and coupling coding, and combined with forward-looking situation inference and trajectory deviation measurement, improves the accuracy and type distinguishability of anomaly identification.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of artificial intelligence, and more specifically, to a method and system for anomaly identification in video monitoring applied to public safety. Background Technology

[0002] Public security video surveillance is a crucial technological means for maintaining social order and preventing and combating illegal and criminal activities. With the large-scale deployment of urban security systems, the real-time analysis of massive amounts of video data has become a key requirement for ensuring public safety. Anomaly detection in video surveillance aims to automatically detect events that deviate from normal scenarios, such as crowds gathering, abnormal running, left-behind objects, and trespassing into restricted areas, from continuous video streams. This is of great significance for timely early warning and rapid response.

[0003] Existing anomaly detection methods in video surveillance mainly include motion detection methods based on background modeling and anomaly detection methods based on deep learning. Background modeling methods detect moving targets by constructing a background model and subtracting the current frame from the background model. These methods are computationally efficient but are sensitive to environmental factors such as lighting changes and shadow interference, and struggle to distinguish between normal and abnormal motion. Deep learning-based methods typically use structures such as autoencoders or generative adversarial networks to learn from normal samples, determining anomalies through reconstruction errors or discrimination scores. These methods improve feature representation capabilities, but mostly analyze single frames or short video clips, lacking modeling of the relationship between static background elements and dynamic foreground targets in the scene, and failing to fully utilize temporal information for forward-looking anomaly situation inference. Furthermore, existing methods have limited ability to distinguish fine-grained anomaly types, typically only providing a binary judgment of whether an anomaly exists, making it difficult to identify specific anomaly situation categories. Summary of the Invention

[0004] In view of the aforementioned problems, and in conjunction with the first aspect of the present invention, embodiments of the present invention provide a video surveillance anomaly identification method applied to public safety, the method comprising: The original monitoring video stream is acquired by the image acquisition equipment deployed in the public safety monitoring area. The original monitoring video stream is divided into a set of video segment units with temporal continuity. Each video segment unit in the set of video segment units carries an acquisition timestamp identifier and an acquisition location identifier. Visual primitive decomposition is performed on each video segment unit in the video segment unit set to obtain the visual primitive component set of the corresponding video segment unit. The visual primitive component set includes static scene primitive components and dynamic target primitive components. The static scene primitive components represent the structured visual element distribution of the background scene in the video segment unit, and the dynamic target primitive components represent the instantaneous motion visual elements of the foreground moving target in the video segment unit. The static scene primitive components and dynamic target primitive components are input into a pre-constructed spatiotemporal primitive coupling coding network. Through the primitive component interaction mapping layer in the spatiotemporal primitive coupling coding network, the static scene primitive components and dynamic target primitive components are co-coded across primitive components to generate a spatiotemporal coupling feature representation with spatiotemporal coupling relationship between static scene primitive components and dynamic target primitive components. The spatiotemporal coupling feature representation is input into a pre-constructed abnormal situation inference network. The spatiotemporal coupling feature representation is used to perform multi-directional situation inference through the situation differentiation branch in the abnormal situation inference network, generating a set of multi-directional situation differentiation trajectories corresponding to video segment units. The set of multi-directional situation differentiation trajectories describes the various situation trends that the visual primitive components in the video segment unit may evolve in the next video segment unit. For each situation differentiation trajectory in the multi-directional situation differentiation trajectory set, a situation deviation measurement is performed on the normal situation trajectory in the preset normal situation trajectory template library, a situation deviation distribution feature is generated, and the abnormal situation type contained in the video segment unit is identified based on the situation deviation distribution feature. The preset normal situation trajectory template library is constructed from normal monitoring video streams that do not contain abnormal events.

[0005] In another aspect, embodiments of the present invention also provide a video monitoring anomaly identification system for public safety, including a processor and a machine-readable storage medium connected to the processor. The machine-readable storage medium is used to store programs, instructions, or code, and the processor is used to execute the programs, instructions, or code in the machine-readable storage medium to implement the above-described method.

[0006] Based on the above, embodiments of the present invention segment the original monitoring video stream acquired by image acquisition devices deployed in public safety monitoring areas into a set of video segment units with temporal sequence continuity, each carrying an acquisition timestamp and an acquisition location identifier. Visual primitive decomposition is then performed on each video segment unit to obtain static scene primitive components representing the distribution of structured visual elements in the background scene and dynamic target primitive components representing the instantaneous motion of visual elements in the foreground moving target. The static scene primitive components and dynamic target primitive components are then input into a pre-constructed spatiotemporal primitive coupling coding network, and co-coding is performed across primitive components to generate a spatiotemporal coupling feature representation with a spatiotemporal coupling relationship between the two. This spatiotemporal coupling feature representation is then input into a pre-constructed abnormal situation inference network, and multi-directional situation inference is performed through situation differentiation branches to generate a set of multi-directional situation differentiation trajectories describing the possible evolution of visual primitive components in the next video segment unit. Finally, each situation differentiation trajectory in the multi-directional situation differentiation trajectory set is compared with the normal situation trajectory template library constructed from normal monitoring video streams that do not contain abnormal events to measure the situation deviation, generate situation deviation distribution features, and identify abnormal situation types accordingly. This invention achieves a unified spatiotemporal representation of scene structure and moving targets through the decomposition and co-coding of static scene primitive components and dynamic target primitive components, avoiding the contextual information loss problem caused by the separation of background and foreground in traditional methods. Through multi-directional situational deduction based on spatiotemporal coupling feature representation, anomaly identification is no longer limited to static discrimination at the current moment, but possesses the ability to predict the possible situational trajectory at the next moment. This allows for early warning through early signs of situational deviation before anomalies fully manifest. By comparing trajectory morphologies with a normal situational trajectory template library, fine-grained anomaly type identification based on situational evolution patterns is achieved. Compared to traditional threshold judgment or reconstruction error judgment methods, this can more accurately characterize the category attributes of anomalies, significantly improving the accuracy, timeliness, and type distinguishability of anomaly identification in public safety video monitoring. Attached Figure Description

[0007] Figure 1 This is a schematic diagram of the execution flow of the video monitoring anomaly identification method for public safety provided in an embodiment of the present invention.

[0008] Figure 2 This is a schematic diagram of exemplary hardware and software components of a video monitoring anomaly identification system for public safety provided in an embodiment of the present invention. Detailed Implementation

[0009] The present invention will now be described in detail with reference to the accompanying drawings. Figure 1This is a flowchart illustrating a video monitoring anomaly identification method for public safety provided in one embodiment of the present invention. The following is a detailed description of this video monitoring anomaly identification method for public safety.

[0010] Step S110: Obtain the original monitoring video stream collected by the image acquisition device deployed in the public safety monitoring area, and divide the original monitoring video stream into a set of video segment units with time sequence continuity. Each video segment unit in the set of video segment units carries an acquisition timestamp identifier and an acquisition location identifier.

[0011] The starting point for the execution of the method in this embodiment is acquiring the raw monitoring video stream generated by the image acquisition device. The image acquisition device is a fixed-view network high-definition camera deployed in the southeast corner of the urban transportation hub square. This camera continuously captures the rectangular monitoring area it covers at a fixed frame rate. The acquired raw monitoring video stream is transmitted in real time to the video analysis server via a dedicated video transmission cable or a secure network protocol.

[0012] The data receiving and preprocessing module inside the video analytics server first receives the raw monitoring video stream and establishes a circular buffer in memory to temporarily store continuously arriving video frame data. The module continuously reads video frames from the circular buffer and performs video segment unit extraction based on a preset fixed time window length. In this embodiment, the time window length is calculated based on the frame rate to cover a time period encompassing several consecutive video frames. The extraction operation is performed progressively with a fixed sliding step size, which is smaller than the time window length. This ensures that there is a temporal overlap between adjacent video segment units. This overlap design effectively prevents events occurring at the boundary of a video segment unit from being interrupted due to crossing the boundary, ensuring the complete capture of the continuity of event evolution. In each extraction operation, the data receiving and preprocessing module retrieves all consecutive video frames within the corresponding time window from the circular buffer and packages them into an independent video segment unit.

[0013] For each newly generated video segment unit, the data receiving and preprocessing module synchronously generates two sets of structured identification information. The first set is the acquisition timestamp identifier, which consists of a start timestamp field and an end timestamp field. The start timestamp field accurately records the acquisition time of the first video frame of the video segment unit, and the end timestamp field accurately records the acquisition time of the last video frame of the video segment unit, with timestamp accuracy down to the millisecond level. The second set is the acquisition location identifier, which is a globally unique string. Its encoding rule is a combination of the camera's area type, the camera's physical deployment location description, and the camera's hardware number. This identifier is "Fixed camera No. 03, Southeast corner of the transportation hub square". This acquisition location identifier not only uniquely points to a specific physical camera, but is also pre-bound in the server's internal configuration file to the camera's three-dimensional spatial coordinates, installation tilt angle, horizontal rotation angle, and lens focal length, among other internal and external parameters.

[0014] The data receiving and preprocessing module combines each newly generated video segment unit, its acquisition timestamp identifier, and its acquisition location identifier into a structured data object, and appends this data object to a global ordered list. When the original monitoring video stream finishes playing or reaches the preset analysis duration, this ordered list becomes the generated set of video segment units. This set maintains the temporal continuity of the original monitoring video stream, and each element in the set carries explicit spatiotemporal tracing information.

[0015] Step S120: Perform visual primitive decomposition on each video segment unit in the video segment unit set to obtain the visual primitive component set of the corresponding video segment unit. The visual primitive component set includes static scene primitive components and dynamic target primitive components. The static scene primitive components represent the structured visual element distribution of the background scene in the video segment unit, and the dynamic target primitive components represent the instantaneous motion visual elements of the foreground moving target in the video segment unit.

[0016] After generating the video segment unit set in step S110, the video analysis server begins to perform deep visual analysis on each video segment unit in the set sequentially. The first core step of visual analysis is visual primitive decomposition. The purpose of this step is to decompose the complex visual information contained in each video segment unit into two mutually exclusive and complementary primitive components: static scene primitive components and dynamic target primitive components. The static scene primitive components are used to fully describe the spatially structured distribution of all background scene visual elements in the camera field whose visual attributes remain stable and whose spatial positions do not change over time during the duration of the video segment unit. These visual elements include, but are not limited to, the paving texture of the square ground, the vegetation outline of the green belt, the facade structure of buildings, and fixed signs. The dynamic target primitive components are used to fully describe the motion state of all instantaneous visual change elements in the camera field caused by the movement of foreground objects during the duration of the video segment unit. These moving visual elements include, but are not limited to, the instantaneous position, instantaneous direction of movement, and instantaneous amplitude of a single pedestrian walking, and the instantaneous position, instantaneous direction of movement, and instantaneous amplitude of a moving vehicle.

[0017] This decomposition process is automatically completed by a pre-built and trained deep neural network model, namely the Visual Primitive Separation Network. The Visual Primitive Separation Network is designed with three processing paths: a static primitive decomposition branch, a dynamic primitive decomposition branch, and a primitive co-constraint layer. The static and dynamic primitive decomposition branches process the input video segment units in parallel, while the primitive co-constraint layer receives the initial outputs from the two branches and performs cross-validation and correction.

[0018] Step S121: Perform frame sequence parsing on each video segment unit in the video segment unit set to obtain the video frame time sequence that constitutes the video segment unit.

[0019] Before sending any video segment unit into the visual primitive separation network, the video analysis server first performs frame sequence parsing on the video segment unit. The video segment unit is stored in memory in a compressed encoding format, encoded using a common video coding standard. The frame sequence parsing process calls the video decoder corresponding to the encoding format to restore the compressed video data block into a series of uncompressed two-dimensional digital image frames arranged in chronological order. Each digital image frame consists of three independent two-dimensional matrices, corresponding to the red, green, and blue color channels, respectively. The dimensions of the three two-dimensional matrices are all based on the preset resolution of the camera's output video, with a certain number of pixel sampling points in both the width and height directions. All the decoded two-dimensional digital image frames are organized into a linear sequence according to their original acquisition time; this linear sequence is the video frame time series. This video frame time series completely preserves all the spatiotemporal visual information of the original video segment unit.

[0020] Step S122: Call the static primitive decomposition branch in the pre-built visual primitive separation network to perform scene structure analysis on each video frame in the video frame time series, extract the distribution pattern of visual elements in each video frame that do not change position over time, and use them as the initial static scene primitive components of the video segment unit. Scene structure analysis includes dividing each video frame into spatial regions and determining the repetition pattern of visual elements in each divided region.

[0021] After obtaining the video frame time series, the static primitive decomposition branch in the visual primitive separation network begins to process it. The static primitive decomposition branch is a feature extraction network with a deep convolutional neural network as its backbone. It contains a series of computational modules specifically designed for static scene analysis to extract those structured visual patterns that persist in the time dimension.

[0022] Step S1221: Perform hierarchical spatial partitioning on each video frame in the video frame time series. Divide each video frame into a set of spatial grid units with different scales according to a multi-level spatial partitioning strategy. The multi-level spatial partitioning strategy includes spatial region segmentation with different granularities.

[0023] The static primitive decomposition branch first performs a hierarchical spatial partitioning on each video frame in the video frame time series. This hierarchical spatial partitioning is accomplished by a spatial pyramid mesh generator. The spatial pyramid mesh generator receives the pixel matrix of a single video frame as input and aggregates spatial regions of different receptive field ranges from the pixel matrix through multiple built-in parallel pooling operation branches.

[0024] The spatial pyramid mesh generator comprises a first spatial partitioning path, a second spatial partitioning path, and a third spatial partitioning path. The first spatial partitioning path uses the largest possible pooling window, sliding it across the pixel matrix of the video frame with a predetermined step size. It maps the pixel region covered by each pooling window to a first-level spatial mesh cell, which corresponds to a relatively large rectangular region on the input video frame. The output of the first spatial partitioning path is an array composed of several first-level spatial mesh cells.

[0025] The second spatial partitioning branch uses a centrally sized pooling window that slides across the pixel matrix of the video frame with a predetermined step size. The pixel region covered by each pooling window is mapped to a second-level spatial grid cell, which corresponds to a rectangular region of medium area on the input video frame. The output of the second spatial partitioning branch is an array consisting of several second-level spatial grid cells.

[0026] The third spatial branch uses the smallest pooling window, and the pixel area covered by each pooling window is mapped to a third-level spatial grid cell, which corresponds to a relatively small rectangular area on the input video frame. The output of the third spatial branch is an array of several third-level spatial grid cells.

[0027] The outputs of the three spatial partitioning branches collectively constitute the multi-scale spatial grid cell set for that video frame. The first-level spatial grid cells capture macroscopic scene structures, such as a complete green belt area or the full width of a road. The second-level spatial grid cells capture mesoscopic structural information, such as the boundaries between different plants within a green belt or the shape of a single stone slab on a sidewalk. The third-level spatial grid cells capture microscopic details and textures, such as the rough texture of a stone slab surface or the veining of plant leaves. For each video frame in the video frame time series, the spatial pyramid grid generator independently executes the above three parallel spatial partitioning operations, generating the corresponding three-level spatial grid cell set for that frame.

[0028] Step S1222: Perform visual attribute statistics on the visual elements within each spatial grid cell, and extract the texture direction distribution, color distribution, and edge direction distribution within each spatial grid cell as local visual attribute descriptors.

[0029] After generating a multi-scale spatial grid cell set for each video frame, the visual attribute descriptor extraction module within the static primitive decomposition branch begins feature extraction for each individual grid cell at each level. This module contains three parallel convolutional filter banks, used to extract three visual attributes: texture orientation distribution, color distribution, and edge direction distribution.

[0030] The first group is the texture orientation distribution extraction filter bank. This filter bank consists of multiple convolutional kernels with different orientation sensitivity characteristics, including a horizontal texture extraction filter with the maximum response to horizontal texture, a vertical texture extraction filter with the maximum response to vertical texture, a first oblique texture extraction filter with the maximum response to a first diagonal texture, and a second oblique texture extraction filter with the maximum response to a second diagonal texture. Each of these filters performs convolution operations on the pixel matrix corresponding to the target spatial grid cell, generating its own corresponding texture response map. The value of each element in the texture response map represents the intensity of the texture pattern at that pixel location in the corresponding direction. The visual attribute descriptor extraction module performs spatial average pooling on all pixel response values ​​within each texture response map to obtain a scalar value, which represents the overall response intensity of that grid cell on the texture in that specific orientation. The response intensity values ​​of all directional filters are combined in a predetermined order to obtain a multi-dimensional texture orientation distribution vector. The dimension of this texture orientation distribution vector is equal to the number of directional filters in the filter bank, and the value of each dimension in the vector characterizes the distribution density of a specific orientation texture within that grid cell.

[0031] The second group is the color distribution extraction filter group. This filter group first transforms the pixel matrix of the target space grid cell from the original color space to a hue-saturation-luminance color space. This color space decomposes color information into hue, saturation, and luminance components, each of which is an independent two-dimensional matrix. Then, the filter group statistically analyzes the numerical distribution of the hue component matrix, dividing the hue range into several preset hue intervals and counting the number of pixels falling within each interval, forming a hue distribution histogram. Simultaneously, the filter group performs similar interval division and pixel count analysis on the saturation and luminance component matrices, respectively, forming saturation and luminance distribution histograms. The visual attribute descriptor extraction module concatenates the statistical values ​​of each interval from the hue, saturation, and luminance distribution histograms in a predetermined order, forming a multi-dimensional color distribution vector. The dimension of this color distribution vector is equal to the sum of the number of intervals in the three histograms, and the value of each dimension in the vector represents the frequency of occurrence of a pixel within that grid cell in a specific color attribute interval.

[0032] The third group is the edge direction distribution extraction filter group. This filter group first uses a set of edge detection convolution kernels to convolve the pixel matrix of the target spatial grid cell. This set of kernels includes a horizontal edge detection kernel for extracting horizontal edges and a vertical edge detection kernel for extracting vertical edges. Using the response values ​​of these two kernels, the edge intensity value and edge direction angle value at each pixel location within the grid cell can be calculated. This filter group maps the edge direction angle values ​​of all pixel locations to several preset edge direction angle intervals. Then, it accumulates the edge intensity values ​​corresponding to the pixel locations within each angle interval to form an edge direction intensity distribution histogram. The visual attribute descriptor extraction module arranges the accumulated intensity values ​​of each interval of this edge direction intensity distribution histogram in a predetermined order to form a multi-dimensional edge direction distribution vector. The dimension of this edge direction distribution vector is equal to the total number of preset edge direction angle intervals, and the value of each dimension in the vector represents the cumulative intensity of the edge contour along a specific direction within the grid cell.

[0033] Finally, the visual attribute descriptor extraction module sequentially concatenates the texture direction distribution vector, color distribution vector, and edge direction distribution vector along the vector dimension to generate a unified, higher-dimensional feature vector. This feature vector is the local visual attribute descriptor of the grid cell. This local visual attribute descriptor is a multi-dimensional feature vector, where each dimension describes a specific visual attribute feature within the grid cell, rather than a single, isolated value.

[0034] Step S1223: Perform temporal consistency analysis on the local visual attribute descriptors of spatial grid cells at the same spatial location in all video frames in the video frame time series, and select stable spatial grid cells whose temporal volatility of local visual attribute descriptors is lower than a preset volatility threshold within the time span covered by the video frame time series, as stable units of scene structure.

[0035] The visual attribute descriptor extraction module generates corresponding local visual attribute descriptor vectors for all spatial grid cells at all levels in all video frames of the video frame time series. The temporal consistency analysis module of the static primitive decomposition branch then performs a time-dimensional stability assessment on these descriptor vectors.

[0036] The temporal consistency analysis module first collects descriptor sequences for each fixed spatial grid location. For a spatial grid cell at a specific spatial coordinate location under a specific hierarchical division, the module traverses each video frame of the video frame time series, extracts the local visual attribute descriptors of the grid cells at the same spatial coordinate location in that video frame, and arranges them in chronological order to form a multidimensional vector sequence that evolves over time, called the local descriptor time series.

[0037] Then, the temporal consistency analysis module performs temporal volatility quantification calculations on the local descriptor time series. This module first calculates the statistical center vector of the series. The statistical center vector is calculated by averaging the local visual attribute descriptor vectors at all time points in the series along each vector dimension, resulting in an average descriptor vector with the same dimension as the local visual attribute descriptors. Next, the module iterates through each time point in the series, calculating the vector distance between the local visual attribute descriptor vector at that time point and the average descriptor vector. This vector distance is calculated as the sum of the absolute values ​​or the square root of the sum of the squares of the differences between the two vectors of the same dimension along each dimension. The module then compiles a distance sequence based on the vector distance values ​​corresponding to all time points, and calculates the statistical mean of this distance sequence. The resulting scalar value is the temporal volatility metric for that spatial grid cell.

[0038] This module compares the calculated temporal volatility metric with a preset volatility threshold. The preset volatility threshold is a hyperparameter of the static primitive decomposition branch, determined during model training through validation set debugging. If the temporal volatility metric of a spatial grid cell is less than the preset volatility threshold, it indicates that the visual attributes within that spatial grid cell have maintained high stability throughout the entire video segment's time span, and its internal visual content has remained largely unchanged. Such a grid cell is labeled as a scene-structure-stable cell. For example, in a transportation hub plaza scene, spatial grid cells corresponding to the fixed exterior wall area of ​​the plaza buildings, because their texture and color remain unchanged throughout the video segment, have very similar local visual attribute descriptors across frames, resulting in extremely low calculated temporal volatility metrics, and are therefore labeled as scene-structure-stable cells. Conversely, spatial grid cells corresponding to areas frequently traversed by pedestrians, whose visual attributes change continuously with pedestrian occlusion and departure, exhibit drastic fluctuations in their local visual attribute descriptors across frames, resulting in high calculated temporal volatility metrics, and are thus classified as unstable cells.

[0039] Step S1224: Spatially stitch the scene structure stabilization unit and the local visual attribute descriptor corresponding to the scene structure stabilization unit according to the original spatial position of the scene structure stabilization unit in the video frame to construct a static scene primitive component that represents the distribution of structured visual elements of the background scene in the video segment unit. The constructed static scene primitive component is used as the initial static scene primitive component of the video segment unit.

[0040] After the temporal consistency analysis module completes the stability labeling of all spatial grid cells, the scene structure stitching module of the static primitive decomposition branch begins to construct the initial static scene primitive components. The scene structure stitching module creates a two-dimensional descriptor graph grid corresponding to the spatial size of the input video frame. Each spatial location of this grid reserves a feature vector storage bit with a length equal to the dimension of the local visual attribute descriptor.

[0041] For spatial grid cells marked as stable units of scene structure, the scene structure stitching module directly fills the corresponding storage location of the two-dimensional descriptor graph grid with the local visual attribute descriptor vector of the cell, according to the spatial coordinates of the cell in the original video frame. For spatial locations not marked as stable units of scene structure, i.e., those regions with drastic temporal fluctuations, the scene structure stitching module uses a spatial neighborhood interpolation completion strategy to generate background descriptors for them. Specifically, for each spatial location to be completed, the module extracts the local visual attribute descriptor vectors of several spatial locations in its surrounding spatial neighborhood that have been marked as stable units of scene structure. These extracted vectors are then weighted and averaged according to their vector dimensions, with the weights set to be inversely proportional to the spatial distance to the location to be completed; that is, the closer the stable unit is, the greater its contribution weight to the descriptor vector. The weighted average result vector is used as the background descriptor for the location to be completed and filled into the two-dimensional descriptor graph grid.

[0042] After all spatial locations have been filled, the two-dimensional descriptor graph mesh output by the scene structure stitching module is the initial static scene primitive component of this video segment unit. This initial static scene primitive component is a three-dimensional tensor in terms of data structure. Its height and width spatial dimensions correspond to the spatial dimensions of the original video frame, while its channel dimension is equal to the dimension of the local visual attribute descriptor vector. It completely records the spatial distribution of background structured visual elements in the entire monitored scene within the time span of this video segment unit.

[0043] Step S123: Call the dynamic primitive decomposition branch in the pre-built visual primitive separation network to extract the instantaneous representation of motion visual elements in the visual difference region between adjacent video frames in the video frame time series, and generate the initial dynamic target primitive components of the video segment unit.

[0044] While the static primitive decomposition branch operates in parallel, the dynamic primitive decomposition branch in the visual primitive separation network processes the same video frame time series. The dynamic primitive decomposition branch is a network branch specifically designed to capture inter-frame motion information, and its core task is to extract the instantaneous state description of the foreground moving target from the continuous frame differences.

[0045] Step S1231: Perform pixel-level difference on two temporally consecutive video frames in the video frame time series to generate a difference response map. The pixel-level difference captures the change in pixel brightness at each pixel position between the two video frames.

[0046] The frame difference calculation module of the dynamic primitive decomposition branch traverses every pair of temporally adjacent video frames in the video frame time series. For a pair of adjacent frames consisting of the preceding and following video frames, the frame difference calculation module performs pixel-level difference operations. This module first converts both the preceding and following video frames from their original color images into single-channel luminance images. Each pixel value in the luminance image is obtained by weighted combination of the red, green, and blue components of that pixel. Then, for all pixel pairs with the same spatial coordinates in the two luminance images, the module calculates the absolute difference in pixel luminance values. This is achieved by subtracting the luminance value at the same coordinate in the preceding luminance image from the luminance value at a certain coordinate in the following luminance image, and taking the absolute value of the difference. The absolute differences calculated for all spatial coordinates are arranged according to their original spatial coordinates to form a two-dimensional matrix with the same spatial size as the original video frames. This two-dimensional matrix is ​​the difference response map of the adjacent frame pair. The larger the value of an element in the difference response map, the more drastic the luminance change at that pixel position between the two frames, and the higher the probability of motion.

[0047] Step S1232: Enhance the response region of the differential response map. Determine the set of pixel positions in the differential response map where the pixel brightness change exceeds the preset response threshold as the target motion response region. The target motion response region represents the concentrated area of ​​pixel brightness change caused by the foreground moving target between two video frames.

[0048] After generating the differential response map, the response region enhancement module of the dynamic primitive decomposition branch performs post-processing on the map to eliminate weak responses caused by non-motion factors such as sensor thermal noise and random illumination flicker.

[0049] The response region enhancement module first performs a threshold filtering operation. This module compares the response value of each pixel in the differential response map with a preset response threshold. This preset response threshold is a hyperparameter of the dynamic primitive decomposition branch. For pixel locations with response values ​​greater than the preset response threshold, the module marks them as candidate moving pixels; for pixel locations with response values ​​less than or equal to the preset response threshold, the module marks them as background pixels and directly sets their response values ​​to zero.

[0050] After thresholding, the response region enhancement module performs morphological filtering on the resulting binary mask of candidate moving pixels. The morphological filtering first uses a structuring element to perform erosion on the binary mask, removing isolated noise points with small areas. Then, it uses the same or slightly larger structuring element to perform dilation, connecting neighboring candidate moving pixels and filling small holes inside the moving target caused by texture uniformity. After erosion and dilation, several connected white regions are formed in the binary mask; each connected region is a target motion response region. The outer contours of these target motion response regions surround the areas of concentrated pixel brightness changes caused by the foreground moving target.

[0051] Step S1233: Estimate the instantaneous motion direction of the target motion response area. Based on the position offset of the corresponding pixel positions in the two video frames between frames, determine the instantaneous motion direction of each pixel in the target motion response area. Perform direction distribution statistics on the instantaneous motion direction of all pixels in the target motion response area and determine the main motion direction vector as the instantaneous spatial displacement direction of the target motion response area.

[0052] After determining the spatial location and contour of the target motion response region, the motion direction estimation module of the dynamic primitive decomposition branch estimates the instantaneous motion direction of each target motion response region.

[0053] The motion direction estimation module first calculates the inter-frame displacement vector of each pixel within the target's motion response region, from the previous video frame to the next. This displacement vector is calculated using a sparse optical flow tracing algorithm. Specifically, for a pixel within the target's motion response region, the module extracts a small pixel block centered on that pixel in the previous video frame. Then, within a search window near the same spatial location in the next video frame, it searches for the pixel block most similar to that pixel block in terms of brightness distribution pattern using pattern matching. The difference between the center coordinates of the two matching pixel blocks is the inter-frame displacement vector of that pixel. This displacement vector is a two-dimensional vector containing horizontal and vertical displacement components; its direction indicates the instantaneous motion direction of the pixel, and its length indicates the instantaneous motion amplitude of the pixel.

[0054] After obtaining the displacement vectors of all pixels within the target motion response area, the motion direction estimation module constructs a direction distribution histogram for main motion direction analysis. This module uniformly divides the circumferential space into several preset direction intervals. For each pixel within the target motion response area, the module determines its direction interval based on the direction angle of its displacement vector, and uses the length of the displacement vector as a voting weight, accumulating it into the statistical value of the corresponding direction interval. After traversing all pixels within the area, a direction distribution histogram is obtained, where the height of each bar in the histogram represents the cumulative motion amplitude of all moving pixels within that direction interval.

[0055] This module scans the direction distribution histogram to find the peak direction interval with the largest cumulative motion amplitude. The module then takes the center direction angle corresponding to this peak direction interval and constructs a two-dimensional unit vector using this angle. The horizontal and vertical components of this unit vector are determined by the cosine and sine values ​​of this direction angle, respectively. This two-dimensional unit vector is the main motion direction vector of the target's motion response region and is used as the instantaneous spatial displacement direction of the target's motion response region.

[0056] Step S1234: Estimate the instantaneous motion amplitude of the target motion response region. Based on the spatial distribution range of pixel brightness change in the target motion response region along the main motion direction vector, calculate the spatial extension scale of the target motion response region in the main motion direction vector direction, which is used as the instantaneous spatial displacement amplitude of the target motion response region.

[0057] After obtaining the instantaneous spatial displacement direction of the target motion response region, the motion amplitude estimation module estimates the instantaneous motion amplitude.

[0058] The motion amplitude estimation module first projects the spatial coordinates of all pixels within the target motion response region onto a projection axis parallel to the instantaneous spatial displacement direction. The projection is calculated by taking the dot product of each pixel's two-dimensional spatial coordinates with the corresponding two-dimensional unit vector along the instantaneous spatial displacement direction, thus obtaining the pixel's one-dimensional projected coordinate value on the projection axis. Then, the module statistically analyzes the projected coordinate values ​​of all pixels, identifying the maximum and minimum projected coordinate values. The difference between these two values ​​is then calculated; this difference represents the spatial extension scale of the target motion response region along the instantaneous spatial displacement direction.

[0059] This spatial extension scale comprehensively reflects the spatial dimensions of the target motion response region and the magnitude of its displacement in the direction of motion. This module directly uses this spatial extension scale as the instantaneous spatial displacement amplitude of the target motion response region. The instantaneous spatial displacement amplitude, together with the instantaneous spatial displacement direction, fully describes the instantaneous motion vector of the target motion response region between two video frames.

[0060] Step S1235: Based on the spatial contour, instantaneous spatial displacement direction and instantaneous spatial displacement amplitude of the target motion response region, generate the initial dynamic target primitive components of the video segment unit. The initial dynamic target primitive components describe the instantaneous state of the visual change elements caused by motion between two video frames.

[0061] For a processed target motion response region, the motion primitive packaging module of the dynamic primitive decomposition branch assembles its three core attributes into a motion primitive data structure. This motion primitive data structure contains three fields. The first field is the contour mask field, which stores a two-dimensional binary matrix with the same spatial size as the original video frame. The value is the maximum value at pixel positions within the target motion response region and the minimum value at pixel positions outside the region. The second field is the motion direction vector field, which stores a two-dimensional floating-point vector recording the instantaneous spatial displacement direction of the target motion response region. The third field is the motion amplitude scalar field, which stores a floating-point scalar recording the instantaneous spatial displacement amplitude of the target motion response region.

[0062] The dynamic primitive decomposition branch performs steps S1231 to S1234 on all adjacent frame pairs in the video frame time series, thereby generating a series of motion primitive data structures. The same moving object is continuously detected in several consecutive adjacent frame pairs, generating a series of corresponding motion primitive data structures. These motion primitive data structures are arranged in the order of their respective frame pair timestamps, collectively constituting the initial dynamic target primitive component of the video segment unit. The initial dynamic target primitive component is a time-seriesd set of motion state descriptions, recording the instantaneous state of each movement of all detected foreground moving targets in the camera field within the time range of the video segment unit.

[0063] Step S124: Input the initial static scene primitive components and the initial dynamic target primitive components into the primitive cooperative constraint layer in the pre-constructed visual primitive separation network. Use the primitive cooperative constraint layer to perform a decomposition consistency check on the initial static scene primitive components and the initial dynamic target primitive components, which are mutually constrained. According to the spatial distribution boundary of the initial static scene primitive components, remove the moving visual elements in the initial dynamic target primitive components that exceed the spatial distribution boundary to obtain the dynamic target primitive components. Based on the spatial area occupied by the removed dynamic target primitive components, complete the visual elements in the initial static scene primitive components that overlap with the spatial area to obtain the static scene primitive components. The decomposition consistency check makes the static scene primitive components and the dynamic target primitive components complementary and non-overlapping in spatial composition.

[0064] After the static primitive decomposition branch and the dynamic primitive decomposition branch output the initial static scene primitive components and the initial dynamic target primitive components respectively, the primitive co-constraint layer of the visual primitive separation network receives these two initial results and performs a decomposition consistency check on them with mutual constraints.

[0065] The primitive collaborative constraint layer first performs dynamic-to-static constraint verification. This layer traverses each motion primitive data structure in the initial dynamic target primitive component, extracting its contour mask field, which defines the spatial region occupied by the moving object in the foreground. For each pixel position within this contour mask, the layer queries the 2D descriptor graph mesh of the initial static scene primitive component to see if that position is marked as a scene structure stable unit. If, within the area covered by the contour mask, more than a certain proportion of pixel positions correspond to scene structure stable units in the initial static scene primitive component, it indicates that the detected "moving target" actually falls on a previously confirmed stable background region. This inconsistency suggests that the motion detection result may be a false detection, such as a false motion caused by changes in lighting or swaying leaves. For such motion primitive data structures, the primitive collaborative constraint layer directly removes them from the initial dynamic target primitive component.

[0066] Then, the primitive collaborative constraint layer performs static-to-dynamic constraint verification. After the above culling operation, the remaining motion primitive data structures are considered reliable foreground moving target detection results. The pixel regions covered by the contour masks of these motion primitive data structures are the regions that have actually been occupied by foreground moving targets in the entire video segment unit, and there should be no stable background information in these regions. The primitive collaborative constraint layer collects the contour masks of all retained motion primitive data structures and merges them into a global foreground occupancy map, which marks all pixel positions traversed by the moving object. On the two-dimensional descriptor map mesh of the initial static scene primitive components, the primitive collaborative constraint layer forcibly marks the background descriptor vectors corresponding to all pixel positions marked by the foreground occupancy map as invalid or missing. Subsequently, this layer performs spatial neighborhood background completion on these invalid positions, and the completion method is the same as the spatial neighborhood interpolation completion strategy described in step S1224, that is, using the surrounding valid background descriptor vectors for weighted interpolation filling.

[0067] After co-constraints and corrections in two directions, the primitive co-constraint layer outputs corrected static scene primitive components and corrected dynamic target primitive components. The corrected static scene primitive components are more spatially complete, with areas previously occluded by the foreground target correctly filled in. The corrected dynamic target primitive components are purer, eliminating false motion responses. Both achieve a completely complementary and non-overlapping ideal decomposition state in spatial composition.

[0068] Step S125: Integrate static scene primitive components and dynamic target primitive components to generate a set of visual primitive components for the corresponding video segment unit.

[0069] The output integration module of the visual primitive separation network receives the corrected static scene primitive components and the corrected dynamic target primitive components from the primitive cooperative constraint layer, and packages these two components into a structured set of visual primitive components. This set of visual primitive components, as a complete visual decomposition result, is associated with the corresponding video segment unit. At this point, the processing flow of a video segment unit in the visual primitive separation network is complete.

[0070] Step S130: Input the static scene primitive components and the dynamic target primitive components into the pre-constructed spatiotemporal primitive coupling coding network. Through the primitive component interaction mapping layer in the spatiotemporal primitive coupling coding network, perform cross-primary component co-coding on the static scene primitive components and the dynamic target primitive components to generate a spatiotemporal coupling feature representation with spatiotemporal coupling relationship between the static scene primitive components and the dynamic target primitive components.

[0071] After completing the visual primitive decomposition of video segment units and obtaining the set of visual primitive components, the method in this embodiment enters the spatiotemporal coupling coding stage. The core objective of this stage is to perform correlation analysis on independent static scene primitive components and dynamic target primitive components, capture the deep spatiotemporal coupling relationship between them, and generate a unified spatiotemporal coupling feature representation. This process is completed by a pre-built and trained spatiotemporal primitive coupling coding network.

[0072] Step S131: Input the static scene primitive components and dynamic target primitive components into the first feature mapping channel and the second feature mapping channel of the spatiotemporal primitive coupling coding network. Map the static scene primitive components to the static primitive feature space through the first feature mapping channel to obtain the static primitive feature vector. Map the dynamic target primitive components to the dynamic primitive feature space through the second feature mapping channel to obtain the dynamic primitive feature vector. The static primitive feature space represents the attribute distribution of static visual elements, and the dynamic primitive feature space represents the instantaneous state of dynamic visual elements.

[0073] The spatiotemporal primitive coupling coding network comprises a first feature mapping channel and a second feature mapping channel, which receive input in parallel. The first feature mapping channel specifically processes static scene primitive components. This channel contains a deep convolutional neural network module, which consists of several convolutional layers, batch normalization layers, and modified linear unit activation layers stacked alternately. This module receives a three-dimensional tensor of the static scene primitive components as input and, through layer-by-layer convolution and nonlinear activation operations, gradually maps the static scene primitive components from their original visual attribute descriptor feature space to a higher-dimensional, more abstract static primitive feature space. During this mapping process, the spatial resolution of the feature maps may gradually decrease, while the number of feature channels gradually increases. Finally, the first feature mapping channel outputs a static primitive feature vector. This static primitive feature vector is a three-dimensional tensor whose spatial dimension distribution corresponds to the video frame space. Each feature map in its channel dimension encodes a high-level distribution of static scene semantic attributes, such as "accessible road surface area attributes," "building facade area attributes," and "green vegetation area attributes."

[0074] The second feature mapping channel specifically handles dynamic target primitive components. Since the dynamic target primitive components are time-series data based on motion primitive data structures, the second feature mapping channel needs to convert them into dense feature tensors compatible with the spatial dimensions of static primitive feature vectors. Internally, this channel first uses a motion feature encoder to independently encode each motion primitive data structure. This motion feature encoder is a small fully connected network that receives the motion direction vector field and motion amplitude scalar field from the motion primitive data structure as input. After transformation through several fully connected layers, it generates a fixed-length motion state embedding vector. Then, this channel uses a spatial scattering module to fill the corresponding region of a two-dimensional feature map with the same spatial size as the static primitive feature vector with the motion state embedding vector of each motion primitive data structure, according to the spatial position indicated by its contour mask field. Each pixel within the region is filled with the same motion state embedding vector. For regions without any moving targets, their feature values ​​are filled with zeros. This channel then passes the entire spatially scattered two-dimensional feature map through several convolutional layers for spatial context smoothing and feature transformation, and the final output is the dynamic primitive feature vector. The dynamic primitive feature vector is also a three-dimensional tensor with the same spatial dimension as the static primitive feature vector. In the channel dimension, it encodes the instantaneous motion state and spatial distribution of moving visual elements in the scene.

[0075] Step S132: Input the static primitive feature vector and the dynamic primitive feature vector into the primitive component interaction mapping layer. In the primitive component interaction mapping layer, perform interaction weight calculation on the static primitive feature vector and the dynamic primitive feature vector to generate an interaction response intensity distribution matrix that characterizes the degree of spatial positional correlation between static scene primitive components and dynamic target primitive components.

[0076] The outputs of the first and second feature mapping channels, namely the static primitive feature vector and the dynamic primitive feature vector, are fed together into the core computational module of the spatiotemporal primitive coupling coding network—the primitive component interaction mapping layer. The core task of the primitive component interaction mapping layer is to explicitly calculate the degree of correlation between each pair of spatial locations of static and dynamic primitives, and to quantify this correlation using an interaction response intensity distribution matrix.

[0077] Step S1321: Perform spatial dimension embedding transformation on the static primitive feature vector to generate a static primitive embedding query vector, and perform spatial dimension embedding transformation on the dynamic primitive feature vector to generate a dynamic primitive embedding key vector.

[0078] The primitive component interaction mapping layer includes an embedding transformation module. This module performs a first embedding transformation on the input static primitive feature vector. This first embedding transformation is implemented through a single-layer convolutional operation. This convolutional layer has several convolutional kernels, each with a spatial size equal to the size of its center point; that is, a linear transformation is performed independently only for each spatial location. This convolutional operation traverses each spatial location of the static primitive feature vector, multiplying the feature vector at that location across the entire channel dimension with the parameter matrix of the convolutional kernel to generate a feature vector with a new channel dimension. This newly generated feature vector is named the static primitive embedding query vector. The spatial dimension of the static primitive embedding query vector is completely identical to that of the static primitive feature vector, while the channel dimension is transformed to a preset embedding dimension.

[0079] Simultaneously, the embedding transformation module performs a second embedding transformation on the input dynamic primitive feature vector. This second embedding transformation is also implemented through a single-layer convolution operation with different parameter matrices. This convolution operation traverses each spatial position of the dynamic primitive feature vector, multiplying the channel feature vector at that position with the parameter matrices of the second set of convolution kernels to generate a dynamic primitive embedding key vector. The channel dimension of the dynamic primitive embedding key vector is also transformed to the same preset embedding dimension.

[0080] Step S1322: Calculate the similarity response between the static primitive embedding query vector and the dynamic primitive embedding key value vector in the feature space, and generate an initial interaction weight map based on the similarity response. Each weight position in the initial interaction weight map corresponds to the point-to-point association degree of the spatial position in the static primitive feature vector and the dynamic primitive feature vector.

[0081] After the embedding transformation is completed, the similarity calculation module of the primitive component interaction mapping layer begins to calculate the degree of association between spatial locations. This module uses a dot product attention mechanism to calculate the similarity response. Specifically, this module constructs a two-dimensional output matrix, called the initial interaction weight graph. The height of this graph is determined by the spatial height of the static primitive embedding query vector, and the width of this graph is determined by the spatial width of the dynamic primitive embedding key value vector.

[0082] For each element position in the initial interaction weight graph, this position corresponds to a spatial coordinate on a static primitive and a spatial coordinate on a dynamic primitive. This module extracts an embedded query vector fragment from the static primitive's embedded query vector at the corresponding static spatial coordinate, and simultaneously extracts an embedded key-value vector fragment from the dynamic primitive's embedded key-value vector at the corresponding dynamic spatial coordinate. Both vector fragments have the same dimension. This module calculates the dot product of these two vector fragments, that is, multiplies the elements of the corresponding dimensions of the two vector fragments, and then sums the results of the multiplications across all dimensions to obtain a scalar value. This scalar value directly reflects the similarity between the visual features of these two spatial positions from different sources in the embedding space; the higher the similarity, the stronger the association between the static and dynamic positions. After the dot product calculation is completed for all element positions, the resulting complete two-dimensional matrix is ​​the initial interaction weight graph.

[0083] Step S1323: Perform spatial neighborhood context aggregation on the initial interaction weight graph. Taking each weight position in the initial interaction weight graph as the center, extract the weight distribution within its preset spatial neighborhood range, and perform centralized aggregation on the weight distribution to generate a neighborhood-enhanced interaction weight graph that takes into account spatial neighborhood information.

[0084] To make weight estimation more robust to local spatial shifts, the spatial context aggregation module of the primitive component interaction mapping layer performs a neighborhood enhancement operation on the initial interaction weight map. This module uses a two-dimensional convolutional kernel with fixed weights to perform a convolution operation on the initial interaction weight map. The size of this two-dimensional convolutional kernel is a preset small square, for example, covering nine weight positions in three rows and three columns. The weight values ​​at the center of the convolutional kernel are set relatively high, and the weight values ​​at the edge positions are set relatively low. All weight values ​​are normalized so that the sum is a single unit value. At each weight position, the convolution operation multiplies the nine weight values ​​within the kernel's coverage area by the weight coefficient of the corresponding position in the kernel, and then sums the nine products. The sum is used as the new weight value after aggregation at that center position. This process smooths the interaction weights between adjacent positions, ensuring that the weight value at any point incorporates the interaction strength information of its surrounding neighborhood. After neighborhood aggregation, the output weight map is called the neighborhood-enhanced interaction weight map.

[0085] Step S1324: Normalize the neighborhood enhanced interaction weight graph and map the weight values ​​in the neighborhood enhanced interaction weight graph to a preset numerical distribution range to generate an interaction response intensity distribution matrix that represents the distribution law of spatial position correlation between static scene primitive components and dynamic target primitive components.

[0086] To ensure comparability of weights corresponding to different dynamic locations, the primitive component interaction mapping layer normalizes the neighborhood enhancement interaction weight graph. This normalization operation is performed along the spatial dimension of the dynamic primitive feature vectors. For each row in the neighborhood enhancement interaction weight graph (corresponding to a fixed static spatial location), all element values ​​in that row (corresponding to the associated weights with all dynamic spatial locations) are processed by a flexible maximum transfer function. This function first exponentiates each element value in the row, then sums the exponents of all elements in the row, and finally divides each element's exponent by this sum. After processing, each element in the row becomes a positive number between the minimum value of zero and the maximum value of a single unit, and the sum of all element values ​​in the row equals the single unit value.

[0087] After applying the flexible maximum transfer function to all rows, the resulting output matrix is ​​the interaction response intensity distribution matrix. The value of each element in this matrix directly represents the probability or strength of the interaction between a spatial location on a static scene primitive and a spatial location on a dynamic target primitive. The closer the value is to a single unit value, the stronger the association. Overall, this interaction response intensity distribution matrix reflects the spatial distribution pattern of the association between static and dynamic primitives.

[0088] Step S133: Based on the interaction response intensity distribution matrix, the static primitive feature vector and the dynamic primitive feature vector are co-modulated. The dynamic primitive information is injected into the static primitive feature vector using the interaction response intensity distribution matrix to obtain the modulated static primitive feature vector. At the same time, the static primitive information is injected into the dynamic primitive feature vector using the interaction response intensity distribution matrix to obtain the modulated dynamic primitive feature vector.

[0089] After generating the interaction response intensity distribution matrix, the collaborative modulation module of the primitive component interaction mapping layer uses this matrix to perform information injection and feature modulation on the original static primitive feature vector and dynamic primitive feature vector, respectively.

[0090] The first step is the modulation process of injecting dynamic information into static features. This module performs a third linear transformation on the dynamic primitive embedding key vector to generate a dynamic primitive embedding value vector. This transformation is achieved through a single-layer convolution, with the output having the same spatial size as the input, and the channel dimension transformed to another preset dimension. Next, the module uses the interaction response intensity distribution matrix as weights to perform weighted aggregation on the dynamic primitive embedding value vector. Specifically, for each spatial location on the static primitive feature vector, the module extracts the row of weight vectors corresponding to that static location from the interaction response intensity distribution matrix. Each element of this row of weight vectors is used as a coefficient and multiplied by the embedding value vector of the corresponding spatial location in the dynamic primitive embedding value vector. Then, all weighted embedding value vectors are summed along the spatial dimension to obtain an aggregated global dynamic context vector that condenses all dynamic information related to that static location. This global dynamic context vector is dimension-adapted through a fully connected layer and then injected into the original static primitive feature vector at that static location through a concatenation operation along the channel dimension, forming a modulated static primitive feature vector.

[0091] Next is the modulation process of injecting static information into dynamic features. This process is symmetrical to the one described above. The co-modulation module performs another linear transformation on the static primitive embedding query vector to generate a static primitive embedding value vector. Then, the module transposes the interaction response intensity distribution matrix and uses the transposed matrix as weights to perform weighted aggregation of the static primitive embedding value vectors. For each spatial location on the dynamic primitive feature vector, the module uses the corresponding column of weights in the transposed weight matrix to perform a weighted summation of the static primitive embedding value vectors at all static spatial locations, obtaining a condensed global static context vector. This global static context vector, after dimensional adaptation, is concatenated by channel dimension and injected into the original dynamic primitive feature vector at the corresponding dynamic location, forming the modulated dynamic primitive feature vector.

[0092] After co-modulation, the two feature vectors obtained are no longer isolated descriptions of the background or motion, but rather, while retaining their own information, they incorporate rich contextual information that interacts with each other.

[0093] Step S134: The modulated static primitive feature vector and the modulated dynamic primitive feature vector are spliced ​​and fused along the feature channel dimension to generate a spatiotemporal coupling feature representation that contains coupling information of static scene primitive components and dynamic target primitive components. The spatiotemporal coupling feature representation contains the spatial dependency and temporal coordination relationship between static primitives and dynamic primitives, which is enhanced by the interaction response intensity distribution matrix.

[0094] After generating the modulated static primitive feature vector and the modulated dynamic primitive feature vector, the feature fusion module of the spatiotemporal primitive coupling coding network performs final feature fusion on the two feature vectors. This embodiment uses channel-dimensional concatenation for fusion, rather than element-wise addition. Channel-dimensional concatenation refers to concatenating the modulated static primitive feature vector and the modulated dynamic primitive feature vector along their respective channel dimensions. Since the spatial height and spatial width dimensions of the two feature vectors are pre-consistent, they can be directly concatenated along the channel dimensions. The spatial size of the resulting new feature tensor remains unchanged, while its channel dimension is equal to the sum of the number of channels in the modulated static primitive feature vector and the number of channels in the modulated dynamic primitive feature vector. This new feature tensor is the spatiotemporal coupling feature representation of the final output of the spatiotemporal primitive coupling coding network for this video segment unit. This spatiotemporal coupling feature representation fully encodes the complex spatial dependency and temporal coordination relationship between the static scene background and the dynamic moving target in the video segment unit, enhanced by the interaction response intensity distribution matrix.

[0095] Step S140: Input the spatiotemporal coupling feature representation into the pre-constructed abnormal situation inference network. Perform multi-directional situation inference on the spatiotemporal coupling feature representation through the situation differentiation branch in the abnormal situation inference network to generate a multi-directional situation differentiation trajectory set corresponding to the video segment unit. The multi-directional situation differentiation trajectory set describes the various situation trends that the visual primitive components in the video segment unit may evolve in the next video segment unit.

[0096] After obtaining the spatiotemporal coupling characteristics of the current video segment unit, the anomaly situation inference network performs multi-directional situation inference for the future. The core design concept of the anomaly situation inference network is that instead of making a single deterministic prediction of the future evolution of the current scene, it infers multiple representative and possible situational branches, thereby predicting the occurrence of potential abnormal behaviors.

[0097] Step S141: Input the spatiotemporal coupling feature representation into the situation initialization layer of the abnormal situation inference network, encode the initial situation state of the spatiotemporal coupling feature representation, and generate the initial situation state vector of the video segment unit at the current time step. The initial situation state vector integrates the current spatiotemporal coupling state of the static scene primitive components and the dynamic target primitive components within the video segment unit.

[0098] The situation initialization layer within the abnormal situation simulation network receives spatiotemporal coupled feature representations as input. This layer consists of a global average pooling layer and several fully connected layers connected sequentially. The global average pooling layer averages the input 3D spatiotemporal coupled feature representations in both height and width dimensions, compressing the spatial dimensions into a single value and outputting a feature vector that retains only the channel dimension. This feature vector is then fed into a mapping network composed of multiple fully connected layers and a feedforward activation function, where it is further encoded into a fixed-length feature vector. The length of this vector is predetermined by the model design of the situation simulation network. This encoded fixed-length feature vector is the initial situation state vector. This initial situation state vector comprehensively represents the overall situation state of the video segment unit within the current time step, composed of the static scene primitives and the dynamic target primitives, within a compressed semantic space.

[0099] Step S142: Input the initial situation state vector into multiple parallel situation inference subnetworks in the situation differentiation branch. Each parallel situation inference subnetwork adopts different situation evolution assumption parameters. The different situation evolution assumption parameters correspond to the different possible motion trends of dynamic target primitive components in the next time step and different interaction modes with static scene primitive components.

[0100] After the situation initialization layer generates the initial situation state vector, this vector is copied and fed into multiple parallel situation extrapolation subnetworks in the situation differentiation branch. The number of these parallel situation extrapolation subnetworks is pre-defined. Each parallel situation extrapolation subnetwork has an identical network topology, but contains a unique set of learnable situation evolution hypothesis parameters. The situation evolution hypothesis parameters are a set of network weight parameters; different parameter configurations cause the subnetwork to tend to predict different evolution trajectories when extrapolating future situations. Some subnetworks' hypothesis parameters tend to predict a situation trajectory where the moving target continues to move at a constant velocity in a straight line; some subnetworks' hypothesis parameters tend to predict a situation trajectory where the moving target is about to change direction; and some subnetworks' hypothesis parameters tend to predict a situation trajectory where the moving target is about to come into contact with a static scene element. Each subnetwork independently performs extrapolation based on its own hypothesis parameters.

[0101] Step S143: In each parallel situation inference sub-network, based on the corresponding situation evolution assumption parameters, a one-step situation evolution calculation is performed on the initial situation state vector to generate the next time step situation inference vector under the corresponding situation evolution assumption. The one-step situation evolution calculation simulates the position changes of dynamic target primitive components and the changes in the interaction relationship between dynamic target primitive components and static scene primitive components under the situation evolution assumption parameters.

[0102] After receiving the initial situation state vector, each parallel situation inference subnetwork performs a situation evolution calculation to predict the situation state for the next future time step from the current situation state.

[0103] Step S1431: Obtain the situation evolution hypothesis parameters corresponding to the parallel situation inference sub-network. The situation evolution hypothesis parameters include the directional offset hypothesis of the simulated motion trend and the interaction intensity change hypothesis of the simulated interaction mode change with the static scene primitive components.

[0104] A parallel situational inference subnetwork reads its own situational evolution hypothesis parameter tensor from its parameter storage area. This situational evolution hypothesis parameter tensor is a one-dimensional vector consisting of two contiguous vector segments. The first vector segment is the direction offset hypothesis, which contains several floating-point numbers and represents the subnetwork's potential tendency to respond to changes in the heading of a dynamic target. The second vector segment is the interaction intensity change hypothesis, also containing several floating-point numbers and representing the subnetwork's potential tendency to respond to changes in the interaction pattern between the dynamic target and the static scene.

[0105] Step S1432: The dynamic feature components corresponding to the dynamic target primitive components in the initial situational state vector are fused with the direction offset hypothesis to generate the offset dynamic feature components. The fusion simulates the feature changes of the dynamic target primitive components after position offset in the direction indicated by the direction offset hypothesis.

[0106] The parallel situational awareness simulation subnetwork contains a feature decoupling module that linearly decomposes the initial situational state vector into a dynamic feature component vector and a static feature component vector. The dynamic feature component vector is considered to encode state information closely related to the moving target in the current situation.

[0107] The motion migration prediction module of this sub-network fuses the dynamic feature component vector with the orientation migration hypothesis. The fusion operation involves concatenating the two vectors through channels, and then inputting the concatenated long vector into a small neural network consisting of fully connected layers and activation functions. The output of this small neural network is the offset dynamic feature component. This neural network acts as a conditional transformer, learning how to perform corresponding spatial migration simulation transformations on the features of the current motion state based on the motion trend indicated by the orientation migration hypothesis.

[0108] Step S1433: The static feature components corresponding to the static scene primitives in the initial situational state vector are fused with the interaction intensity change hypothesis to generate the changed static feature components. The fusion simulates the corresponding adjustment of the feature attributes of the static scene primitives under the assumption of a change in the interaction mode.

[0109] Parallel to step S1432, the state interaction adjustment module of this sub-network fuses the static feature component vectors extracted in step S1432 with the interaction intensity change hypothesis. The fusion operation is also performed by concatenating the two vectors through channels and inputting them into another small transform neural network. This small transform neural network learns the ability to modify static scene features based on the interaction intensity change hypothesis. For example, when the interaction intensity change hypothesis tends to simulate a scenario where "a moving target is about to enter a certain area," this network adjusts the response pattern of the features in that area within the static feature components, generating the modified static feature components.

[0110] Step S1434: Perform secondary interactive coding on the offset dynamic feature components and the changed static feature components to generate simulated coupling features of dynamic target primitive components and static scene primitive components in the next time step under the situation evolution assumption parameters. The secondary interactive coding recalculates the coupling relationship between the two in the new assumed position and interaction mode based on the offset dynamic feature components and the changed static feature components.

[0111] After obtaining the hypothetical shifted dynamic feature components and changed static feature components, the secondary interactive coding module within this sub-network recouples them. The construction of this secondary interactive coding module references the design of the primitive component interaction mapping layer in a spatiotemporal primitive coupling coding network, also using an attention mechanism to calculate the new interaction relationship generated by the two components under the hypothetical new state. This secondary interactive coding module calculates the interaction response intensity distribution matrix between the shifted dynamic feature components and the changed static feature components, and uses this matrix to aggregate information from the changed static feature components, injecting it into the shifted dynamic feature components, ultimately generating a simulated coupling feature. This simulated coupling feature characterizes the cooperative state of the dynamic target and the static scene in the next time step under the situational evolution assumptions of this sub-network.

[0112] Step S1435: Map the simulated coupling features back to the situation state space to generate the situation inference vector for the next time step under the situation evolution assumption parameters. The situation inference vector for the next time step represents the situation state that may occur in the next time step of the video segment unit under a specific situation evolution assumption.

[0113] The output mapping module of the parallel situational inference subnetwork inputs the simulated coupling features into a fully connected output layer. The output dimension of this fully connected layer is completely consistent with the dimension of the initial situational state vector. After mapping by this fully connected layer, the simulated coupling features are transformed back into the situational state space, forming a situational inference vector for the next time step. This situational inference vector for the next time step is semantically and structurally compatible with the initial situational state vector, representing the possible situational state of the video segment unit in the next future time step under the assumed situational evolution parameters.

[0114] Step S144: Integrate the next time step situational inference vectors generated by all parallel situational inference sub-networks to form a multi-directional situational inference state set corresponding to the video segment unit in the next time step.

[0115] After all parallel situational awareness subnetworks have completed one step of the scenario, all the generated situational awareness vectors for the next time step are collected into a set container. This set container is the multi-directional situational awareness state set. Each element in this multi-directional situational awareness state set corresponds to a future situational state under a different situational evolution assumption. For example, for a pedestrian walking in a transportation hub plaza, this multi-directional situational awareness state set may include a situational state of continuing to walk straight forward, a situational state of turning left, a situational state of turning right, a situational state of increasing walking speed, and a situational state of decreasing walking speed, etc.

[0116] Step S145: Connect each next time step situation projection vector in the multi-directional situation projection state set with the initial situation state vector of the current time step of the video segment unit to form multiple situation differentiation trajectories that start from the current situation state and point to different next time step situation states. All situation differentiation trajectories constitute the multi-directional situation differentiation trajectory set corresponding to the video segment unit.

[0117] The trajectory assembly module of the abnormal situation simulation network organizes all discrete situational state vectors in the multi-directional situational simulation state set into directed trajectory forms. For each situational simulation vector at the next time step in the set, the trajectory assembly module combines it with the initial situational state vector at the current time step to form a trajectory data pair. This trajectory data pair is an ordered binary tuple, where the first element is the starting state vector, i.e., the initial situational state vector; and the second element is the ending state vector, i.e., the situational simulation vector at a certain next time step. Each such binary tuple constitutes a brief situational differentiation trajectory, representing a possible situational evolution path starting from the current state. The situational differentiation trajectories generated by all parallel situational simulation sub-networks are aggregated to form the complete set of situational differentiation trajectories for this video segment unit, i.e., the multi-directional situational differentiation trajectory set.

[0118] Step S150: Perform situation deviation measurement on each situation differentiation trajectory in the multi-directional situation differentiation trajectory set and the normal situation trajectory in the preset normal situation trajectory template library, generate situation deviation distribution features, and identify the abnormal situation types contained in the video segment unit based on the situation deviation distribution features. The preset normal situation trajectory template library is constructed from normal monitoring video streams that do not contain abnormal events.

[0119] After the abnormal situation simulation network outputs a set of multi-directional situation differentiation trajectories corresponding to the current video segment unit, the method in this embodiment enters the final situation deviation measurement and anomaly identification stage. The core of this anomaly identification stage is to compare all possible future situation evolution paths deduced by the model with a pre-built normal situation trajectory template library representing "normal" behavior, and determine whether an anomaly exists and its category by quantifying the degree of deviation.

[0120] Step S151: Obtain a preset normal situation trajectory template library. The normal situation trajectory template library stores multiple normal situation trajectory templates. Each normal situation trajectory template corresponds to a normal situation path in a monitoring video stream that does not contain abnormal events, showing the evolution of visual primitive components over time.

[0121] Before performing anomaly identification, the method in this embodiment pre-constructs a normal situation trajectory template library. The construction process of this normal situation trajectory template library is independent of the online inference process. During construction, the video analysis server acquires a large number of historical normal monitoring video streams from the same or similar monitoring scenarios. These video streams are manually reviewed or pre-screened by algorithms to confirm that they do not contain any abnormal events. For these normal monitoring video streams, the same processing flow from steps S110 to S140 is executed to extract the standard situation differentiation trajectory of each video segment unit after multi-directional situation differentiation. For each typical normal event evolution pattern, such as "pedestrians walking straight along the sidewalk at a constant speed," "vehicles slowing down and stopping along the roadway," and "pedestrians standing still and waiting," the server extracts representative situation state change sequences from a large number of normal trajectories using a clustering algorithm as templates and stores them in the normal situation trajectory template library. Each normal situation trajectory template is essentially a situation state change sequence, describing a complete normal situation's state evolution path from beginning to end.

[0122] Step S152: For each situation differentiation trajectory in the multi-directional situation differentiation trajectory set, extract the situation state change sequence on its trajectory path.

[0123] During the online inference phase, for the set of multi-directional situation differentiation trajectories generated by the current video segment unit, the situation deviation measurement module traverses each situation differentiation trajectory in the set. Since a situation differentiation trajectory is an ordered pair consisting of a starting vector and an ending vector, it itself constitutes a shortest situational state change sequence containing only two state points. For each such starting-to-ending situational state change sequence, the situation deviation measurement module extracts it and prepares it for morphological comparison with the templates in the normal situational trajectory template library.

[0124] Step S153: Compare the trajectory shape of the situation change sequence with each normal situation trajectory template in the normal situation trajectory template library, calculate the trajectory shape deviation between the situation change sequence and each normal situation trajectory template, and measure the similarity of the path shape of the two trajectories in the situation space.

[0125] The situation deviation measurement module will extract the current situation state change sequence and compare the trajectory shape with each normal situation trajectory template stored in the normal situation trajectory template library.

[0126] For example, step S1531: Extract key trajectory state points from the situation state change sequence. Extract several key trajectory state points that can represent the overall trajectory shape and direction from the situation state change sequence to form a key situation state change sequence.

[0127] Since the current situational state change sequence in this embodiment only contains two state points, the starting point and the ending point, and these two state points are themselves key state points that characterize the overall shape and direction of the short trajectory, the key situational state change sequence is the sequence composed of the two state vectors of the starting point and the ending point.

[0128] Step S1532: For each normal situation trajectory template in the normal situation trajectory template library, extract the template key state point sequence corresponding to the normal situation trajectory template.

[0129] For a normal situation trajectory template in the normal situation trajectory template library, it is a relatively long time series containing situation state vectors for multiple consecutive time steps. The situation deviation measurement module performs key state point extraction on this normal situation trajectory template. This key state point extraction can be accomplished through uniform sampling or extraction of local extrema, ultimately forming a template key state point sequence. The template key state point sequence is a sparse representation of the original template sequence, but it still retains the main trajectory and shape of the original template in the situation state space.

[0130] Step S1533: Perform inter-sequence state correspondence matching between the key situation state change sequence and each template key state point sequence, find the best matching point of each trajectory key state point in the key situation state change sequence in the template key state point sequence, and establish the state point correspondence relationship.

[0131] The situation deviation measurement module employs a dynamic time warping algorithm or a variant thereof to align the current key situation state change sequence (containing two state points, a start point and an end point) with a template key state point sequence (containing multiple state points). The dynamic time warping algorithm automatically establishes the correspondence between state points in two sequences by finding the curved path that minimizes the cumulative distance between them. For example, the start point of the current trajectory might match an intermediate state point in the template sequence, while the end point of the current trajectory might match another subsequent state point in the template sequence.

[0132] Step S1534: For each pair of trajectory key state points and template key state points that establish a correspondence between state points, calculate the state vector deviation in the situational state space. The state vector deviation includes the direction deviation and amplitude deviation between the trajectory key state points and the template key state points in the situational state space.

[0133] For each pair of state points matched by the dynamic time warping algorithm, the situation deviation measurement module calculates their state vector deviation in the situation state space. This module calculates two deviation components. The first is the direction deviation, which is obtained by calculating the cosine similarity between the current trajectory point's state vector and the matching template point's state vector, and then subtracting this cosine similarity from the unit value. The direction deviation quantifies the difference in direction between the two state vectors in the situation state space. The second is the amplitude deviation, which is obtained by calculating the absolute value of the difference in magnitude between the two state vectors, dividing it by the magnitude of one of the vectors, and then normalizing. The amplitude deviation quantifies the difference in energy or intensity between the two state vectors. The weighted sum of the two components constitutes the comprehensive state vector deviation for the pair of state points.

[0134] Step S1535: Integrate the state vector deviations of all state point correspondences to generate the trajectory shape deviation of the situation state change sequence relative to the normal situation trajectory template. The trajectory shape deviation reflects the degree to which the overall situation differentiation trajectory deviates from the normal situation trajectory template.

[0135] The situation deviation measurement module accumulates or averages the comprehensive state vector deviations between all state points in the current key situation state change sequence and their corresponding matching template points to obtain a scalar value. This scalar value represents the deviation of the current situation differentiation trajectory from the current normal situation trajectory template in terms of trajectory shape. The larger the value of this trajectory shape deviation, the greater the difference between the derived situation differentiation trajectory and the normal template, and the less it conforms to the normal pattern.

[0136] Step S154: Construct the situation deviation distribution characteristics of all trajectory shape deviations into video segment units. The situation deviation distribution characteristics reflect the overall degree of deviation between the situation differentiation trajectories in the multi-directional situation differentiation trajectory set and the normal situation trajectory.

[0137] The situational deviation measurement module calculates the deviation in trajectory shape from each normal situational trajectory template in the normal situational trajectory template library for each situational differentiation trajectory in the multi-directional situational differentiation trajectory set of the current video segment unit. This results in a set of trajectory shape deviation values. This set of trajectory shape deviation values ​​forms a distribution, which is the situational deviation distribution characteristic of the video segment unit. The situational deviation distribution characteristic comprehensively reflects the overall deviation of multiple future projection paths from all known normal behavior patterns.

[0138] Step S155: Analyze the situation deviation distribution characteristics. If the situation deviation distribution characteristics indicate that the trajectory shape deviation of all situation differentiation trajectories in the multi-directional situation differentiation trajectory set exceeds the preset normal deviation tolerance range, then it is determined that the video segment unit contains an abnormal situation. The abnormal situation type corresponding to the abnormal situation is determined according to the maximum distribution direction of the trajectory shape deviation. The abnormal situation type is associated with the normal situation type corresponding to the normal situation trajectory template with the largest deviation.

[0139] The anomaly detection module receives the situational deviation distribution features and performs anomaly identification based on preset judgment rules. This module first checks the minimum value in the situational deviation distribution features. If this minimum value is still greater than a preset normal deviation tolerance threshold, it means that even the future path most similar to all normal templates deduced has deviated beyond the normal range. This indicates that the dynamic situation of the current scene cannot be explained by any known normal behavior pattern, and the anomaly detection module therefore determines that the current video segment unit contains an abnormal situation.

[0140] After determining the existence of an abnormal situation, the anomaly detection module further identifies the type of abnormal situation. This module examines the distribution of trajectory deviation and identifies the normal situation trajectory template with the smallest average trajectory deviation (although still greater than a threshold). In other words, it determines which type of normal situation the current situation differs from the least, or in other words, which type of normal situation the current situation has deviated from. Since each type of normal situation trajectory template is associated with a semantic tag, such as "pedestrian walking normally," "vehicles driving normally," or "crowds gathering normally," the anomaly detection module combines the semantic tag corresponding to the closest normal situation trajectory template with "abnormal deviation" to generate the final anomaly type tag, such as "pedestrian walking normally with abnormal deviation" or "vehicles driving normally with abnormal deviation." Finally, the anomaly detection module packages and outputs this anomaly type tag along with the corresponding video segment unit's identifier information, completing the anomaly identification process for the entire video segment unit.

[0141] Figure 2 The illustration shows exemplary hardware and software components of a video surveillance anomaly detection system 100 for public safety, which can implement the ideas of this application, according to some embodiments of this application. For example, a processor 120 can be used in the video surveillance anomaly detection system 100 for public safety and to perform the functions described in this application.

[0142] The video surveillance anomaly identification system 100 applied to public safety can be a general-purpose server or a special-purpose server; both can be used to implement the video surveillance anomaly identification method for public safety of this application. Although only one server is shown in this application, for convenience, the functions described in this application can be implemented in a distributed manner on multiple similar platforms to balance the load.

[0143] For example, a video surveillance anomaly detection system 100 applied to public safety may include a network port 110 connected to a network, one or more processors 120 for executing program instructions, a communication bus 130, and various forms of storage media 140, such as a disk, ROM, or RAM, or any combination thereof. Exemplarily, the video surveillance anomaly detection system 100 applied to public safety may also include program instructions stored in ROM, RAM, or other types of non-transitory storage media, or any combination thereof. The methods of this application can be implemented according to these program instructions. The video surveillance anomaly detection system 100 applied to public safety also includes an I / O interface 150 between the computer and other input / output devices.

[0144] For ease of explanation, only one processor is described in the video surveillance anomaly detection system 100 for public safety applications. However, it should be noted that the video surveillance anomaly detection system 100 for public safety applications in this application may also include multiple processors. Therefore, the steps performed by one processor described in this application may also be performed jointly or individually by multiple processors. For example, if the processor of the video surveillance anomaly detection system 100 for public safety applications performs steps A and B, it should be understood that steps A and B may also be performed jointly by two different processors or individually by one processor. For example, the first processor performs step A, the second processor performs step B, or the first processor and the second processor jointly perform steps A and B.

[0145] Furthermore, this embodiment of the invention also provides a readable storage medium, wherein computer-executable instructions are preset in the readable storage medium, and when the processor executes the computer-executable instructions, the above-mentioned video monitoring anomaly identification method applied to public safety is implemented.

[0146] It should be noted that, in order to simplify the description of the present invention and thus help to understand one or more embodiments of the invention, multiple features may sometimes be grouped into one embodiment, drawing or description thereof in the foregoing description of the embodiments of the present invention.

Claims

1. A video surveillance anomaly identification method applied to public safety, characterized in that, The method includes: The original monitoring video stream is acquired by the image acquisition equipment deployed in the public safety monitoring area. The original monitoring video stream is divided into a set of video segment units with temporal continuity. Each video segment unit in the set of video segment units carries an acquisition timestamp identifier and an acquisition location identifier. Visual primitive decomposition is performed on each video segment unit in the video segment unit set to obtain the visual primitive component set of the corresponding video segment unit. The visual primitive component set includes static scene primitive components and dynamic target primitive components. The static scene primitive components represent the structured visual element distribution of the background scene in the video segment unit, and the dynamic target primitive components represent the instantaneous motion visual elements of the foreground moving target in the video segment unit. The static scene primitive components and dynamic target primitive components are input into a pre-constructed spatiotemporal primitive coupling coding network. Through the primitive component interaction mapping layer in the spatiotemporal primitive coupling coding network, the static scene primitive components and dynamic target primitive components are co-coded across primitive components to generate a spatiotemporal coupling feature representation with spatiotemporal coupling relationship between static scene primitive components and dynamic target primitive components. The spatiotemporal coupling feature representation is input into a pre-constructed abnormal situation inference network. The spatiotemporal coupling feature representation is used to perform multi-directional situation inference through the situation differentiation branch in the abnormal situation inference network, generating a set of multi-directional situation differentiation trajectories corresponding to video segment units. The set of multi-directional situation differentiation trajectories describes the various situation trends that the visual primitive components in the video segment unit may evolve in the next video segment unit. For each situation differentiation trajectory in the multi-directional situation differentiation trajectory set, a situation deviation measurement is performed on the normal situation trajectory in the preset normal situation trajectory template library, a situation deviation distribution feature is generated, and the abnormal situation type contained in the video segment unit is identified based on the situation deviation distribution feature. The preset normal situation trajectory template library is constructed from normal monitoring video streams that do not contain abnormal events.

2. The video monitoring anomaly identification method for public safety according to claim 1, characterized in that, For each video segment unit in the set of video segment units, visual primitive decomposition is performed to obtain the set of visual primitive components for the corresponding video segment unit. The set of visual primitive components includes static scene primitive components and dynamic target primitive components, including: Frame sequence parsing is performed on each video segment unit in the set of video segment units to obtain the time sequence of video frames that constitute the video segment unit; The static primitive decomposition branch in the pre-built visual primitive separation network is invoked to perform scene structure analysis on each video frame in the video frame time series. The distribution pattern of visual elements in each video frame that do not change position over time is extracted as the initial static scene primitive components of the video segment unit. Scene structure analysis includes dividing each video frame into spatial regions and determining the repetition pattern of visual elements in each divided region. The dynamic primitive decomposition branch in the pre-built visual primitive separation network is invoked to extract the instantaneous representation of moving visual elements in the visual difference region between adjacent video frames in the video frame time series, generating the initial dynamic target primitive components of the video segment unit. The instantaneous representation extraction includes capturing the region where the pixel brightness changes between adjacent video frames and determining the instantaneous spatial displacement direction and instantaneous spatial displacement amplitude of the region. The initial static scene primitives and the initial dynamic target primitives are input into the primitive co-constraint layer in the pre-constructed visual primitive separation network. The primitive co-constraint layer is used to perform a decomposition consistency check on the initial static scene primitives and the initial dynamic target primitives, which are mutually constrained. Based on the spatial distribution boundary of the initial static scene primitives, moving visual elements in the initial dynamic target primitives that exceed the spatial distribution boundary are removed to obtain the dynamic target primitives. Based on the spatial area occupied by the removed dynamic target primitives, visual elements in the initial static scene primitives that overlap with the spatial area are filled in to obtain the static scene primitives. The decomposition consistency check makes the static scene primitives and the dynamic target primitives complementary and non-overlapping in spatial composition. Integrate static scene primitives and dynamic target primitives to generate a set of visual primitives for the corresponding video segment unit.

3. The video monitoring anomaly identification method for public safety according to claim 2, characterized in that, The dynamic primitive decomposition branch in the pre-built visual primitive separation network is invoked to extract instantaneous representations of moving visual elements in the visual difference regions between adjacent video frames in the video frame time series, generating the initial dynamic target primitive components of the video segment unit, including: Pixel-level difference is performed on two temporally consecutive video frames in a video frame time series to generate a difference response map. Pixel-level difference captures the change in pixel brightness at each pixel position between the two video frames. The response region of the differential response map is enhanced. The set of pixel locations in the differential response map where the pixel brightness change exceeds the preset response threshold is determined as the target motion response region. The target motion response region represents the concentrated area of ​​pixel brightness change caused by the foreground moving target between two video frames. The instantaneous motion direction of the target motion response area is estimated. Based on the position offset of the corresponding pixel positions in two video frames between frames, the instantaneous motion direction of each pixel in the target motion response area is determined. The instantaneous motion direction of all pixels in the target motion response area is statistically distributed to determine the main motion direction vector as the instantaneous spatial displacement direction of the target motion response area. The instantaneous motion amplitude of the target motion response region is estimated. Based on the spatial distribution range of pixel brightness change in the target motion response region in the main motion direction vector direction, the spatial extension scale of the target motion response region in the main motion direction vector direction is calculated as the instantaneous spatial displacement amplitude of the target motion response region. Based on the spatial contour, instantaneous spatial displacement direction, and instantaneous spatial displacement amplitude of the target motion response region, the initial dynamic target primitive components of the video segment unit are generated. The initial dynamic target primitive components describe the instantaneous state of the visual change elements caused by motion between two video frames.

4. The video monitoring anomaly identification method for public safety according to claim 2, characterized in that, The static primitive decomposition branch of the pre-built visual primitive separation network is invoked to perform scene structure analysis on each video frame in the video frame time series. The distribution pattern of visual elements whose positions do not change over time is extracted from each video frame and used as the initial static scene primitive components of the video segment unit, including: Hierarchical spatial partitioning is performed on each video frame in the video frame time series. Each video frame is divided into a set of spatial grid units with different scales according to a multi-level spatial partitioning strategy. The multi-level spatial partitioning strategy includes spatial region segmentation with different granularities. Visual attribute statistics are performed on visual elements within each spatial grid cell, and the texture direction distribution, color distribution, and edge direction distribution within each spatial grid cell are extracted as local visual attribute descriptors. Temporal consistency analysis is performed on the local visual attribute descriptors of spatial grid cells at the same spatial location in all video frames in the video frame time series. Stable spatial grid cells with temporal volatility of local visual attribute descriptors below a preset volatility threshold within the time span covered by the video frame time series are selected as stable units of scene structure. The scene structure stable unit and the local visual attribute descriptor corresponding to the scene structure stable unit are spatially spliced ​​according to the original spatial position of the scene structure stable unit in the video frame to construct a static scene primitive component that represents the distribution of structured visual elements of the background scene in the video segment unit. The static scene primitive component records the distribution of scene elements whose visual attributes remain unchanged within the time span of the video segment unit. The constructed static scene primitives are used as the initial static scene primitives for video segment units.

5. The video monitoring anomaly identification method for public safety according to claim 1, characterized in that, Static scene primitives and dynamic target primitives are input into a pre-constructed spatiotemporal primitive coupling coding network. Through a primitive component interaction mapping layer within this network, cross-primitive co-coding is performed on the static scene primitives and dynamic target primitives, generating a spatiotemporal coupling feature representation that reflects the spatiotemporal coupling relationship between them. This representation includes: The static scene primitive components and the dynamic target primitive components are input into the first feature mapping channel and the second feature mapping channel of the spatiotemporal primitive coupling coding network. The static scene primitive components are mapped to the static primitive feature space through the first feature mapping channel to obtain the static primitive feature vector. The dynamic target primitive components are mapped to the dynamic primitive feature space through the second feature mapping channel to obtain the dynamic primitive feature vector. The static primitive feature space represents the attribute distribution of static visual elements, and the dynamic primitive feature space represents the instantaneous state of dynamic visual elements. The static primitive feature vector and the dynamic primitive feature vector are input into the primitive component interaction mapping layer. In the primitive component interaction mapping layer, the interaction weights of the static primitive feature vector and the dynamic primitive feature vector are calculated to generate an interaction response intensity distribution matrix that characterizes the degree of spatial positional correlation between the static scene primitive components and the dynamic target primitive components. Based on the interaction response intensity distribution matrix, the static primitive feature vector and the dynamic primitive feature vector are co-modulated. The dynamic primitive information is injected into the static primitive feature vector using the interaction response intensity distribution matrix to obtain the modulated static primitive feature vector. At the same time, the static primitive information is injected into the dynamic primitive feature vector using the interaction response intensity distribution matrix to obtain the modulated dynamic primitive feature vector. The modulated static primitive feature vector and the modulated dynamic primitive feature vector are spliced ​​and fused along the feature channel dimension to generate a spatiotemporal coupled feature representation that contains coupling information of static scene primitive components and dynamic target primitive components. The spatiotemporal coupled feature representation contains the spatial dependency and temporal coordination relationship between static primitives and dynamic primitives, which is enhanced by the interaction response intensity distribution matrix.

6. The video monitoring anomaly identification method for public safety according to claim 5, characterized in that, The static and dynamic primitive feature vectors are input into the primitive component interaction mapping layer. In this layer, interaction weights are calculated for both vectors to generate an interaction response intensity distribution matrix that characterizes the spatial correlation between static scene primitive components and dynamic target primitive components. This matrix includes: Perform spatial dimension embedding transformation on static primitive feature vectors to generate static primitive embedding query vectors, and perform spatial dimension embedding transformation on dynamic primitive feature vectors to generate dynamic primitive embedding key vectors. Calculate the similarity response between the static primitive embedding query vector and the dynamic primitive embedding key vector in the feature space, and generate an initial interaction weight map based on the similarity response. Each weight position in the initial interaction weight map corresponds to the point-to-point association degree of the spatial position in the static primitive feature vector and the dynamic primitive feature vector. Spatial neighborhood context aggregation is performed on the initial interaction weight graph. Taking each weight position in the initial interaction weight graph as the center, the weight distribution within its preset spatial neighborhood is extracted, and the weight distribution is centrally aggregated to generate a neighborhood-enhanced interaction weight graph that takes into account spatial neighborhood information. The neighborhood enhanced interaction weight graph is normalized and mapped to a preset numerical distribution range, generating an interaction response intensity distribution matrix that represents the distribution law of spatial position correlation between static scene primitive components and dynamic target primitive components. The value of the interaction response intensity distribution matrix directly reflects the strength of the interaction between a certain spatial position in the static scene primitive component and a certain spatial position in the dynamic target primitive component. The interaction response intensity distribution matrix is ​​used as the output of the primitive component interaction mapping layer.

7. The video monitoring anomaly identification method for public safety according to claim 1, characterized in that, The spatiotemporal coupling feature representation is input into a pre-constructed abnormal situation inference network. The network then performs multi-directional situation inference on the spatiotemporal coupling feature representation through situation differentiation branches, generating a set of multi-directional situation differentiation trajectories corresponding to each video segment unit, including: The spatiotemporal coupling feature representation is input into the situation initialization layer of the abnormal situation inference network. The spatiotemporal coupling feature representation is encoded into an initial situation state to generate an initial situation state vector of the video segment unit at the current time step. The initial situation state vector integrates the current spatiotemporal coupling state of the static scene primitive components and the dynamic target primitive components within the video segment unit. The initial situation state vector is input into multiple parallel situation inference subnetworks in the situation differentiation branch. Each parallel situation inference subnetwork adopts different situation evolution assumption parameters. The different situation evolution assumption parameters correspond to the different possible motion trends of dynamic target primitive components in the next time step and different interaction modes with static scene primitive components. In each parallel situational simulation subnetwork, based on the corresponding situational evolution assumption parameters, a one-step situational evolution calculation is performed on the initial situational state vector to generate the next time step situational simulation vector under the corresponding situational evolution assumption. The one-step situational evolution calculation simulates the positional changes of dynamic target primitive components and the changes in the interaction relationship between dynamic target primitive components and static scene primitive components under the situational evolution assumption parameters. Integrate the next time step situational projection vectors generated by all parallel situational projection sub-networks to form a multi-directional situational projection state set corresponding to the video segment unit in the next time step; Each next time step situation projection vector in the multi-directional situation projection state set is connected to the initial situation state vector of the current time step of the video segment unit to form multiple situation differentiation trajectories that start from the current situation state and point to different next time step situation states. All situation differentiation trajectories constitute the multi-directional situation differentiation trajectory set corresponding to the video segment unit.

8. The video monitoring anomaly identification method for public safety according to claim 7, characterized in that, In each parallel situational estimation subnetwork, based on the corresponding situational evolution assumption parameters, a one-step situational evolution calculation is performed on the initial situational state vector to generate the next time-step situational estimation vector under the corresponding situational evolution assumption, including: Obtain the situation evolution hypothesis parameters corresponding to the parallel situation inference subnetwork. The situation evolution hypothesis parameters include the directional offset hypothesis of simulating motion trend and the interaction intensity change hypothesis of simulating the interaction mode change with static scene primitive components. The dynamic feature components corresponding to the dynamic target primitive components in the initial situational state vector are fused with the direction offset hypothesis to generate the offset dynamic feature components. The fusion simulates the feature changes of the dynamic target primitive components after their position shifts in the direction indicated by the direction offset hypothesis. The static feature components corresponding to the static scene primitives in the initial situational state vector are fused with the interaction intensity change hypothesis to generate the changed static feature components. The fusion simulates the corresponding adjustment of the feature attributes of the static scene primitives under the assumption of a change in the interaction mode. The offset dynamic feature components and the changed static feature components are subjected to secondary interactive coding to generate simulated coupling features of dynamic target primitive components and static scene primitive components in the next time step under the situation evolution assumption parameters. The secondary interactive coding recalculates the coupling relationship between the two in the new assumed position and interaction mode based on the offset dynamic feature components and the changed static feature components. The simulated coupling features are mapped back to the situation state space to generate the situation inference vector for the next time step under the situation evolution assumption parameters. The situation inference vector for the next time step represents the situation state that may occur in the next time step of the video segment unit under a specific situation evolution assumption.

9. The video monitoring anomaly identification method for public safety according to claim 1, characterized in that, For each situation differentiation trajectory in the multi-directional situation differentiation trajectory set, a situation deviation measurement is performed between it and the normal situation trajectory in the preset normal situation trajectory template library. Situation deviation distribution features are generated, and based on these features, the abnormal situation types contained in the video segment unit are identified, including: Obtain a preset normal situation trajectory template library. The normal situation trajectory template library stores multiple normal situation trajectory templates. Each normal situation trajectory template corresponds to a normal situation path in a monitoring video stream that does not contain abnormal events, showing the evolution of visual primitive components over time. For each situation differentiation trajectory in the multi-directional situation differentiation trajectory set, extract the situation state change sequence along its trajectory path; The trajectory shape is compared with each normal situation trajectory template in the normal situation trajectory template library for the sequence of changes in situation status. The deviation of trajectory shape between the sequence of changes in situation status and each normal situation trajectory template is calculated. The trajectory shape comparison measures the similarity of the path shape of the two trajectories in the situation status space. The situation deviation distribution characteristics of all trajectory shape deviations constitute video segment units. The situation deviation distribution characteristics reflect the overall degree of deviation between the situation differentiation trajectory and the normal situation trajectory in the multi-directional situation differentiation trajectory set. The distribution characteristics of the situation deviation are analyzed. If the distribution characteristics of the situation deviation indicate that the deviation of the trajectory shape of all situation differentiation trajectories in the multi-directional situation differentiation trajectory set exceeds the preset normal deviation tolerance range, then the video segment unit is determined to contain an abnormal situation. The abnormal situation type corresponding to the abnormal situation is determined according to the maximum distribution direction of the trajectory shape deviation. The abnormal situation type is associated with the normal situation type corresponding to the normal situation trajectory template with the largest deviation.

10. A video monitoring anomaly identification system for public safety applications, characterized in that, The video surveillance anomaly identification system for public safety includes a processor and a memory, the memory and the processor being connected. The memory is used to store programs, instructions or code, and the processor is used to execute the programs, instructions or code in the memory to implement the video surveillance anomaly identification method for public safety as described in any one of claims 1-9.