Network camera monitoring identification method and system based on artificial intelligence

By performing spatiotemporal context encoding processing and dynamic feature matching comparison on network camera monitoring data, anomaly confidence scores are generated, which solves the problem of accurate identification of abnormal behavior patterns in existing technologies and achieves accurate judgment and timely warning of abnormal situations.

CN120673315AActive Publication Date: 2025-09-19SICHUAN XINSAIHU INTERNET OF THINGS TECHNOLOGY CO LTD

Patent Information

Application Number
CN202510786095.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-12
Publication Date
2025-09-19
Estimated Expiration
2045-06-12

AI Technical Summary

Technical Problem

Existing network camera surveillance recognition methods ignore information in the time dimension and time series metadata, making it difficult to accurately identify abnormal behavior patterns. In addition, surveillance data under different environmental conditions is difficult to standardize, analyze, and compare, reducing the accuracy and reliability of surveillance recognition.

Method used

The system acquires a surveillance data stream consisting of video frame sequences and corresponding time-series metadata sequences, performs spatiotemporal context encoding on the data stream, and maps the spatial features of the video frames with the temporal features of the time-series metadata to generate a contextual feature cube. Next, it performs normal behavior pattern learning, extracting stable feature patterns across multiple time periods using an unsupervised learning algorithm to generate a baseline feature library. Finally, it dynamically compares the contextual feature cube for the current time period with the baseline feature library, calculates the feature matching deviation, and generates an anomaly confidence score.

Benefits of technology

It achieves accurate judgment of abnormal situations in monitoring scenes, generates timely and accurate monitoring warning instructions, and improves the accuracy and reliability of network camera monitoring identification.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120673315A_ABST
    Figure CN120673315A_ABST
Patent Text Reader

Abstract

The invention provides a network camera monitoring identification method and system based on artificial intelligence, and the method comprises the steps: firstly obtaining a monitoring data stream which is outputted by a network camera and comprises a video frame sequence and a corresponding time sequence metadata sequence, and then carrying out the spatial-temporal context coding processing of the monitoring data stream; the method comprises the following steps: generating a context feature cube containing spatial position information and time evolution information, then executing a normal behavior mode learning operation based on the context feature cube, and generating a reference feature library containing typical scene feature templates and feature evolution rule description; and dynamically matching and comparing the context feature cube of the current time period with the reference feature library, calculating a feature matching deviation value and generating an abnormal confidence score, and finally generating a monitoring early warning instruction containing abnormal occurrence time, a space coordinate range and a confidence level identifier according to the abnormal confidence score and corresponding space-time position information. And the accuracy and the early warning effect of monitoring and identification of the network camera are effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and in particular to an artificial intelligence-based network camera monitoring and recognition method and system. Background Art

[0002] Network cameras, as essential data acquisition devices, play a key role in numerous fields, including security surveillance, traffic management, and industrial production monitoring. Existing network camera surveillance recognition methods primarily focus on simple analysis of video frames, typically extracting only spatial features such as object shape and color, while ignoring temporal information and camera-related temporal metadata, such as acquisition timestamps, camera pose information, and ambient light intensity.

[0003] This single analysis approach has numerous limitations. First, due to the lack of correlation with temporal features, it is difficult to accurately identify abnormal behavior patterns that evolve over time. For example, slowly occurring abnormal changes or periodic anomalies can be easily missed. Second, the inability to fully utilize temporal metadata makes it difficult to effectively standardize and compare surveillance data under different environmental conditions (such as varying light intensities and camera positions), significantly compromising the accuracy and reliability of surveillance recognition. Therefore, a network camera surveillance recognition method that comprehensively considers spatial, temporal, and temporal metadata is needed to improve surveillance recognition effectiveness. Summary of the Invention

[0004] In view of the above-mentioned problems, in combination with the first aspect of the present invention, an embodiment of the present invention provides an artificial intelligence-based network camera monitoring and recognition method, the method comprising: Obtain a monitoring data stream output by a network camera, wherein the monitoring data stream includes a continuously acquired video frame sequence and a corresponding time-series metadata sequence, wherein the time-series metadata sequence includes an acquisition timestamp of each frame of video, camera posture information, and a description of the ambient light intensity; Performing spatiotemporal context coding on the monitoring data stream, associating and mapping the spatial features of the video frame sequence with the temporal features of the temporal metadata sequence, and generating a context feature cube containing spatial position information and temporal evolution information; Performing normal behavior pattern learning operations based on the context feature cube, extracting feature patterns that appear stably over multiple time periods through an unsupervised learning algorithm, and generating a baseline feature library containing typical scene feature templates and feature evolution law descriptions; Dynamically matching and comparing the context feature cube of the current period with the benchmark feature library, calculating the feature matching deviation value and generating an anomaly confidence score; A monitoring and warning instruction is generated according to the abnormality confidence score and the corresponding spatiotemporal location information, wherein the monitoring and warning instruction includes the abnormality occurrence time, spatial coordinate range and confidence level identifier.

[0005] On the other hand, an embodiment of the present invention also provides an artificial intelligence-based network camera monitoring and identification system, including a processor and a machine-readable storage medium, wherein the machine-readable storage medium is connected to the processor, the machine-readable storage medium is used to store programs, instructions or codes, and the processor is used to execute the programs, instructions or codes in the machine-readable storage medium to implement the above method.

[0006] Based on the above aspects, an embodiment of the present invention obtains a monitoring data stream including a video frame sequence and a corresponding time-series metadata sequence, performs spatiotemporal context coding on the monitoring data stream, associates and maps the spatial features of the video frame sequence with the temporal features of the time-series metadata sequence, and generates a context feature cube including spatial location information and temporal evolution information. This effectively integrates multi-dimensional information and improves the integrity and accuracy of feature expression. Normal behavior pattern learning operations are performed based on the context feature cube. Feature patterns that appear stably in multiple time periods are extracted through an unsupervised learning algorithm. A baseline feature library containing typical scene feature templates and feature evolution law descriptions is generated. The context feature cube of the current time period is dynamically matched and compared with the baseline feature library, a feature matching deviation value is calculated, and an anomaly confidence score is generated. This allows accurate judgment of abnormal situations in the monitoring scene. Based on the anomaly confidence score and the corresponding spatiotemporal location information, a monitoring warning instruction is generated, including the time of anomaly occurrence, spatial coordinate range, and confidence level identifier. This achieves timely discovery, precise positioning, and accurate assessment of abnormal events, thereby improving the accuracy and reliability of network camera monitoring recognition. BRIEF DESCRIPTION OF THE DRAWINGS

[0007] Figure 1 The figure is a schematic diagram of the execution flow of the network camera monitoring and recognition method based on artificial intelligence provided by an embodiment of the present invention.

[0008] Figure 2 Schematic diagram of exemplary hardware and software components of an artificial intelligence-based network camera monitoring and recognition system provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0009] The present invention will be described in detail below with reference to the accompanying drawings. Figure 1 The figure is a flow chart of an artificial intelligence-based network camera monitoring and recognition method provided by an embodiment of the present invention. The artificial intelligence-based network camera monitoring and recognition method is introduced in detail below.

[0010] Step S110: Acquire a monitoring data stream output by a network camera, wherein the monitoring data stream includes a continuously acquired video frame sequence and a corresponding time series metadata sequence, wherein the time series metadata sequence includes an acquisition timestamp of each frame of video, camera posture information, and an ambient light intensity description.

[0011] Specifically, in a common security surveillance scenario, network cameras are deployed in a specific area for continuous monitoring. The network cameras capture the monitored area at a pre-set frame rate, generating continuous video frames. These frames are then arranged in the order they were captured, forming a video frame sequence.

[0012] For each video frame, the network camera system can synchronously record the corresponding time-series metadata. The acquisition timestamp is accurate to a specific point in time and is generated based on the camera's internal clock system. It is used to clearly identify when the video frame was captured. The camera's pose information is obtained by the positioning and attitude sensors installed on the camera. It contains the camera's position coordinates in three-dimensional space and the shooting angle information. This information can be represented by a set of vectors, where each dimension of the vector corresponds to a different direction and angle in space. The description of the ambient light intensity is measured by the light sensor. It reflects the lighting conditions in the monitored area at that moment and may be expressed as different levels or a continuous range of values. These video frame sequences and the corresponding time-series metadata sequences together constitute the monitoring data stream.

[0013] Step S120: performing spatiotemporal context coding processing on the monitoring data stream, associating and mapping the spatial features of the video frame sequence with the temporal features of the temporal metadata sequence, and generating a context feature cube containing spatial position information and temporal evolution information.

[0014] After acquiring the surveillance data stream, it needs to be subjected to spatiotemporal context encoding. The goal is to associate and map the spatial features contained in the video frame sequence with the temporal features embodied in the temporal metadata sequence. This process generates a context feature cube that contains the spatial location information of the surveillance scene and its evolution over time.

[0015] Step S121: performing spatial feature extraction processing on the video frame sequence, using a convolutional neural network to extract local area features and global scene features of each frame of video, and generating a spatial feature vector sequence.

[0016] Using a convolutional neural network to extract spatial features from a video frame sequence is a key step. Convolutional neural networks are deep learning models specifically designed to process grid-structured data, making them ideal for processing two-dimensional image data such as video frames. For each frame in the sequence, the convolutional neural network performs a layer-by-layer feature extraction operation.

[0017] Step S1211: input the video frame sequence into a pre-trained basic feature extraction network, and extract shallow visual features through a convolutional layer. The shallow visual features include edge contours, texture details and color distribution information.

[0018] The pre-trained basic feature extraction network is trained based on a large amount of image data and has learned many common image features. When a video frame sequence is input into the basic feature extraction network, the convolution layer in the network performs a convolution operation on the video frame. The convolution operation slides a series of convolution kernels on the video frame and performs a weighted summation on each local area to extract features of different scales and directions. In this process, the convolution layer extracts shallow visual features of the video frame, such as edge contour features, which can identify the boundaries of objects; texture detail features, which can capture the texture information of the object surface; color distribution information, which reflects the distribution of colors in the video frame.

[0019] Step S1212: performing multi-scale aggregation processing on the shallow visual features through a spatial pyramid pooling module to generate local area features with different receptive fields.

[0020] The spatial pyramid pooling module performs multi-scale aggregation on the shallow visual features extracted previously. It divides the shallow visual features into regions of different sizes and performs a pooling operation on each region. Pooling can reduce the dimensionality of features while retaining important feature information. By dividing the features at different scales, the spatial pyramid pooling module can generate local region features with different receptive fields. The receptive field refers to the size of the input image that a neuron in a convolutional neural network can understand. Receptive fields of different sizes can capture object features at different scales. For example, a small receptive field can focus on the detailed features of an object, while a large receptive field can capture the overall shape and structure of the object. Thus, the spatial pyramid pooling module generates a series of local region features with different scales and receptive fields.

[0021] Step S1213: performing global information compression processing on the shallow visual features through a global average pooling layer to generate global scene features that describe the overall scene layout.

[0022] The global average pooling layer compresses global information based on shallow visual features. It averages the entire shallow visual feature map, averaging the feature values ​​of each channel to obtain a single value. In this way, the global average pooling layer compresses the global information in the shallow visual features to generate a global scene feature that describes the overall scene layout. This global scene feature reflects the overall distribution and structure of all objects and elements in the video frame, ignoring some local details and focusing on the overall scene characteristics.

[0023] Step S1214: cascade and splice the local region features with the global scene features to generate a spatial feature vector corresponding to each video frame.

[0024] After obtaining the local region features and global scene features, they need to be concatenated. Cascade concatenation involves concatenating two feature vectors in a set order to form a longer feature vector. For each video frame, the corresponding local region feature vector and global scene feature vector are concatenated to obtain the corresponding spatial feature vector. This spatial feature vector combines the detailed features of the local region with the overall features of the global scene, providing a more comprehensive description of the spatial characteristics of the video frame.

[0025] Step S1215: Arrange the spatial feature vectors according to the acquisition order of the video frames to generate a spatial feature vector sequence.

[0026] Finally, the spatial feature vectors corresponding to each video frame are arranged according to the acquisition order of the video frames. Since the video frames are acquired in chronological order, the spatial feature vector sequence obtained in this arrangement reflects the temporal changes in the spatial features of the video frame sequence.

[0027] Step S122: Perform time feature encoding processing on the temporal metadata sequence, convert the acquisition timestamp into a time interval sequence, convert the camera posture information into a spatial coordinate offset sequence, convert the ambient light intensity description into a light change gradient sequence, and generate a time feature vector sequence.

[0028] For time series metadata sequences, time feature encoding is required. By performing specific conversion operations on different types of time series metadata, they are converted into feature sequences suitable for subsequent processing.

[0029] Step S1221: Convert the acquisition timestamp into a time interval sequence.

[0030] The acquisition timestamp records the specific capture time of each video frame. To better reflect temporal changes, it is converted into a time interval sequence. Specifically, the time difference between two adjacent acquisition timestamps is calculated to obtain a series of time interval values. These time interval values ​​form a time interval sequence, which reflects the temporal rhythm and interval of video frame capture.

[0031] Step S1222: Convert the camera pose information into a sequence of spatial coordinate offsets.

[0032] Camera pose information includes the camera's position and angle in three-dimensional space. By calculating the difference between the camera pose information corresponding to two adjacent video frames, we can obtain a spatial coordinate offset. These spatial coordinate offsets are arranged in the order of the video frames to generate a spatial coordinate offset sequence. This spatial coordinate offset sequence reflects the changes in the camera's position and angle during the recording process.

[0033] Step S1223: Convert the ambient light intensity description into a light change gradient sequence.

[0034] The ambient light intensity description reflects the lighting conditions in the monitored area. The change between the ambient light intensity descriptions corresponding to two adjacent video frames is calculated to obtain the light change gradient. These light change gradients are arranged in the order of the video frames to generate a light change gradient sequence. This light change gradient sequence can reflect the temporal trend of the light intensity in the monitored area.

[0035] Step S1224: splicing the time interval sequence, the spatial coordinate offset sequence, and the illumination change gradient sequence to generate a time feature vector sequence.

[0036] The time interval sequence, spatial coordinate offset sequence, and illumination change gradient sequence obtained previously are spliced ​​together to form a comprehensive time feature vector sequence, which contains various change information of the time series metadata in the time dimension.

[0037] Step S123: establishing a time alignment relationship between the spatial feature vector sequence and the temporal feature vector sequence, taking the acquisition timestamp of each frame of video as the alignment reference, splicing and fusing the spatial feature vectors and temporal feature vectors of corresponding time points to generate a spatiotemporal fusion feature vector.

[0038] To correlate spatial and temporal features, it's necessary to establish a temporal alignment between the spatial and temporal feature vector sequences. Using the acquisition timestamp of each frame as the alignment reference, the spatial and temporal feature vectors are ensured to be temporally aligned. Spatial and temporal feature vectors at the same time point are concatenated and fused. This concatenation process concatenates two vectors in a predetermined order to form a new vector. The resulting spatiotemporal fusion feature vector combines both spatial and temporal features, providing a more comprehensive description of the spatial and temporal characteristics of the surveillance scene.

[0039] Step S124: Arrange the spatiotemporal fusion feature vectors in chronological order to construct a three-dimensional feature matrix, where the first dimension is the time step, the second dimension is the spatial position coordinate, and the third dimension is the feature dimension, thereby generating a context feature cube containing spatial position information and time evolution information.

[0040] The resulting spatiotemporal fusion feature vectors are arranged in chronological order to construct a three-dimensional feature matrix. The first dimension of this matrix represents the time step, reflecting the acquisition sequence and temporal changes of the video frames. The second dimension represents the spatial coordinates, corresponding to the different spatial locations in the monitoring area. The third dimension represents the feature dimension, which contains the individual feature components of the spatiotemporal fusion feature vectors. This construction generates a contextual feature cube that contains both spatial location information and temporal evolution information. This contextual feature cube can intuitively display the spatial and temporal feature changes of the monitoring scene.

[0041] Step S130: performing normal behavior pattern learning operations based on the context feature cube, extracting feature patterns that appear stably in multiple time periods through an unsupervised learning algorithm, and generating a benchmark feature library, which contains typical scene feature templates and feature evolution law descriptions.

[0042] After acquiring the contextual feature cube, we need to perform normal behavior pattern learning based on it. Specifically, we use an unsupervised learning algorithm to extract characteristic patterns that appear consistently over multiple time periods from the contextual feature cube. This generates a baseline feature library, which serves as an important reference for determining whether the monitored scene is abnormal.

[0043] Step S131: performing time window division processing on the context feature cube, and selecting sub-cubes of continuous time steps as learning units.

[0044] To facilitate analysis of characteristic patterns within the context feature cube, it must first be partitioned into time windows. Based on the time dimension, the context feature cube is divided into multiple sub-cubes at consecutive time steps. Each sub-cube is a learning unit. The size of the time window should be determined based on the actual monitoring scenario and data characteristics. An appropriate time window size ensures that a learning unit contains sufficient feature information to capture stable patterns, while not being too large, causing patterns to be blurred or ignored. For example, in a high-traffic shopping mall monitoring scenario, if the time window is too small, it may not cover a complete customer behavior cycle. If the time window is too large, different types of customer behavior may be mixed together, making it difficult to extract stable characteristic patterns. Each learning unit contains spatial and temporal feature information within a specific time period.

[0045] Step S132: Perform feature clustering analysis on each learning unit, and use a density clustering algorithm to identify significantly dense feature clusters, where the feature clusters represent stable feature patterns that recur over multiple time periods.

[0046] For each learning unit, a density clustering algorithm is used to perform feature clustering analysis. The core idea of ​​the density clustering algorithm is to perform clustering based on the density of data points. In the learning unit of the context feature cube, data points represent different feature vectors, which contain spatial and temporal feature information. The density clustering algorithm calculates the density around each data point and divides the high-density areas into different feature clusters. Significantly dense feature clusters represent stable feature patterns that recur over multiple time periods. For example, in a shopping mall monitoring scenario, there may be some fixed customer flow paths and stay areas. The feature vectors corresponding to these areas will cluster together in different time periods to form significantly dense feature clusters. These feature clusters can be accurately identified through the density clustering algorithm.

[0047] Step S133: performing feature statistical processing on the feature cluster, calculating the central vector of the feature cluster as a typical scene feature template, and calculating the distribution variance of the feature vectors in the feature cluster as a feature stability index.

[0048] After identifying feature clusters, they need to be statistically processed. For each feature cluster, its center vector is calculated. The center vector is obtained by summing the dimensions of all feature vectors within the feature cluster and taking the average. This center vector can be used as a typical scene feature template, representing the typical scene characteristics corresponding to the feature cluster. For example, in a shopping mall surveillance scenario, a feature cluster corresponds to customer activity near the checkout counter. Its center vector represents the typical characteristics of customer activity near the checkout counter, such as dwell time and range of activity. At the same time, the distribution variance of the feature vectors within the feature cluster is calculated. The distribution variance reflects the degree of dispersion of the feature vectors within the feature cluster. The smaller the variance, the more stable the feature vectors within the feature cluster. The distribution variance can be used as a feature stability indicator to measure the stability of the feature pattern. If the variance of a feature cluster is small, it means that the feature pattern is very stable over multiple time periods and is more likely to be part of a normal behavior pattern.

[0049] Step S134: performing correlation analysis on the feature clusters of adjacent time windows, tracking the evolution trajectory of the same feature cluster in the time dimension, calculating the time change rate and directional offset of the feature vector, and generating a description of the feature evolution law.

[0050] In order to understand the evolution of feature patterns over time, it is necessary to perform correlation analysis on feature clusters in adjacent time windows.

[0051] Step S1341: Acquire the feature cluster set of the previous time window and the feature cluster set of the current time window.

[0052] The corresponding feature cluster sets are extracted from two adjacent time windows respectively. These feature cluster sets contain the feature cluster information identified in different time windows, and each feature cluster consists of its center vector and internal feature vectors.

[0053] Step S1342: performing similarity calculation processing on each feature cluster of the previous time window and the feature cluster of the current time window, and using the cosine similarity algorithm to calculate the similarity of the cluster center vectors.

[0054] For each feature cluster in the previous time window and each feature cluster in the current time window, the cosine similarity algorithm is used to calculate the similarity between their cluster center vectors. Cosine similarity measures the similarity between two vectors by calculating the cosine of the angle between them. The calculation process involves first calculating the dot product of the two cluster center vectors and then dividing that dot product by the product of the two vectors' moduli. A cosine similarity value closer to 1 indicates that the two vectors are more similar; a value closer to -1 indicates that the two vectors are more opposite; and a value closer to 0 indicates that there is little correlation between the two vectors. This calculation can determine whether two feature clusters are likely manifestations of the same feature pattern at different times.

[0055] Step S1343: establishing a temporal association relationship between feature clusters based on the similarity calculation results, and marking feature clusters with a similarity greater than a preset threshold as temporal continuations of the same feature.

[0056] Based on the similarity calculation results, a preset threshold is set. If the similarity between two feature clusters exceeds the preset threshold, they are considered to be the continuation of the same feature at different times and are marked as having a temporal correlation. The preset threshold setting needs to be adjusted based on actual conditions. An appropriate threshold can accurately establish the temporal correlation between feature clusters. For example, in a shopping mall monitoring scenario, if the similarity between two feature clusters is greater than 0.8, they are considered to be the continuation of the same customer behavior pattern at different time periods.

[0057] Step S1344: For feature clusters with a time-continuing relationship, extract the sequence data of the central vector of the feature cluster in the time dimension to generate a feature evolution trajectory.

[0058] For feature clusters with temporal correlations, we extract the time-series data of their central vectors. These central vectors are arranged in chronological order to obtain the feature evolution trajectory of the feature cluster. This trajectory reflects the temporal changes in the feature pattern. For example, in a shopping mall monitoring scenario, a feature cluster may correspond to the movement paths of customers within the mall. Its trajectory can show the changes in the customer's movement direction and speed over different time periods.

[0059] Step S1345: performing first-order difference calculation processing on the feature evolution trajectory to obtain the feature change amount of adjacent time steps, and generating the feature time change rate in combination with the time interval information.

[0060] A first-order difference calculation is performed on the feature evolution trajectory. This difference is the difference between the feature vectors of two adjacent time steps. This difference represents the feature change between the adjacent time steps. Combined with the time interval information, the feature change is divided by the time interval to obtain the feature time rate of change. The feature time rate of change reflects the rate of change of the feature pattern over time. For example, in a shopping mall monitoring scenario, if a feature cluster corresponds to the length of time spent by a customer, the feature time rate of change can reflect the rate of change of the customer's stay time, indicating whether it is gradually increasing or decreasing.

[0061] Step S1346: performing direction vector calculation processing on the feature evolution trajectory, extracting the main direction component of the feature change, and generating a feature direction offset.

[0062] The directional vector of the feature evolution trajectory is calculated, and the main directional component of the feature change is extracted by analyzing the changes in the feature vector in different dimensions. Specifically, two adjacent feature vectors in the feature evolution trajectory are subtracted to obtain a vector, which is then normalized to obtain the directional vector. The main directional component in the directional vector is extracted. This main directional component is the feature directional offset, which reflects the direction of change of the feature pattern over time. For example, in a shopping mall monitoring scenario, if a feature cluster corresponds to the direction of customer movement, the feature directional offset can show the change in the customer's movement direction.

[0063] Step S1347: The feature time change rate and the feature direction offset are integrated to generate a feature evolution law description that describes the feature evolution law over time.

[0064] The feature time rate of change and the feature directional offset are combined to form a comprehensive description. By concatenating the feature time rate of change and the feature directional offset according to the set weights, a new vector is generated. This vector contains the temporal evolution of the feature pattern, including the speed and direction of change. This feature evolution description will be used as part of the baseline feature library for subsequent anomaly detection.

[0065] Step S135: Integrate and store the typical scene feature templates, feature stability indicators, and feature evolution law descriptions to generate a benchmark feature library.

[0066] Finally, the typical scenario feature templates, feature stability indicators, and feature evolution law descriptions are integrated and stored. For example, these can be stored in a database, with each feature cluster corresponding to a record containing the typical scenario feature template, feature stability indicators, and feature evolution law descriptions. Together, these records form a baseline feature library, which contains the typical features and evolution laws of normal behavior patterns in monitoring scenarios.

[0067] Step S140: dynamically matching and comparing the context feature cube of the current period with the reference feature library, calculating the feature matching deviation value and generating an anomaly confidence score.

[0068] After the baseline feature library is generated, the context feature cube of the current period needs to be dynamically matched and compared with the baseline feature library to detect whether there are any abnormalities.

[0069] Step S141: extracting a current feature vector sequence from the context feature cube of the current period, wherein the current feature vector sequence includes a spatial feature vector and a temporal feature vector of the current time step.

[0070] Extract the current feature vector sequence from the context feature cube for the current time period. Specifically, based on the time dimension of the context feature cube, extract the feature vector corresponding to the current time step. These feature vectors contain both spatial and temporal feature information at the current point in time. For example, in a shopping mall monitoring scenario, the spatial feature vector might include the distribution of people in different areas within the mall, while the temporal feature vector might include the current trend in foot traffic. Arrange these feature vectors in chronological order to obtain the current feature vector sequence.

[0071] Step S142: performing feature block processing on the current feature vector sequence, dividing the feature vectors of consecutive time steps into comparison units of the same size as the benchmark feature library learning units.

[0072] To facilitate matching and comparison with the baseline feature library, the current feature vector sequence is segmented. The feature vectors of consecutive time steps are divided into comparison units of the same size as the baseline feature library learning units. This ensures that the comparison units and learning units are consistent in time and feature dimensions, facilitating effective matching and comparison. For example, in a shopping mall surveillance scenario, if the baseline feature library learning units are segmented according to 10-minute time steps, the current feature vector sequence is also segmented according to 10-minute time steps, resulting in multiple comparison units.

[0073] Step S143: performing a benchmark matching operation on each comparison unit, calculating the Euclidean distance between the feature vector in the comparison unit and the corresponding typical scene feature template in the benchmark feature library, and generating a feature matching deviation value.

[0074] For each comparison unit, a benchmark matching operation is performed. The Euclidean distance between the feature vector within the comparison unit and the corresponding typical scene feature template in the benchmark feature library is calculated. Euclidean distance is a commonly used method for measuring the distance between two vectors, reflecting the degree of similarity between the two vectors in feature space. The calculation process involves first calculating the sum of the squares of the differences in the corresponding dimensions of the two vectors (dimensions must be aligned and unified before calculation), then taking the square root of the sum of squares to obtain the Euclidean distance. The calculated Euclidean distance represents the feature matching deviation. A larger deviation indicates a greater difference between the current feature vector and the typical scene feature template. For example, in a shopping mall surveillance scenario, if a comparison unit corresponds to human activity in a certain area within the mall, the Euclidean distance between the feature vector within the comparison unit and the corresponding typical scene feature template in the benchmark feature library is calculated. A larger Euclidean distance indicates that human activity in that area deviates significantly from normal patterns.

[0075] Step S144: performing time-weighted processing on the feature matching deviation values, allocating enhanced weights to the deviation values ​​of the most recent time steps, and generating a weighted deviation value sequence.

[0076] In order to pay more attention to the recent feature changes, the feature matching deviation value is time-weighted.

[0077] Step S1441: Determine the length parameter of the time-weighted window, where the length parameter is consistent with the time step of the benchmark feature library learning unit.

[0078] First, determine the length of the time-weighted window. This length should be consistent with the time step of the baseline feature library learning unit to ensure reasonable comparisons in the temporal dimension. For example, in a shopping mall surveillance scenario, if the baseline feature library learning unit is divided into 10-minute time steps, then the time-weighted window length should also be set to 10 minutes.

[0079] Step S1442: Generate an exponential decay weight coefficient according to the time distance between the time step and the current time point, wherein the exponential decay weight coefficient shows an exponentially decreasing trend as the time distance increases.

[0080] An exponentially decaying weight coefficient is generated based on the temporal distance between the time step and the current time point. Specifically, an exponential function is used to generate the weight coefficient, with the independent variable being the temporal distance between the time step and the current time point. As the temporal distance increases, the value of the exponential function decreases rapidly, thus achieving exponential decay of the weight. For example, time steps closer to the current time point have a larger corresponding weight coefficient, while time steps farther from the current time point have a smaller corresponding weight coefficient.

[0081] Step S1443: multiplying the feature matching deviation value by the exponential decay weight coefficient corresponding to the time step to generate a time-weighted deviation value.

[0082] After obtaining the exponentially decaying weight coefficient, each feature matching deviation value is multiplied by the weight coefficient of the corresponding time step. This gives more weight to recent feature matching deviation values, while less weight to more distant deviation values. This weighting process more accurately reflects the difference between the current monitoring scene and the normal pattern, as recent changes are often more valuable as a reference.

[0083] Step S1444: Arrange the time-weighted deviation values ​​in chronological order to generate a weighted deviation value sequence.

[0084] The deviation values ​​after time weighting processing are arranged in chronological order to form a weighted deviation value sequence. This weighted deviation value sequence comprehensively considers the feature matching deviation value and time factors, and can more comprehensively display the matching differences between the current monitoring scene and the benchmark feature library over a period of time.

[0085] Step S145: performing statistical analysis on the weighted deviation value sequence, calculating the mean, maximum value and variance of the deviation values, and generating a comprehensive deviation index.

[0086] Statistical analysis of the weighted deviation value sequence is performed to assess the degree of deviation from the normal pattern from multiple perspectives. The mean of the weighted deviation value sequence is calculated by summing all deviation values ​​in the sequence and dividing by the number of deviation values. The mean provides an average deviation level, reflecting the overall variance. The maximum value is calculated, which is the largest deviation value in the sequence. This identifies the point with the greatest deviation from the normal pattern during a period, helping to identify possible serious anomalies. The variance is calculated by first calculating the square of the difference between each deviation value and the mean, then summing these squared values ​​and dividing them by the number of deviation values. The variance measures the dispersion of the deviation values. A larger variance indicates greater fluctuation in the deviation values, potentially indicating unstable anomalies. These statistical values ​​are combined to generate a comprehensive deviation index, which more comprehensively reflects the degree of anomalies in the current monitoring scenario.

[0087] Step S146: Input the comprehensive deviation index into a preset confidence mapping function, and generate an abnormality confidence score based on the positive correlation between the deviation value and the abnormality possibility.

[0088] The preset confidence mapping function is based on extensive experimental data and practical application experience. It establishes a mapping relationship between the comprehensive deviation index and the anomaly confidence score. Since the deviation value is positively correlated with the probability of an anomaly, the larger the comprehensive deviation index, the greater the difference between the current monitoring scenario and the normal mode, and the higher the probability of an anomaly.

[0089] Step S1461: Obtain the comprehensive deviation index distribution data under historical normal scenarios, and calculate the deviation mean and deviation standard deviation under normal scenarios.

[0090] To accurately map the comprehensive deviation index to the anomaly confidence score, it is necessary to obtain historical data on the distribution of the comprehensive deviation index under normal circumstances. This data is collected when the monitoring scenario is normal and contains a large number of comprehensive deviation index values. By statistically analyzing this data, the mean deviation and standard deviation under normal circumstances are calculated. The mean deviation is calculated by summing all comprehensive deviation index values ​​and dividing by the number of data points. The standard deviation is calculated by first calculating the square of the difference between each comprehensive deviation index value and the mean deviation, then summing these squared values, dividing them by the number of data points, and finally taking the square root. The mean deviation reflects the average level of the comprehensive deviation index under normal circumstances, while the standard deviation reflects the dispersion of these deviation values ​​around the mean.

[0091] Step S1462: Taking the deviation mean as a reference, the difference between the current comprehensive deviation index and the deviation mean is divided by the deviation standard deviation to generate a standardized deviation value.

[0092] Using the mean deviation under normal scenarios as a benchmark, the difference between the current comprehensive deviation index and the mean deviation is divided by the standard deviation to obtain the standardized deviation value. This standardization process converts comprehensive deviation indicators of varying magnitudes into comparable values, eliminating the impact of differences in comprehensive deviation index magnitude across different monitoring scenarios or time periods. The standardized deviation value more accurately reflects the degree of deviation of the current monitoring scenario from the normal pattern.

[0093] Step S1463: converting the standardized deviation value into a confidence score in the interval [0, 1] by a nonlinear mapping method, wherein the larger the standardized deviation value, the higher the anomaly confidence score.

[0094] A nonlinear mapping method is used to convert the standardized deviation value into an anomaly confidence score in the interval [0, 1]. The nonlinear mapping function can be designed based on the actual situation. Typically, a monotonically increasing function is used, so that the larger the standardized deviation value, the higher the corresponding anomaly confidence score. For example, an S-shaped curve function can be used as a nonlinear mapping function. This nonlinear mapping function increases slowly when the standardized deviation value is small. As the standardized deviation value increases, the function value rises rapidly, eventually approaching 1. Through this nonlinear mapping, the standardized deviation value can be more reasonably converted into an anomaly confidence score, so that the score can accurately reflect the possibility of an anomaly occurring.

[0095] Step S150: Generate a monitoring warning instruction based on the abnormality confidence score and the corresponding spatiotemporal location information, wherein the monitoring warning instruction includes the abnormality occurrence time, spatial coordinate range and confidence level identifier.

[0096] After obtaining the anomaly confidence score and the corresponding spatiotemporal location information, it is necessary to generate monitoring and warning instructions based on this information. The monitoring and warning instructions should include the time and spatial coordinate range of the anomaly and the confidence level identification, so as to provide monitoring personnel with a comprehensive and accurate description of the anomaly so that they can take corresponding measures in a timely manner.

[0097] Step S151: extracting the timestamp information corresponding to the anomaly confidence score from the context feature cube to determine the start time and end time of the anomaly.

[0098] The context feature cube records the feature information corresponding to each time step and the associated anomaly confidence score. To extract the timestamp information corresponding to the anomaly confidence score, we first set a threshold for the anomaly confidence score. When the anomaly confidence score corresponding to a time step exceeds the threshold, it is considered a time point at which an anomaly may have occurred. Among these time points exceeding the threshold, the earliest time point is found as the start time of the anomaly, and the latest time point is found as the end time of the anomaly. For example, in a shopping mall monitoring scenario, if the anomaly confidence score threshold is set to 0.6, if the anomaly confidence score exceeds 0.6 from the 10th time step and continues to exceed 0.6 until the 20th time step, the timestamp corresponding to the 10th time step is the start time of the anomaly, and the timestamp corresponding to the 20th time step is the end time of the anomaly.

[0099] Step S152: Based on the spatial position information of the context feature cube, locate the spatial region where the feature matching deviation value exceeds a preset threshold, and generate the spatial coordinate range where the anomaly occurs.

[0100] The second dimension of the context feature cube represents the spatial position coordinates. For each spatial position coordinate, the corresponding result has been obtained when calculating the feature matching deviation value. Set a preset feature matching deviation value threshold, traverse the feature matching deviation values ​​corresponding to each spatial position in the context feature cube, and when the feature matching deviation value of a spatial position exceeds the threshold, it is marked as a spatial position where an anomaly may exist. Integrate all marked spatial positions, determine their boundary range, and thus generate the spatial coordinate range where the anomaly occurs. Taking shopping mall monitoring as an example, the shopping mall is divided into multiple areas, each area has corresponding spatial coordinates. When the feature matching deviation value of a certain area exceeds the preset threshold, the area is included in the abnormal spatial range. Finally, all abnormal areas are combined to obtain a spatial coordinate range containing multiple coordinate points, for example, represented by a rectangular area or a polygonal area.

[0101] Step S153: Divide the abnormality confidence score into multiple continuous intervals, each interval corresponds to a confidence level identifier, and the confidence level identifier increases as the confidence score increases.

[0102] Based on extensive experimental data and practical application experience, the anomaly confidence score is divided into multiple continuous intervals within the range of 0 to 1. For example, it can be divided into three intervals: [0, 0.3), [0.3, 0.7), and [0.7, 1]. Each interval is assigned a corresponding confidence level identifier, such as "mild anomaly" corresponds to the interval [0, 0.3), "moderate anomaly" corresponds to the interval [0.3, 0.7), and "severe anomaly" corresponds to the interval [0.7, 1]. Thus, the confidence level identifier increases as the anomaly confidence score increases, intuitively reflecting the severity of the anomaly.

[0103] Step S154: Determine a corresponding confidence level identifier according to the interval to which the abnormality confidence score belongs.

[0104] Compare the calculated anomaly confidence score with the defined intervals to determine the interval to which the score belongs. Once the interval is determined, the severity of the anomaly can be determined based on the confidence level associated with that interval. For example, if the anomaly confidence score is 0.8, which falls within the interval [0.7, 1], the corresponding confidence level is labeled "Severe Anomaly."

[0105] Step S155: The start time and end time of the abnormality, the spatial coordinate range and the confidence level identifier are integrated to generate a monitoring warning instruction containing spatiotemporal positioning information and severity information.

[0106] The identified anomaly start and end times, spatial coordinate range, and confidence level are integrated. A structured approach can be used to generate monitoring warning instructions. For example, a textual statement might read: "During the period [start time] - [end time], an anomaly of [confidence level] occurred in the [spatial coordinate range] area. Please address it promptly." In a shopping mall monitoring scenario, a monitoring warning instruction might read: "During the period [start time] - [end time], a serious anomaly of [confidence level] occurred in the [spatial coordinate range] area. Please address it promptly." Such monitoring warning instructions provide monitoring personnel with clear spatiotemporal location information and anomaly severity information, enabling them to quickly make decisions and take appropriate measures.

[0107] Step S210: Acquire a historical monitoring calibration data set, where the historical monitoring calibration data set includes a monitoring data stream of a known normal scenario and a monitoring data stream of a known abnormal scenario.

[0108] To accurately calibrate the surveillance recognition model, a historical surveillance calibration dataset is required. In practice, data can be collected from multiple data sources. For surveillance data streams from known normal scenarios, collection can be performed during periods when the surveillance scene is stable and free of anomalies. For example, in a shopping mall surveillance scenario, during normal business hours on weekdays, when personnel flow and equipment operation are normal, network cameras continuously record video frame sequences and corresponding time-series metadata sequences. This data constitutes the surveillance data stream for normal scenarios. For surveillance data streams from known abnormal scenarios, data can be collected by simulating various possible abnormal situations. For example, abnormal events such as fire and theft can be simulated in the mall, and the surveillance data streams during these events can be recorded. During data collection, data integrity and accuracy must be ensured. Preliminary screening and cleaning of the collected data is performed to remove invalid or erroneous data to ensure the quality of the historical surveillance calibration dataset.

[0109] Step S220: performing context encoding processing on the historical monitoring calibration data set to generate a normal scene context feature cube and an abnormal scene context feature cube.

[0110] Context encoding is performed on the historical monitoring calibration dataset, a process similar to that used for real-time monitoring data streams. For normal scene monitoring data streams, spatial features are first extracted from the video frame sequence. Using a convolutional neural network, the video frames are input into a pre-trained basic feature extraction network. The pre-trained network performs multiple convolution operations on the video frames through convolutional layers to extract shallow visual features containing edge contours, texture details, and color distribution information. Next, the shallow visual features are aggregated at multiple scales using a spatial pyramid pooling module to generate local region features with different receptive fields. A global average pooling layer is then used to perform global information compression on the shallow visual features, resulting in global scene features that describe the overall scene layout. The local region features are concatenated with the global scene features to generate a spatial feature vector corresponding to each video frame, which is then arranged into a sequence of spatial feature vectors in the order in which they were acquired.

[0111] At the same time, temporal feature encoding is performed on the temporal metadata sequence. The acquisition timestamp is converted into a time interval sequence, and the difference between adjacent timestamps is calculated. The camera pose information is converted into a spatial coordinate offset sequence, and the pose information difference between adjacent frames is calculated. The ambient light intensity description is converted into a light gradient sequence, and the change in light intensity between adjacent frames is calculated. These three sequences are concatenated into a temporal feature vector sequence.

[0112] Then, using the acquisition timestamp of each frame as the alignment benchmark, the spatial feature vectors and temporal feature vectors at the corresponding time points are concatenated and fused to generate a spatiotemporal fusion feature vector. Finally, the spatiotemporal fusion feature vectors are arranged in chronological order to construct a three-dimensional feature matrix, generating a normal scene context feature cube.

[0113] For the monitoring data stream of abnormal scenarios, the same processing flow is adopted to finally generate the abnormal scenario context feature cube.

[0114] Step S230: performing a normal behavior pattern learning operation using the normal scene context feature cube to generate an initial benchmark feature library.

[0115] Normal behavior pattern learning is performed using the Normal Scene Context Feature Cube. First, the Normal Scene Context Feature Cube is divided into time windows. The appropriate time window length is determined based on the characteristics and requirements of the monitoring scenario. For example, in shopping mall monitoring, the time window can be set to 10 minutes, and sub-cubes with consecutive 10-minute time steps can be selected as learning units.

[0116] Each learning unit is subjected to feature clustering analysis using a density clustering algorithm. This algorithm first defines a radius parameter and a minimum number of points. For each feature point in the learning unit, the number of points within the radius is calculated. If the number of points exceeds the minimum number of points, the point is considered a core point, and the surrounding points are clustered into the same cluster. By continuously expanding the neighborhood of the core point, significantly dense feature clusters are identified. These feature clusters represent stable feature patterns that recur over multiple time periods.

[0117] Perform feature statistics on the feature cluster and calculate the cluster's central vector. Add the dimensions of all feature vectors within the cluster and divide by the number of feature vectors to obtain the central vector, which serves as the typical scene feature template. Simultaneously, calculate the distribution variance of the feature vectors within the cluster to measure the degree of dispersion of the feature vectors relative to the central vector, which serves as an indicator of feature stability.

[0118] Perform correlation analysis on feature clusters in adjacent time windows. Obtain the feature cluster sets for the previous time window and the current time window. Calculate the similarity between each feature cluster in the previous time window and the feature cluster in the current time window using the cosine similarity algorithm. Calculate the dot product of the center vectors of the two clusters and divide it by the product of the module lengths of the two vectors to obtain the similarity value. Based on the similarity calculation results, establish a temporal correlation between the feature clusters. Set a similarity threshold. When the similarity exceeds this threshold, mark the two feature clusters as a temporal continuation of the same feature.

[0119] For feature clusters with temporal continuity, the time-series data of their central vectors are extracted to generate feature evolution trajectories. First-order differences are performed on the feature evolution trajectories to calculate the feature changes between adjacent time steps. Combined with the time interval information, the feature temporal rate of change is obtained. Directional vector calculation is performed on the feature evolution trajectories. By analyzing the changing directions of adjacent feature vectors, the main directional component of the feature changes is extracted and the feature directional offset is generated. The feature temporal rate of change and the feature directional offset are combined to generate a description of the feature evolution pattern.

[0120] Finally, the typical scene feature templates, feature stability indicators and feature evolution law descriptions are integrated and stored to generate the initial benchmark feature library.

[0121] Step S240: matching and comparing the abnormal scene context feature cube with the initial reference feature library, and calculating the feature matching deviation value distribution under the abnormal scene.

[0122] The abnormal scene context feature cube is matched and compared with the initial baseline feature library. The current feature vector sequence is extracted from the abnormal scene context feature cube. This current feature vector sequence contains the spatial feature vector and temporal feature vector of the current time step. The current feature vector sequence is subjected to feature block processing, dividing the feature vectors of consecutive time steps into comparison units of the same size as the learning units of the baseline feature library.

[0123] A benchmark matching operation is performed on each comparison unit, calculating the Euclidean distance between the feature vector within the comparison unit and the corresponding typical scene feature template in the benchmark feature library. Specifically, the feature vector is subtracted from the corresponding dimension of the typical scene feature template, the square root of the sum is taken, and the feature matching deviation value is obtained. A statistical analysis of the feature matching deviation values ​​for all comparison units is performed, plotting a histogram or calculating a probability density function to determine the distribution of feature matching deviation values ​​in abnormal scenarios.

[0124] Step S250: adjusting the parameters of the confidence mapping function according to the difference between the feature matching deviation value distribution in the abnormal scene and the deviation value distribution in the normal scene, so that the abnormal confidence score of the abnormal scene is higher than the abnormal confidence score of the normal scene.

[0125] Step S251: Analyze the comprehensive deviation indicator set corresponding to the normal scene to determine the deviation value distribution range of the normal scene.

[0126] Analyze the set of comprehensive deviation indicators corresponding to normal scenarios. The comprehensive deviation indicators are obtained after time-weighted processing and statistical analysis of the feature matching deviation values ​​of normal scenarios. First, review the calculation process of the feature matching deviation values ​​in normal scenarios, perform time-weighted processing on the deviation values, and assign enhanced weights to the deviation values ​​of the recent time steps. Determine the length parameter of the time-weighted window to make it consistent with the time step of the baseline feature library learning unit. Based on the time distance between the time step and the current time point, generate an exponentially decaying weight coefficient. Multiply the feature matching deviation value with the corresponding weight coefficient to obtain the time-weighted deviation value, and arrange it into a weighted deviation value sequence in chronological order.

[0127] Perform statistical analysis on the weighted deviation value sequence, calculate the mean, maximum, and variance of the deviation values, and generate a comprehensive deviation index. Collect comprehensive deviation indicators from a large number of normal scenarios, plot a histogram, or fit a probability distribution curve to determine the distribution range of deviation values ​​in normal scenarios. For example, determine the distribution range by calculating the mean plus or minus a set multiple of the standard deviation.

[0128] Step S252: Analyze the feature matching deviation value distribution in the abnormal scene to determine the deviation value distribution range of the abnormal scene.

[0129] Analyze the distribution of feature matching deviation values ​​in abnormal scenarios. Using the method previously described for calculating feature matching deviation values ​​for abnormal scenarios, obtain a large number of feature matching deviation values. Similarly, perform time-weighted processing and statistical analysis on these deviation values ​​to generate a comprehensive deviation index for abnormal scenarios. Draw a histogram of the comprehensive deviation index for abnormal scenarios or fit a probability distribution curve to determine the distribution range of deviation values ​​for abnormal scenarios. Similar to the analysis process for normal scenarios, calculate statistics such as the mean and standard deviation to determine the distribution range.

[0130] Step S253: adjusting the parameters of the confidence mapping function according to the degree of overlap between the deviation value distribution range of the normal scene and the deviation value distribution range of the abnormal scene.

[0131] Compare the degree of overlap between the deviation value distribution ranges for normal and abnormal scenarios. If there is a large overlap, it indicates that the current confidence mapping function cannot effectively distinguish between normal and abnormal scenarios, and the function parameters need to be adjusted. The confidence mapping function is typically a nonlinear function, and its parameters determine the mapping relationship between the comprehensive deviation index and the anomaly confidence score. You can adjust the parameters iteratively, recalculating the anomaly confidence scores for normal and abnormal scenarios after each adjustment and observing changes in the degree of overlap. For example, increasing the slope of the function in areas with large deviation values ​​can result in a higher anomaly confidence score corresponding to the comprehensive deviation index of abnormal scenarios.

[0132] Step S254: Through an iterative verification process, the parameters are gradually optimized until the abnormality confidence scores of normal scenarios are concentrated in the first interval, and the abnormality confidence scores of abnormal scenarios are concentrated in the second interval, wherein the highest score value in the first interval is less than the lowest score value in the second interval, and the difference between the highest score value and the lowest score value is greater than the set difference.

[0133] Perform an iterative verification process to continuously optimize the parameters of the confidence mapping function. After each parameter adjustment, use the historical monitoring calibration data set for verification. Input the comprehensive deviation indicators of normal scenarios and abnormal scenarios into the adjusted confidence mapping function respectively to obtain the corresponding abnormal confidence scores. Statistically analyze the distribution of abnormal confidence scores for normal scenarios and abnormal scenarios to observe whether the abnormal confidence scores of normal scenarios are concentrated in the first interval, the abnormal confidence scores of abnormal scenarios are concentrated in the second interval, and there is sufficient spacing between the two intervals. Set the difference value based on actual application requirements, for example, it can be set to 0.2. If the conditions are not met, continue to adjust the parameters and repeat the verification process until satisfactory results are achieved.

[0134] Step S255: outputting the adjusted confidence mapping function parameters.

[0135] After multiple iterations of verification, if the anomaly confidence scores for normal scenarios are concentrated in the first interval, the anomaly confidence scores for abnormal scenarios are concentrated in the second interval, and the difference between the highest and lowest scores is greater than the set difference, the adjusted confidence mapping function parameters are output. These parameters are used in subsequent real-time monitoring data stream processing, enabling the monitoring recognition model to more accurately identify anomalies.

[0136] Step S260: combining the adjusted confidence mapping function with the initial reference feature library to generate a calibrated monitoring recognition model.

[0137] The adjusted confidence mapping function is combined with the initial baseline feature library to form a calibrated monitoring recognition model. In subsequent real-time monitoring, when the real-time monitoring data stream is acquired, the context feature cube is generated according to the previous process. This is then matched and compared with the initial baseline feature library to calculate a comprehensive deviation index. This comprehensive deviation index is then input into the adjusted confidence mapping function to obtain a more accurate anomaly confidence score. The calibrated monitoring recognition model combines the characteristic information of normal scenarios with the optimized confidence mapping relationship, enabling more effective identification of anomalies.

[0138] Step S270: Process the real-time monitoring data stream using the calibrated monitoring recognition model to generate anomaly confidence scores.

[0139] The calibrated surveillance recognition model is used to process the real-time surveillance data stream. First, the real-time surveillance data stream is acquired, including a sequence of video frames and a sequence of temporal metadata. The real-time surveillance data stream is subjected to spatiotemporal context encoding to generate a real-time context feature cube. The real-time context feature cube is then dynamically matched and compared with the initial baseline feature library. The current feature vector sequence is extracted from the real-time context feature cube and partitioned into comparison units of the same size as the baseline feature library learning units. The Euclidean distance between the feature vectors within the comparison unit and the corresponding typical scene feature template in the baseline feature library is calculated to obtain the feature matching deviation value. The feature matching deviation value is time-weighted to generate a weighted deviation value sequence, which is then statistically analyzed to obtain a comprehensive deviation index.

[0140] Finally, the comprehensive deviation index is input into the adjusted confidence mapping function, and an anomaly confidence score is generated according to the mapping relationship of the function. The anomaly confidence score can more accurately reflect the possibility of anomalies occurring in real-time monitoring scenarios.

[0141] Throughout the data processing process, various privacy protection and anti-leakage technologies are employed when collecting privacy-sensitive data. The collection of monitoring data streams must be carried out under the conditions of ensuring compliance with all legal requirements and obtaining authorized permissions. Encryption technology is employed to store monitoring data streams in encrypted form on the server, accessible only to authorized users. During data transmission, secure transport protocols such as SSL / TLS are used to encrypt data and prevent theft or tampering during transmission. Furthermore, strict identity authentication and permission management are implemented for users accessing the data, ensuring that only users with appropriate authorization can access and process privacy-sensitive data.

[0142] In the above embodiments, a convolutional neural network is used as a key model for spatial feature extraction, which has multiple convolutional layers, pooling layers and fully connected layers. The convolutional layer is used to extract the features of the video frame, the pooling layer is used to reduce the dimension of the features, and the fully connected layer is used to integrate and classify the features. During the training process, a large amount of image data is used for training, and the parameters of the model are continuously adjusted through the back propagation algorithm so that the model can accurately extract the local area features and global scene features of the video frame. As for the density clustering algorithm, it is an unsupervised learning algorithm that divides high-density areas into different clusters by calculating the density between data points. During the training process, there is no need to label the data, and the density clustering algorithm will automatically identify the feature clusters in the data.

[0143] Figure 2The following diagram illustrates exemplary hardware and software components of an AI-based network camera surveillance and recognition system 100 that can implement the concepts of the present application, as provided in some embodiments of the present application. For example, the processor 120 can be used in the AI-based network camera surveillance and recognition system 100 to perform the functions described in the present application.

[0144] The AI-based network camera surveillance recognition system 100 can be a general-purpose server or a special-purpose server, both of which can be used to implement the AI-based network camera surveillance recognition method of this application. Although only one server is shown in this application, for convenience, the functions described in this application can be implemented in a distributed manner on multiple similar platforms to balance the processing load.

[0145] For example, the AI-based network camera surveillance and recognition system 100 may include a network port 110 connected to a network, one or more processors 120 for executing program instructions, a communication bus 130, and storage media 140 in various forms, such as a disk, ROM, or RAM, or any combination thereof. For example, the AI-based network camera surveillance and recognition system 100 may also include program instructions stored in ROM, RAM, or other types of non-transitory storage media, or any combination thereof. The method of the present application may be implemented based on these program instructions. The AI-based network camera surveillance and recognition system 100 also includes an I / O interface 150 between the computer and other input and output devices.

[0146] For ease of explanation, only one processor is described in the artificial intelligence-based network camera surveillance and recognition system 100. However, it should be noted that the artificial intelligence-based network camera surveillance and recognition system 100 in this application may also include multiple processors, and therefore the steps performed by one processor described in this application may also be performed jointly or individually by multiple processors. For example, if the processor of the artificial intelligence-based network camera surveillance and recognition system 100 performs steps A and B, it should be understood that steps A and B may also be performed jointly by two different processors or individually in one processor. For example, the first processor performs step A and the second processor performs step B, or the first processor and the second processor perform steps A and B together.

[0147] In addition, an embodiment of the present invention further provides a readable storage medium, in which computer-executable instructions are preset. When a processor executes the computer-executable instructions, the above-mentioned artificial intelligence-based network camera monitoring and recognition method is implemented.

[0148] It should be noted that in order to simplify the description of the present invention and thus help understand one or more embodiments of the invention, in the foregoing description of the embodiments of the present invention, multiple features are sometimes combined into one embodiment, figure or description thereof.

Claims

1. A network camera monitoring and recognition method based on artificial intelligence, characterized in that: The method comprises: Obtain a monitoring data stream output by a network camera, wherein the monitoring data stream includes a continuously acquired video frame sequence and a corresponding time-series metadata sequence, wherein the time-series metadata sequence includes an acquisition timestamp of each frame of video, camera posture information, and a description of the ambient light intensity; Performing spatiotemporal context coding on the monitoring data stream, associating and mapping the spatial features of the video frame sequence with the temporal features of the temporal metadata sequence, and generating a context feature cube containing spatial position information and temporal evolution information; Performing normal behavior pattern learning operations based on the context feature cube, extracting feature patterns that appear stably over multiple time periods through an unsupervised learning algorithm, and generating a baseline feature library containing typical scene feature templates and feature evolution law descriptions; Dynamically matching and comparing the context feature cube of the current period with the benchmark feature library, calculating the feature matching deviation value and generating an anomaly confidence score; A monitoring and warning instruction is generated according to the abnormality confidence score and the corresponding spatiotemporal location information, wherein the monitoring and warning instruction includes the abnormality occurrence time, spatial coordinate range and confidence level identifier.

2. The network camera monitoring and recognition method based on artificial intelligence according to claim 1 is characterized in that: The spatiotemporal context coding processing is performed on the monitoring data stream, and the spatial features of the video frame sequence are associated and mapped with the temporal features of the time series metadata sequence to generate a context feature cube containing spatial position information and temporal evolution information, including: Performing spatial feature extraction processing on the video frame sequence, using a convolutional neural network to extract local region features and global scene features of each frame of video, and generating a spatial feature vector sequence; Performing time feature encoding processing on the time series metadata sequence, converting the acquisition timestamp into a time interval sequence, converting the camera pose information into a spatial coordinate offset sequence, and converting the ambient light intensity description into a light change gradient sequence, thereby generating a time feature vector sequence; Establishing a time alignment relationship between the spatial feature vector sequence and the temporal feature vector sequence, taking the acquisition timestamp of each frame of video as the alignment reference, splicing and fusing the spatial feature vectors and temporal feature vectors of corresponding time points to generate a spatiotemporal fusion feature vector; The spatiotemporal fusion feature vectors are arranged in chronological order to construct a three-dimensional feature matrix, wherein the first dimension is the time step, the second dimension is the spatial position coordinate, and the third dimension is the feature dimension, thereby generating a context feature cube containing spatial position information and time evolution information.

3. The network camera monitoring and recognition method based on artificial intelligence according to claim 2 is characterized in that: The spatial feature extraction process is performed on the video frame sequence, and a convolutional neural network is used to extract local area features and global scene features of each frame of video to generate a spatial feature vector sequence, including: Inputting the video frame sequence into a pre-trained basic feature extraction network, and extracting shallow visual features through a convolutional layer, wherein the shallow visual features include edge contours, texture details and color distribution information; Performing multi-scale aggregation processing on the shallow visual features through a spatial pyramid pooling module to generate local area features with different receptive fields; Performing global information compression processing on the shallow visual features through a global average pooling layer to generate global scene features that describe the overall scene layout; Concatenate the local region features with the global scene features to generate a spatial feature vector corresponding to each video frame; The spatial feature vectors are arranged in the order of acquisition of the video frames to generate a spatial feature vector sequence.

4. The network camera monitoring and recognition method based on artificial intelligence according to claim 1 is characterized in that: The normal behavior pattern learning operation is performed based on the context feature cube, and feature patterns that appear stably in multiple time periods are extracted through an unsupervised learning algorithm to generate a baseline feature library. The baseline feature library contains typical scene feature templates and feature evolution law descriptions, including: Performing time window division processing on the context feature cube, and selecting sub-cubes of consecutive time steps as learning units; Perform feature clustering analysis on each learning unit and use density clustering algorithm to identify significant dense feature clusters, which represent stable feature patterns that recur over multiple time periods. Performing feature statistical processing on the feature cluster, calculating the central vector of the feature cluster as a typical scene feature template, and calculating the distribution variance of the feature vectors within the feature cluster as a feature stability indicator; Perform correlation analysis on feature clusters in adjacent time windows, track the evolution trajectory of the same feature cluster in the time dimension, calculate the time change rate and directional offset of the feature vector, and generate a description of the feature evolution law; The typical scene feature templates, feature stability indicators and feature evolution law descriptions are integrated and stored to generate a benchmark feature library.

5. The method for network camera monitoring and recognition based on artificial intelligence according to claim 4 is characterized in that: The process of performing correlation analysis on feature clusters in adjacent time windows, tracking the evolution trajectory of the same feature cluster in the time dimension, calculating the time change rate and directional offset of the feature vector, and generating a description of the feature evolution law includes: Get the feature cluster set of the previous time window and the feature cluster set of the current time window; Perform similarity calculation on each feature cluster of the previous time window and the feature cluster of the current time window, and use the cosine similarity algorithm to calculate the similarity of the cluster center vector; Establish the temporal correlation relationship of feature clusters based on the similarity calculation results, and mark the feature clusters with similarity greater than a preset threshold as the temporal continuation of the same feature; For feature clusters with time-continuing relationships, the sequence data of the central vector of the feature cluster in the time dimension is extracted to generate the feature evolution trajectory; Performing first-order difference calculation on the feature evolution trajectory to obtain the feature change amount of adjacent time steps, and generating the feature time change rate in combination with the time interval information; Performing direction vector calculation processing on the feature evolution trajectory, extracting the main direction component of the feature change, and generating a feature direction offset; The feature time change rate and the feature direction offset are integrated to generate a feature evolution law description that describes the feature evolution law over time.

6. The network camera monitoring and recognition method based on artificial intelligence according to claim 1 is characterized in that: The dynamically matching and comparing the context feature cube of the current period with the reference feature library, calculating the feature matching deviation value and generating anomaly confidence score includes: Extracting a current feature vector sequence from the context feature cube of the current period, wherein the current feature vector sequence includes a spatial feature vector and a temporal feature vector of the current time step; Performing feature block processing on the current feature vector sequence, dividing the feature vectors of consecutive time steps into comparison units of the same size as the benchmark feature library learning units; Perform a benchmark matching operation on each comparison unit, calculate the Euclidean distance between the feature vector in the comparison unit and the corresponding typical scene feature template in the benchmark feature library, and generate a feature matching deviation value; Performing time-weighted processing on the feature matching deviation values, assigning enhanced weights to the deviation values ​​of recent time steps, and generating a weighted deviation value sequence; Performing statistical analysis on the weighted deviation value sequence, calculating the mean, maximum value and variance of the deviation values, and generating a comprehensive deviation index; The comprehensive deviation index is input into a preset confidence mapping function, and an abnormality confidence score is generated based on the positive correlation between the deviation value and the abnormality possibility.

7. The method for network camera monitoring and identification based on artificial intelligence according to claim 6, characterized in that: The step of performing time-weighted processing on the feature matching deviation values, assigning enhanced weights to the deviation values ​​of recent time steps, and generating a weighted deviation value sequence includes: Determining a length parameter of the time-weighted window, wherein the length parameter is consistent with a time step of a reference feature library learning unit; Generate an exponential decay weight coefficient based on the time distance between the time step and the current time point, wherein the exponential decay weight coefficient shows an exponentially decreasing trend as the time distance increases; Multiplying the feature matching deviation value by the exponential decay weight coefficient of the corresponding time step to generate a time-weighted deviation value; Arranging the time-weighted deviation values ​​in chronological order to generate a weighted deviation value sequence; The integrated deviation index is input into a preset confidence mapping function, and an abnormality confidence score is generated according to the positive correlation between the deviation value and the abnormality possibility, including: Obtain the comprehensive deviation indicator distribution data under historical normal scenarios, and calculate the deviation mean and standard deviation under normal scenarios; Taking the deviation mean as a benchmark, the difference between the current comprehensive deviation index and the deviation mean is divided by the deviation standard deviation to generate a standardized deviation value; The standardized deviation value is converted into a confidence score in the interval [0, 1] by a nonlinear mapping method, wherein the larger the standardized deviation value, the higher the anomaly confidence score.

8. The method for network camera monitoring and identification based on artificial intelligence according to claim 6, characterized in that: The method comprises: Acquire a historical monitoring calibration data set, where the historical monitoring calibration data set includes a monitoring data stream of a known normal scenario and a monitoring data stream of a known abnormal scenario; Performing context encoding processing on the historical monitoring calibration data set to generate a normal scene context feature cube and an abnormal scene context feature cube; Performing a normal behavior pattern learning operation using the normal scene context feature cube to generate an initial benchmark feature library; Matching and comparing the abnormal scene context feature cube with the initial reference feature library, and calculating the feature matching deviation value distribution under the abnormal scene; Adjusting the parameters of the confidence mapping function according to the difference between the feature matching deviation value distribution in the abnormal scenario and the deviation value distribution in the normal scenario so that the abnormal confidence score of the abnormal scenario is higher than the abnormal confidence score of the normal scenario; Combining the adjusted confidence mapping function with the initial benchmark feature library to generate a calibrated monitoring recognition model; Processing the real-time monitoring data stream using the calibrated monitoring recognition model to generate anomaly confidence scores; The step of adjusting the parameters of the confidence mapping function based on the difference between the feature matching deviation value distribution in the abnormal scenario and the deviation value distribution in the normal scenario so that the abnormal confidence score of the abnormal scenario is higher than the abnormal confidence score of the normal scenario includes: Analyze the comprehensive deviation indicator set corresponding to the normal scene to determine the deviation value distribution range of the normal scene; Analyze the distribution of feature matching deviation values ​​in the abnormal scenario to determine the deviation value distribution range of the abnormal scenario; Adjusting the parameters of the confidence mapping function according to the degree of overlap between the deviation value distribution range of the normal scene and the deviation value distribution range of the abnormal scene; Through an iterative verification process, the parameters are gradually optimized until the abnormality confidence scores of normal scenarios are concentrated in a first interval, and the abnormality confidence scores of abnormal scenarios are concentrated in a second interval, wherein the highest score value in the first interval is less than the lowest score value in the second interval, and the difference between the highest score value and the lowest score value is greater than a set difference; Outputs the adjusted confidence mapping function parameters.

9. The network camera monitoring and recognition method based on artificial intelligence according to claim 1 is characterized in that: The generating of a monitoring and warning instruction according to the abnormality confidence score and the corresponding spatiotemporal location information, wherein the monitoring and warning instruction includes the abnormality occurrence time, spatial coordinate range and confidence level identifier, includes: Extracting timestamp information corresponding to the anomaly confidence score from the context feature cube to determine the start time and end time of the anomaly; Based on the spatial position information of the context feature cube, locate the spatial area where the feature matching deviation value exceeds a preset threshold, and generate the spatial coordinate range where the anomaly occurs; Dividing the anomaly confidence score into a plurality of continuous intervals, each interval corresponding to a confidence level identifier, and the confidence level identifier increases as the confidence score increases; Determining a corresponding confidence level identifier based on the interval to which the abnormality confidence score belongs; The start time and end time of the abnormality, the spatial coordinate range and the confidence level identifier are integrated to generate a monitoring and early warning instruction containing spatiotemporal positioning information and severity information.

10. An artificial intelligence-based network camera monitoring and recognition system, characterized in that: It includes a processor and a memory, the memory is connected to the processor, the memory is used to store programs, instructions or codes, and the processor is used to execute the programs, instructions or codes in the memory to implement the artificial intelligence-based network camera monitoring and recognition method described in any one of claims 1 to 9.

Citation Information

Patent Citations

  • Video human action reorganization method based on sparse subspace clustering

    CN104732208A

  • Character action recognition analysis method and system based on infrared laser and deep learning

    CN118747911A

  • Image processing method and system for intelligent security and protection monitoring

    CN118887622A

  • Early warning analysis method based on intelligent vision and server

    CN119810757A

  • Method and system for semantically segmenting scenes of a video sequence

    GB0406512D0

Cited By

  • Industrial computer data processing method and device

    CN120875483A

  • Beidou video multi-behavior analysis early warning system and terminal

    CN120954208A

  • Abnormality identification method for intelligent monitoring of power distribution room

    CN121686349A

  • An abnormality identification method for intelligent monitoring of a power distribution room

    CN121686349B

  • Network camera adding method and device and storage medium

    CN121691769A