Image analysis early warning method and system based on visual large model

By preprocessing image data with timestamps and calling a large visual model for analysis, the problems of low efficiency of manual observation and insufficient adaptability of simple algorithms in existing technologies are solved, and efficient and accurate image analysis and early warning are achieved.

CN120655982AActive Publication Date: 2025-09-16GUANGXI CHUANGXUAN TECH CO LTD

Patent Information

Application Number
CN202510775166.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-11
Publication Date
2025-09-16
Estimated Expiration
2045-06-11

AI Technical Summary

Technical Problem

Existing image analysis and early warning methods rely on manual observation, which is inefficient and easily affected by human factors. Detection systems based on simple image processing algorithms have insufficient detection capabilities in complex and changing scenes, and it is difficult to accurately identify abnormal areas and types.

Method used

By obtaining a set of original image data with timestamps, preprocessing operations are performed to enhance the characteristics of the target area and stabilize the background environment characteristics, and a pre-trained visual large model is called to perform image analysis, generate abnormal area positioning information and type identification, and generate scene warning instructions.

Benefits of technology

It achieves timely discovery, precise positioning and rapid response to abnormal events, and improves the efficiency and accuracy of image analysis and early warning.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120655982A_ABST
    Figure CN120655982A_ABST
Patent Text Reader

Abstract

The invention provides an image analysis early warning method and system based on a visual large model, and the method comprises the steps: firstly obtaining an original image data set of a to-be-analyzed scene comprising a plurality of image units with timestamps, then carrying out the preprocessing of the original image data set, and obtaining a preprocessed image set with the features of an enhanced target region and the features of a stable background environment; calling a pre-trained visual large model to carry out image analysis processing on the preprocessed image set, generating an image analysis result containing abnormal region positioning information and an abnormal type identifier, and determining abnormal event attributes and time-space distribution feature information according to the image analysis result; and finally, generating a scene early warning instruction containing event positioning coordinates based on the abnormal event attributes and the time-space distribution feature information, and sending the scene early warning instruction to the target early warning terminal to trigger a response operation, thereby effectively improving the accuracy and response speed of image analysis early warning.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer vision technology, and in particular to an image analysis and early warning method and system based on a large visual model. Background Art

[0002] Real-time, accurate image analysis and early warning are crucial in numerous application scenarios, such as security surveillance, industrial production monitoring, and natural disaster warning. Currently, traditional image analysis and early warning methods rely primarily on manual observation and detection systems based on simple image processing algorithms. Manual observation is not only inefficient but also susceptible to human factors such as fatigue and inattention, leading to missed and false detections of abnormal events. Furthermore, manually processing large amounts of image data consumes significant time and labor costs.

[0003] While detection systems based on simple image processing algorithms achieve a degree of automation, these algorithms are typically limited in functionality and can only process specific types of image features. Their detection capabilities and adaptability are severely limited for complex and changing scenarios and diverse abnormal events. For example, in situations such as lighting variations and object occlusion, traditional algorithms often struggle to accurately identify abnormal areas and types of anomalies. Therefore, there is an urgent need for more intelligent, efficient, and adaptable image analysis and early warning methods to improve the accuracy of image analysis and the timeliness of early warnings. Summary of the Invention

[0004] In view of the above-mentioned problems, in combination with the first aspect of the present invention, an embodiment of the present invention provides an image analysis and early warning method based on a large visual model, the method comprising: Acquire an original image data set of a scene to be analyzed, wherein the original image data set includes a plurality of image units that are continuously acquired and marked with time stamps; Performing a preprocessing operation on the original image data set to obtain a preprocessed image set, wherein the preprocessed image set includes enhanced target area features and stable background environment features; Calling a pre-trained visual large model to perform image analysis processing on the pre-processed image set to generate an image analysis result including abnormal area location information and abnormal type identification; Determine, based on the image analysis results, the attributes of the abnormal events present in the scene to be analyzed and the spatiotemporal distribution feature information of the abnormal events in the image; A scene warning instruction including event location coordinates is generated based on the abnormal event attributes and the spatiotemporal distribution feature information, and the scene warning instruction is sent to a target warning terminal to trigger a response operation.

[0005] On the other hand, an embodiment of the present invention also provides an image analysis and early warning system based on a large visual model, including a processor and a machine-readable storage medium, wherein the machine-readable storage medium is connected to the processor, the machine-readable storage medium is used to store programs, instructions or codes, and the processor is used to execute the programs, instructions or codes in the machine-readable storage medium to implement the above method.

[0006] Based on the above aspects, the embodiment of the present invention obtains a set of original image data that is continuously collected and marked with timestamps for the scene to be analyzed, performs preprocessing operations on the original image data set, obtains enhanced target area features and stable background environment features, effectively improves image quality, highlights key information, and calls a pre-trained visual large model to perform image analysis processing on the preprocessed image set, which can give full play to the powerful feature extraction and pattern recognition capabilities of the visual large model, generate accurate image analysis results containing abnormal area positioning information and abnormal type identification, determine the abnormal event attributes and spatiotemporal distribution feature information in the image picture according to the image analysis results, deeply explore the nature and development laws of abnormal events, generate scene warning instructions containing event positioning coordinates based on this information, and send them to the target warning terminal to trigger a response operation, thereby realizing timely discovery, precise positioning and rapid response of abnormal events, and greatly improving the efficiency and accuracy of image analysis warnings. BRIEF DESCRIPTION OF THE DRAWINGS

[0007] Figure 1 It is a schematic diagram of the execution flow of the image analysis and early warning method based on the visual large model provided by an embodiment of the present invention.

[0008] Figure 2 Schematic diagram of exemplary hardware and software components of an image analysis and early warning system based on a large visual model provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0009] The present invention will be described in detail below with reference to the accompanying drawings. Figure 1 This is a flow chart of an image analysis and early warning method based on a large visual model provided by an embodiment of the present invention. The image analysis and early warning method based on a large visual model is introduced in detail below.

[0010] Step S110: obtaining an original image data set of a scene to be analyzed, wherein the original image data set includes a plurality of image units that are continuously acquired and marked with time stamps.

[0011] In this embodiment, the scene to be analyzed can be an urban security scenario, such as a busy city street. To obtain a set of raw image data for this scene, multiple high-definition surveillance cameras can be installed at key locations on the street, such as intersections, shopping mall entrances, and bus stops. These cameras are capable of continuously capturing images, and each captured image is automatically timestamped.

[0012] Specifically, the camera can capture images at a pre-set acquisition frequency. Assuming the acquisition frequency is a certain number of images per second, when acquisition begins at a certain time, the first image is timestamped to the millisecond of that moment. Each subsequent image is timestamped accordingly. For example, if acquisition begins at 10:00 AM, the first image will be timestamped at 10:00:00:00 milliseconds, and the second image immediately after will be timestamped at 10:0 ...

[0013] Over time, these cameras continuously capture street images, ultimately forming a raw image data set consisting of a large number of continuously captured, time-stamped image units. The image units in this raw image data set can fully record various scene information on city streets at different times, such as pedestrian movement, vehicle traffic, and the status of street furniture.

[0014] Step S120: performing a preprocessing operation on the original image data set to obtain a preprocessed image set, wherein the preprocessed image set includes enhanced target region features and stable background environment features.

[0015] In this embodiment, during the acquisition process, the raw image data set may contain noise interference and uneven brightness due to environmental factors (such as lighting changes and dust) and the camera's inherent performance, which can adversely affect subsequent image analysis. Therefore, it is necessary to preprocess the raw image data set to obtain a preprocessed image set more suitable for analysis.

[0016] In urban security scenarios, the target area may be pedestrians, vehicles, or suspicious objects on the street, while the background environment includes the street pavement, buildings, and streetlights. The purpose of preprocessing is to highlight the characteristics of the target area while stabilizing the background environment. For example, by enhancing pedestrian facial features and license plate features, subsequent analysis can more accurately identify pedestrians and vehicles. A more stable street background can reduce interference from environmental factors in target area analysis.

[0017] Step S121: performing noise suppression processing on the original image data set, reducing random noise interference in the image unit through an adaptive filtering algorithm, and obtaining a noise-suppressed image unit.

[0018] In this embodiment, during the image acquisition process of urban streets, random noise interference is inevitably present in the raw image units due to factors such as camera sensor noise and unstable external light. This noise blurs the image, affecting subsequent recognition and analysis of the target area. To reduce the impact of this noise, an adaptive filtering algorithm is used for noise suppression.

[0019] The core idea of ​​the adaptive filtering algorithm is to dynamically adjust the filtering parameters according to the noise conditions in different areas of the image. The specific implementation process is as follows: Step S1211: Calculate statistical features of local neighborhood pixel values ​​of each pixel point in a single image unit, where the statistical features include the mean and variance of the neighborhood pixel values.

[0020] For each pixel in a single image unit, a local neighborhood can be selected around it. The size of this local neighborhood can be selected based on the actual situation, with common sizes being 3×3, 5×5, or 7×7. Taking a 5×5 local neighborhood as an example, for a pixel in an image, the pixels in the 5 rows and 5 columns surrounding it can be considered as a local neighborhood.

[0021] Next, the mean and variance of all pixel values ​​within the local neighborhood are calculated. The mean is calculated by adding the values ​​of all pixels within the local neighborhood and dividing by the total number of pixels. For example, in a 5×5 local neighborhood, there are 25 pixels. Adding the values ​​of these 25 pixels and dividing by 25 gives the mean of the local neighborhood. The variance is calculated by first taking the square of the difference between each pixel value and the mean, adding these squared values, and finally dividing by the total number of pixels. By calculating the mean and variance, we can understand the distribution of pixel values ​​within the local neighborhood. The mean reflects the average level of pixel values ​​within the local neighborhood, while the variance reflects the degree of dispersion of the pixel values.

[0022] Step S1212: Determine the noise intensity estimation value of each pixel point according to the mean and variance of the local neighborhood pixel values.

[0023] After obtaining the mean and variance of the local neighborhood pixel values ​​for each pixel, the noise intensity estimate for that pixel can be determined according to the set rules. Generally speaking, the larger the variance, the greater the degree of dispersion of the pixel values ​​in the local neighborhood, which means that the possibility of noise is greater, and the corresponding noise intensity estimate is higher. For example, the variance can be mapped to a range of noise intensity estimates through a preset functional relationship. Assuming that the variance ranges from 0 to a certain maximum value, the variance value is converted into a noise intensity estimate through a linear or nonlinear function. The noise intensity estimate can range from 0 to 1, and the closer the value is to 1, the greater the noise intensity.

[0024] Step S1213: Dynamically adjust the size of the filter window and the filter weight based on the noise intensity estimation value, wherein the area with a larger noise intensity estimation value adopts a larger filter window and a smaller center pixel weight, and the area with a smaller noise intensity estimation value adopts a smaller filter window and a larger center pixel weight.

[0025] The filter window size and filter weight are dynamically adjusted based on the noise intensity estimate for each pixel. Regions with larger noise intensity estimates indicate severe noise, requiring a larger filter window to cover more pixels for more effective noise removal. Furthermore, to reduce the influence of central pixels on the filtering results, the weight of central pixels can be reduced. For example, in a region with a large noise intensity estimate, the filter window can be increased from 3×3 to 5×5 or 7×7, and the weight of central pixels can be reduced from a higher value to a lower value.

[0026] Conversely, for areas with a small noise intensity estimate, indicating relatively low noise levels, a smaller filter window can be used to meet the filtering requirements, while increasing the weight of the center pixel to preserve the original information of that pixel. For example, in an area with a small noise intensity estimate, the filter window can be kept at 3×3 and the weight of the center pixel can be increased.

[0027] Step S1214: performing weighted average calculation on the neighborhood pixel values ​​of each pixel point to generate a filtered pixel value, wherein the weighted average calculation uses the dynamically adjusted filtering window and filtering weight.

[0028] After determining the dynamic filtering window and filtering weight for each pixel, a weighted average is calculated for the pixel's neighborhood. Specifically, for each pixel in its neighborhood, its pixel value is multiplied by the corresponding filtering weight, and all products are added together. The resulting sum is the filtered pixel value. For example, in a 5×5 filtering window, there are 25 pixels, each with a corresponding filtering weight. Each pixel's pixel value is multiplied by its corresponding filtering weight, and these 25 products are added together to obtain the filtered pixel value for that pixel.

[0029] Step S1215: combining the filtered pixel values ​​of all pixels into a single noise-suppressed image unit, performing the above processing on each image unit in the original image data set, and obtaining a noise-suppressed image unit set.

[0030] By combining the filtered pixel values ​​of all pixels in a single image unit according to their original positions, a noise-suppressed image unit is formed. The noise suppression process described above is then performed on each image unit in the original image data set, ultimately resulting in a noise-suppressed image unit set. The image unit noise in this set is effectively reduced, resulting in a clearer image.

[0031] Step S122: performing brightness equalization processing on the noise-suppressed image units, adjusting the brightness distribution of the image units based on global histogram statistics, so that the brightness mean values ​​of different image units remain consistent, and obtaining brightness-balanced image units.

[0032] In this embodiment, during the image capture process of urban streets, the brightness of the captured image units may vary significantly due to varying lighting conditions at different times, such as daytime and nighttime, sunny and cloudy days, etc. This brightness difference can affect the subsequent recognition and analysis of target areas in the image. Therefore, it is necessary to perform brightness equalization processing on the image units after noise suppression.

[0033] The brightness equalization method is based on global histogram statistics. First, a global histogram is calculated for each noise-suppressed image unit. The global histogram shows the distribution of the number of pixels at each brightness level in the image. For example, in an 8-bit image, where the brightness levels range from 0 to 255, the global histogram counts the number of pixels at each brightness level.

[0034] Then, based on the statistical results of the global histogram, the brightness distribution of the image cells is adjusted. Specifically, a mapping function is used to map each pixel value in the original image to a new pixel value, making the brightness distribution of the adjusted image more uniform. For example, histogram equalization can be used to stretch the histogram of the original image to make the number of pixels at each brightness level more uniform.

[0035] Through this brightness equalization process, the average brightness of different image units remains consistent. For example, in images of city streets, whether captured during the day or at night, after brightness equalization, the average brightness will be close to a preset standard value. This eliminates the impact of lighting conditions on image analysis, making subsequent analysis more accurate. The result is a brightness-balanced image unit.

[0036] Step S123: performing target region enhancement processing on the brightness-equalized image unit, identifying the target region boundary in the image unit through an edge detection algorithm, and performing nonlinear enhancement on the pixel value of the target region based on the boundary information to obtain an image unit with enhanced target region features.

[0037] In this embodiment, in urban security scenarios, accurately identifying the features of target areas (such as pedestrians and vehicles) is crucial. While image units that have undergone brightness equalization may have more uniform brightness, the features of the target areas may still be less prominent. Therefore, target area enhancement processing is required for these brightness-equalized image units.

[0038] First, an edge detection algorithm is used to identify the boundaries of the target area within the image unit. Edge detection algorithms can detect areas of image brightness with dramatic changes, which often represent the boundaries of the target area. Common edge detection algorithms include the Canny edge detection algorithm and the Sobel edge detection algorithm. For example, the Canny edge detection algorithm uses Gaussian smoothing, gradient calculation, non-maximum suppression, and double thresholding to ultimately obtain edge information within the image.

[0039] In city street images, edge detection algorithms can be used to identify the boundaries of target areas such as pedestrians and vehicles. For example, the body outline of a pedestrian and the outline of a vehicle can be detected.

[0040] Then, based on the detected boundary information of the target region, the pixel values ​​in the target region are nonlinearly enhanced. Nonlinear enhancement can be performed using a variety of methods, such as contrast stretching and grayscale transformation. For example, a nonlinear grayscale transformation function can be used to stretch the pixel value range within the target region, making the features of the target region more distinct. For example, the facial features of pedestrians and license plates can be made more prominent in the image using these nonlinear enhancement methods.

[0041] Finally, after the target region enhancement process, an image unit with enhanced target region features is obtained. In this image unit, the features of the target region are significantly enhanced, facilitating more accurate recognition and analysis of the target region in the future.

[0042] Step S124: performing background stabilization processing on the image unit with enhanced target area features, extracting background area pixel values ​​of continuous image units, eliminating random fluctuations of the background area through a time series smoothing algorithm, and obtaining image units with stable background environment features.

[0043] In this embodiment, in a sequence of city street images, the background environment (e.g., street pavement, buildings, etc.) may experience random fluctuations due to random factors (e.g., wind blowing leaves, temporary occlusion by vehicles, etc.). These random fluctuations can affect the analysis of the target area, necessitating background stabilization processing for the image units whose target area features are enhanced.

[0044] First, extract the pixel values ​​of the background area of ​​consecutive image units. This can be achieved through various background modeling methods, such as Gaussian mixture models and codebook models. For example, the Gaussian mixture model can establish multiple Gaussian distribution models for each pixel in the image sequence and update the parameters of these Gaussian distribution models based on the historical values ​​of each pixel. This method can separate the background area from the target area in the image and extract the pixel values ​​of the background area.

[0045] Then, a time series smoothing algorithm is used to eliminate random fluctuations in the background area. The time series smoothing algorithm can smooth the continuous pixel values ​​of the background area and remove random noise. Common time series smoothing algorithms include moving average and exponential smoothing. Taking the moving average method as an example, for the time series value of each pixel in the background area, the average value within the set time window can be calculated as the smoothed pixel value. For example, if a time window of 5 frames is selected, for a certain pixel point, the average value of the pixel values ​​of the current frame and the previous 4 frames can be calculated as the smoothed pixel value of the pixel point.

[0046] Through this background stabilization process, background features become more stable. In city street images, background areas such as the pavement and buildings become clearer and smoother, reducing the interference of random fluctuations on the analysis of the target area. The result is an image unit with stable background features.

[0047] Step S125: performing feature fusion processing on the image unit with enhanced target area features and the image unit with stable background environment features to generate a preprocessed image set, wherein the preprocessed image set includes enhanced target area features and stable background environment features.

[0048] In this embodiment, after the target region feature enhancement processing and the background stabilization processing, an image unit with enhanced target region features and an image unit with stabilized background environment features are obtained. In order to obtain an image set that contains both enhanced target region features and stabilized background environment features, it is necessary to perform feature fusion processing on these two image units.

[0049] Feature fusion can be performed using a variety of methods, such as weighted fusion and splicing fusion. Taking weighted fusion as an example, different weights can be assigned to image cells with enhanced target region features and image cells with stable background features. For example, a higher weight can be assigned to image cells with enhanced target region features to highlight the target region's features, while a lower weight can be assigned to image cells with stable background features to ensure background stability. The corresponding pixel values ​​of the two image cells are then weighted and added together according to the weights to produce the fused pixel value.

[0050] Through the above feature fusion process, the image units with enhanced target region features are fused with the image units with stable background environment features, ultimately generating a preprocessed image set. The image units in this preprocessed image set contain both enhanced target region features (such as clear facial features of pedestrians and clear license plate features of vehicles) and stable background environment features (such as smooth street pavement and clear building outlines).

[0051] Step S130: calling a pre-trained visual large model to perform image analysis processing on the pre-processed image set, and generating an image analysis result including abnormal area positioning information and abnormal type identification.

[0052] In this embodiment, after obtaining the pre-processed image set, it is necessary to perform image analysis processing on it to identify abnormal areas in the image and determine the abnormality type. To this end, a pre-trained visual large model is called.

[0053] The pre-trained large-scale visual model is trained on a large amount of image data and possesses powerful image feature extraction and analysis capabilities. In urban security scenarios, this large-scale visual model can identify various anomalies on the street, such as unusual pedestrian behavior, illegal vehicle driving, and the presence of suspicious objects.

[0054] The specific image analysis and processing process is as follows: Step S131: Input the preprocessed image set into the feature extraction module of the visual large model, perform hierarchical feature extraction on the pixel values ​​of the image units through a multi-layer convolutional neural network, and generate a multi-scale image feature set including basic texture features, structural contour features and context-related features.

[0055] In this embodiment, in an urban security scenario, a pre-processed urban street image set is input into the feature extraction module of the visual macro model. This feature extraction module uses a multi-layer convolutional neural network to perform hierarchical feature extraction on the pixel values ​​of image units to obtain image features of different scales and types.

[0056] Step S1311: In a shallow convolution layer of the multi-layer convolutional neural network, a convolution kernel of a first size is used to perform local feature extraction on the pixel values ​​of the image unit to generate a basic texture feature map reflecting fine-grained texture changes in the image unit.

[0057] In the shallow layers of a multi-layer convolutional neural network, first-size convolution kernels are used to process the pixel values ​​of image cells. These first-size convolution kernels are typically small, allowing for detailed local scanning of the image. In urban street images, these small convolution kernels can capture subtle textures on the street surface, such as the pattern of floor tiles and fine cracks in the pavement, as well as texture features on building walls, such as the arrangement of bricks and stones. By extracting these local features, a basic texture feature map is generated. This basic texture feature map reflects the fine-grained texture variations within the image cell.

[0058] Step S1312: In the middle convolution layer of the multi-layer convolutional neural network, use a second-size convolution kernel to perform feature fusion processing on the basic texture feature map, extract the edge contour and shape structure information of the target object in the image unit, and generate a structural contour feature map, wherein the second size is larger than the first size.

[0059] Once the base texture feature map is generated, it is input into the middle convolutional layer of the multi-layer convolutional neural network. In this layer, a second-size convolution kernel is used for feature fusion. Because the second-size convolution kernel is larger than the first-size convolution kernel, it can operate on the base texture feature map over a wider range. In urban street images, this processing can extract the edge contours and shape structure information of target objects (such as pedestrians, vehicles, and buildings). For example, for pedestrians, the outline of their bodies can be clearly delineated; for vehicles, the outline and general shape of their exterior can be determined. The resulting structural contour feature map more clearly displays the basic form of the target objects in the image.

[0060] Step S1313: In the deep convolutional layer of the multi-layer convolutional neural network, global pooling operations and cross-layer connections are used to integrate contextual information of the structural contour feature map, extract the spatial association and semantic relationship between different areas in the image unit, and generate a contextual association feature map.

[0061] In the deep convolutional layers of a multi-layer convolutional neural network, global pooling operations and cross-layer connections are used to process the structural contour feature map. Global pooling operations aggregate the feature information of the entire image, integrating local feature information into a global feature representation. Cross-layer connections can fuse feature information from different layers, allowing deep features to be combined with feature information from shallow and mid-layer layers. In urban street images, these operations can extract spatial correlations and semantic relationships between different regions. For example, the relative positional relationships between pedestrians and surrounding buildings and vehicles, as well as the relationship between vehicles and traffic signs, can be analyzed. The resulting contextual feature map reflects the comprehensive information between different regions in the image.

[0062] Step S1314: performing channel dimension splicing processing on the basic texture feature map, the structural contour feature map, and the contextual association feature map to generate a fused feature map with multi-scale feature representation.

[0063] After obtaining the basic texture feature map, structural contour feature map, and contextual association feature map, they are spliced ​​in the channel dimension. Channel dimension splicing is to combine these three feature maps in the channel dimension. For example, the basic texture feature map may have a certain number of channels, and the structural contour feature map and contextual association feature map also have different numbers of channels. Arranging them in sequence in the channel dimension forms a new feature map, which contains information from different types of feature maps and has multi-scale feature representation capabilities. In the application of urban street images, the fused feature map can simultaneously reflect fine-grained texture information, structural contour information of the target object, and contextual association information between different regions.

[0064] Step S1315: normalizing the fused feature map to eliminate the difference in response strength between features at different levels, and generating a multi-scale image feature set including basic texture features, structural contour features, and contextual association features.

[0065] To eliminate the differences in response strength between features at different levels in the fused feature map, the fused feature map needs to be normalized. Normalization ensures that the features in the feature map have similar scales and distributions, preventing some features from masking the information of other features due to excessive response strength. In the context of urban street images, normalization ensures that basic texture features, structural contour features, and contextual features have equal importance in subsequent processing. The resulting multi-scale image feature set includes basic texture features, structural contour features, and contextual features.

[0066] Step S132: Input the multi-scale image feature set into the anomaly detection module of the visual large model, perform weighted aggregation on the multi-scale image features through the attention mechanism, highlight the feature response of the abnormal area, and generate a set of abnormal area candidate frames.

[0067] In this embodiment, the generated multi-scale image feature set is input into the anomaly detection module of the visual macro model. The anomaly detection module uses an attention mechanism to perform weighted aggregation on the multi-scale image features.

[0068] The attention mechanism assigns different weights to different features based on their importance. In images of city streets, the attention mechanism assigns higher weights to features in areas where anomalies are likely, highlighting the characteristic responses of these anomalous areas. For example, if a pedestrian with unusual behavior appears on the street, the features of the area surrounding the pedestrian will be enhanced by the attention mechanism.

[0069] Specifically, the attention mechanism assigns a weight to each feature in a multi-scale image feature set. This weight is calculated based on the relevance and importance of the features. For example, features associated with abnormal regions may have higher weights, while features associated with normal backgrounds may have lower weights. By multiplying each feature by its corresponding weight and then aggregating them, the feature responses in abnormal regions can be highlighted.

[0070] During the aggregation process, attention should be paid to the dimensional consistency and feature dimension matching between different features. Ensure that there is no dimensional mismatch between the features during weighted aggregation. If different features have different dimensions, normalization or standardization may be required to bring them to the same dimension. For example, basic texture features, structural contour features, and contextual features may need to be normalized before weighted aggregation to ensure accuracy.

[0071] After weighted aggregation, a set of abnormal region candidate boxes is generated. The candidate boxes in this set represent areas where anomalies may exist. Each candidate box contains information about the location and size of the area. For example, in an image of a city street, an abnormal region candidate box might include a pedestrian climbing over a fence or a vehicle traveling in the wrong direction.

[0072] When the attention mechanism assigns weights to features, it first calculates the similarity between the features. This similarity can be calculated based on a distance metric between feature vectors, such as Euclidean distance or cosine similarity. By calculating the similarity between features, it is possible to determine which features are highly correlated with those in abnormal areas. For example, for pedestrian features in city street images, a feature with a high similarity to known abnormal pedestrian behavior characteristics will be assigned a higher weight.

[0073] After obtaining feature similarity, you can assign initial weights to each feature based on the similarity. Initial weights can be implemented using a function mapping that maps similarity values ​​to a weight range. For example, you can use a linear or nonlinear function to convert similarity values ​​into weights between 0 and 1.

[0074] To further optimize weights, we need to consider feature importance. Feature importance can be determined based on its role in anomaly detection. For example, when detecting abnormal pedestrian behavior, a pedestrian's posture may be more important than their clothing color. Therefore, posture features are given a higher weight. By evaluating the importance of each feature and then adjusting the initial weights, we can determine the final weights.

[0075] When performing weighted aggregation, a weighted calculation can be performed on each position in each feature map. Specifically, for each position's feature value, it is multiplied by the corresponding weight, and then the weighted feature values ​​of all positions are added together to obtain the weighted aggregation result of the feature map.

[0076] The weighted aggregation results of all feature maps are concatenated to form a new feature map, which highlights the characteristic responses of abnormal areas.

[0077] Furthermore, based on the new feature map, an object detection algorithm is used to generate candidate boxes for abnormal regions. Object detection algorithms can use methods such as sliding windows and anchor boxes to search for areas with possible abnormalities on the feature map and generate corresponding candidate boxes. Each candidate box records its position and size in the image.

[0078] Step S133: performing boundary correction processing on the abnormal region candidate frame set, adjusting the position and size of the candidate frame through regression prediction, so that the candidate frame fits the boundary of the actual abnormal region, and obtaining a corrected abnormal region candidate frame set.

[0079] In this embodiment, the candidate frames in the abnormal region candidate frame set may have a certain deviation from the boundaries of the actual abnormal region. In order to improve the accuracy of abnormal region positioning, it is necessary to perform boundary correction processing on the candidate frames.

[0080] Regression prediction is a commonly used boundary correction method. By predicting and adjusting the position and size of the candidate box of the abnormal area, it can be made to fit the boundary of the actual abnormal area more closely. The specific process is as follows: First, features are extracted from each candidate frame in the set of abnormal region candidates. These features can be image features within the candidate frame, or they can be feature representations of multi-scale image features at the candidate frame location obtained by the feature extraction module. For example, for abnormal region candidates in urban street images, features of pedestrians or vehicles within the candidate frame can be extracted, such as the pedestrian's body outline and the vehicle's license plate.

[0081] The extracted candidate box features are then fed into a regression prediction model. This model is trained to learn the relationship between candidate box features and the actual anomaly region boundaries. During training, a large amount of labeled data is used, including candidate box features and corresponding actual anomaly region boundaries. Through training, the regression prediction model establishes a mapping relationship between features and boundaries.

[0082] The regression prediction model predicts the position and size of the candidate box based on the input candidate box features. For example, it predicts the horizontal and vertical distances the candidate box needs to move, as well as the width and height that need to be adjusted. These predicted values ​​can be expressed as offsets and scaling factors.

[0083] Then, based on the output of the regression prediction model, the position and size of the candidate box are adjusted. The original position and size of the candidate box are added with the predicted offset and scaling factor to obtain the corrected position and size of the candidate box. For example, for a candidate box, its original upper left corner coordinates are (x1, y1), width is w1, and height is h1. The regression prediction model predicts a horizontal offset of dx, a vertical offset of dy, a width scaling factor of sw, and a height scaling factor of sh. Then the corrected upper left corner coordinates are (x1+dx, y1+dy), a width of w1*sw, and a height of h1*sh.

[0084] Finally, the corrected positions and sizes of all candidate frames are combined to form a corrected abnormal region candidate frame set. The candidate frames in this abnormal region candidate frame set are more closely aligned with the boundaries of the actual abnormal region, improving the accuracy of abnormal region positioning.

[0085] Step S134: input the corrected abnormal region candidate frame set into the classification module of the visual large model, extract the local features in the candidate frame and perform similarity comparison with the pre-stored normal feature template to generate the abnormal type confidence corresponding to each candidate frame.

[0086] In this embodiment, the modified abnormal region candidate frame set includes regions where abnormalities may exist. In order to determine the abnormality type corresponding to each candidate frame, it is input into the classification module of the visual large model.

[0087] The classification module first extracts local features within the candidate frame. These local features can be the portion of the multi-scale image features obtained by the feature extraction module that resides within the candidate frame, or they can be features extracted specifically from the image within the candidate frame. For example, in a city street image, for a candidate frame containing a pedestrian, the pedestrian's facial features, body posture characteristics, and other features are extracted as local features.

[0088] The extracted local features are then compared for similarity with pre-stored normal feature templates. These templates are collected and organized during the model training phase and represent features from various normal situations. For example, for pedestrians on city streets, a normal feature template might include features of a pedestrian walking in a normal posture or wearing normal clothing.

[0089] Similarity comparison can be performed using a variety of methods, such as Euclidean distance and cosine similarity. By calculating the similarity between local features and the normal feature template, it is possible to determine whether the situation within the candidate box is abnormal. The higher the similarity, the closer it is to the normal situation; the lower the similarity, the greater the possibility of abnormality.

[0090] Based on the similarity comparison results, the anomaly type confidence score is generated for each candidate box. The anomaly type confidence score indicates the likelihood that the situation within the candidate box belongs to a certain anomaly type. For example, for a candidate box containing a vehicle, after similarity comparison, the confidence score for the vehicle driving against the flow anomaly may be 0.8, while the confidence score for the vehicle speeding anomaly may be 0.2.

[0091] When extracting local features within a candidate frame, you can crop and preprocess the image within the candidate frame. The image area within the candidate frame is cropped, and then similar operations are performed on it as the original image preprocessing, such as normalization and scaling, to ensure the accuracy of local feature extraction.

[0092] Next, local features are extracted from the preprocessed image within the candidate frame. This can be done using a convolutional neural network or other feature extraction methods to extract representative local features. For example, a small convolutional neural network can be used to extract features from the image of a pedestrian within the candidate frame to obtain the pedestrian's facial and body posture feature vectors.

[0093] Based on this, the extracted local features are compared one by one with pre-stored normal feature templates. For each normal feature template, the similarity between the local feature and the template is calculated. Various similarity calculation methods can be used, such as calculating the Euclidean distance or cosine similarity between feature vectors. If Euclidean distance is used, a smaller distance indicates a higher similarity; if cosine similarity is used, a similarity value closer to 1 indicates a higher similarity.

[0094] Based on the similarity calculation results, a confidence score for the anomaly type is generated for each candidate box. A threshold can be used to determine the anomaly type. For example, if the similarity between a local feature and a normal feature template falls below a certain threshold, the situation within the candidate box is considered to be likely an anomaly type corresponding to the normal feature template and a corresponding confidence score is assigned. Confidence calculations can be performed based on similarity values, such as using a linear function to convert similarity values ​​into confidence values.

[0095] Each candidate box may have multiple anomaly type confidences. These confidences can be sorted and the highest confidence anomaly type is selected as the primary anomaly type for the candidate box. For example, for a candidate box containing a pedestrian, the confidence level for the running anomaly might be 0.7, while the confidence level for the holding a dangerous object anomaly might be 0.3. In this case, the running anomaly is selected as the primary anomaly type for the candidate box.

[0096] Step S135: Filter abnormal region candidate frames whose confidence exceeds a preset threshold as final abnormal regions, record the position coordinates and corresponding abnormality types of the final abnormal regions, and generate image analysis results including abnormal region positioning information and abnormality type identification.

[0097] In this embodiment, in order to determine the true abnormal area, it is necessary to screen the confidence of the abnormal area candidate frame. The preset threshold is a pre-set value used to determine whether the situation within the candidate frame is truly abnormal.

[0098] The abnormal region candidate boxes whose confidence exceeds the preset threshold are selected as the final abnormal region. For example, in a city street image, if the preset threshold is 0.6 and the abnormality type confidence of a candidate box is 0.8, then the candidate box will be selected as the final abnormal region.

[0099] For each final abnormal region, record its location coordinates and corresponding abnormality type. The location coordinates can be expressed as the coordinates of the candidate box's upper left and lower right corners in the image, or as the center coordinates and width and height. The abnormality type identifier can be a string or numeric code representing the abnormality type of the abnormal region, such as "pedestrian climbing over a fence" or "vehicle driving against the direction."

[0100] The position coordinates of all the final abnormal regions and the abnormality type identification are combined to generate an image analysis result containing the abnormal region location information and the abnormality type identification. The image analysis result can indicate which areas in the image have abnormalities and the type of abnormality.

[0101] Step S140: determining the attributes of abnormal events existing in the scene to be analyzed and the spatiotemporal distribution feature information of the abnormal events in the image according to the image analysis result.

[0102] In this embodiment, based on the image analysis results, the attributes and spatiotemporal distribution characteristics of abnormal events in the scene to be analyzed can be further determined. Abnormal event attributes can describe the specific characteristics of abnormal events, and spatiotemporal distribution characteristics can reflect the distribution of abnormal events in time and space.

[0103] Step S141: parsing the abnormality type identifier in the image analysis result, matching it with a predefined abnormal event attribute mapping table, and determining the category attribute and severity attribute of the abnormal event.

[0104] In this embodiment, the abnormal type identifiers in the image analysis results represent different abnormal situations. The predefined abnormal event attribute mapping table is established during the system development phase, and associates various abnormal type identifiers with corresponding abnormal event attributes.

[0105] By parsing the anomaly type identifier and searching for the corresponding entry in the anomaly event attribute mapping table, the category and severity attributes of the anomaly event are determined. For example, in a city street scenario, if the anomaly type identifier is "Pedestrian climbing over a fence," the mapping table can determine that its category attribute is "Pedestrian Violation" and its severity attribute is "Moderate." The category attribute categorizes anomalies, facilitating subsequent management and processing; the severity attribute indicates the severity of the anomaly, providing a reference for early warning and processing.

[0106] When parsing the exception type identifier, we first perform word segmentation and semantic understanding on it. We break the exception type identifier into multiple keywords, then analyze the semantics of these keywords to accurately understand the meaning of the exception type. For example, the exception type identifier "pedestrian climbing over fence" is broken down into keywords such as "pedestrian," "climbing," and "fence." By understanding the semantics of these keywords, we can determine the specific nature of the exception type.

[0107] Based on the parsed keywords, a search is performed in a predefined abnormal event attribute mapping table. This mapping table contains keywords for various abnormality types and their corresponding category and severity attributes. By matching the keywords, the most suitable entry is found. For example, searching the mapping table for entries containing the keywords "pedestrian," "climbing," and "fence" reveals the corresponding category attribute of "pedestrian violation" and the severity attribute of "moderate."

[0108] If no exact match is found in the abnormal event attribute mapping table, fuzzy matching or similarity calculation can be used. Calculate the similarity between the abnormal event type identifier and each entry in the mapping table, and select the entry with the highest similarity as the match result. For example, similarity can be calculated using methods such as edit distance or cosine similarity.

[0109] For each matching entry, extract the category and severity attributes. Associate these attributes with the anomaly type identifier and record them in the anomaly event attributes. For example, consider "pedestrian violation" and "moderate" as the category and severity attributes for the "pedestrian climbing over a fence" anomaly event.

[0110] Step S142: extracting abnormal region location information of each image unit in the image analysis result, combining it with the timestamp of the image unit, establishing a corresponding relationship between the abnormal region location and time, and obtaining time series location data of the abnormal region.

[0111] In this embodiment, the image analysis result includes the abnormal area location information of each image unit, and each image unit has a timestamp. By combining the abnormal area location information with the timestamp, a corresponding relationship between the abnormal area position and time can be established.

[0112] For example, in a sequence of city street images, for a particular abnormal event (such as a vehicle traveling against the flow), the corresponding abnormal area location information is available in images captured at different times. The location of the abnormal area at each moment is associated with the timestamp at that moment to form time series location data. This time series location data can show how the location of the abnormal area changes over time.

[0113] First, the abnormal region location information for each image unit is extracted from the image analysis results. This abnormal region location information can be the position coordinates of the abnormal region candidate box, such as the coordinates of the upper left and lower right corners, or the center coordinates and width and height. For example, for an abnormal region candidate box in a city street image, its upper left corner coordinates are extracted as (x1, y1) and its lower right corner coordinates are extracted as (x2, y2).

[0114] Then, obtain the timestamp of each image unit. The timestamp indicates the acquisition time of the image, which can be accurate to seconds or milliseconds. For example, the timestamp of an image is 10:10:20 am.

[0115] Next, associate the abnormal area location information with the corresponding timestamp. You can use a data structure, such as a dictionary or list, with the timestamp as the key and the abnormal area location information as the value. For example, the abnormal area location information (x1, y1, x2, y2) corresponding to 10:10:20 am is stored in a dictionary with the key being 10:10:20 am and the value being (x1, y1, x2, y2).

[0116] Finally, the abnormal region location information and timestamp association results for all image units are combined to form a time series location data for the abnormal region. This time series location data can be arranged in chronological order to facilitate subsequent analysis and processing. For example, the abnormal region location information and timestamps for multiple moments can be stored in a list in chronological order to form a complete time series location data.

[0117] Step S143: performing trajectory analysis on the time series position data, calculating the moving direction and moving speed of the abnormal area in the continuous image units, and generating a motion trend attribute of the abnormal event.

[0118] In this embodiment, by performing trajectory analysis on the time series position data of the abnormal region, the movement of the abnormal region in the continuous image units can be understood, thereby generating the motion trend attribute of the abnormal event.

[0119] Step S1431: converting each abnormal region position in the time series position data into a coordinate point in the image coordinate system, wherein the coordinate point includes a horizontal coordinate value and a vertical coordinate value.

[0120] When converting the abnormal area location to a coordinate point in the image coordinate system, the center coordinates are calculated based on the abnormal area's location information. If the abnormal area location information includes the coordinates of the upper left and lower right corners, the horizontal coordinate value of the center coordinate is the average of the horizontal coordinate values ​​of the upper left and lower right corners, and the vertical coordinate value is the average of the vertical coordinate values ​​of the upper left and lower right corners. For example, for the abnormal area location information (x1, y1, x2, y2), the center coordinates are ((x1+x2) / 2, (y1+y2) / 2).

[0121] Step S1432: Calculate the difference between the coordinates of the abnormal area points in adjacent time-stamp image units to obtain the horizontal displacement and the vertical displacement.

[0122] In this embodiment, the horizontal displacement is the difference in horizontal coordinate values ​​between the center coordinates of the abnormal region in adjacent time-stamped image units, and the vertical displacement is the difference in vertical coordinate values ​​between the center coordinates of the abnormal region in adjacent time-stamped image units. For example, in a city street scene, the horizontal coordinate value of the center coordinate of the abnormal region (a vehicle driving illegally) in the image unit at the previous moment is X1, and the vertical coordinate value is Y1. The horizontal coordinate value of the center coordinate of the abnormal region in the image unit at the next moment is X2, and the vertical coordinate value is Y2. Then the horizontal displacement is X2-X1, and the vertical displacement is Y2-Y1. Special attention should be paid to the consistency of dimensions here. Because the dimensions of the coordinate values ​​are pixels, the dimensions of the displacement are also pixels to ensure the accuracy of subsequent calculations.

[0123] Step S1433: Calculate the time interval between adjacent time stamp image units, and obtain the moving speed value of the abnormal area according to the ratio of the module length of the horizontal displacement and the vertical displacement to the time interval.

[0124] In this embodiment, it can be determined by a trigonometric function relationship. The moving direction angle can reflect the moving direction trend of the abnormal area in the image. For example, in mathematical principles, the inverse tangent function can be used to calculate the moving direction angle, that is, according to the ratio of the horizontal displacement and the vertical displacement, the angle value is obtained by the inverse tangent function. In actual scenarios, if the horizontal displacement is positive and the vertical displacement is positive, then the abnormal area moves toward the lower right of the image; if the horizontal displacement is negative and the vertical displacement is positive, then the abnormal area moves toward the lower left of the image, and so on. The dimension of the moving direction angle is the angular unit, which is different from the dimension of the coordinate value and the displacement, but in this calculation process, it is a logical angle calculation and there will be no dimension mismatch problem.

[0125] Step S1434: Calculate the time interval between adjacent time stamp image units, and obtain the moving speed value of the abnormal area according to the ratio of the module length of the horizontal displacement and the vertical displacement to the time interval.

[0126] In this embodiment, the time interval between adjacent timestamp image units is first calculated. This can be obtained by subtracting the timestamp of the previous moment from the timestamp of the next moment. The dimension of the time interval is a time unit, such as seconds. Next, the modulus of the horizontal and vertical displacements is calculated. This involves performing some processing on the sum of the squares of the horizontal and vertical displacements (this does not involve a square root formula, but the calculation logic is similar to that of a square root) to obtain a value representing the actual distance the abnormal region has moved in the image, measured in pixels. Finally, this modulus is divided by the time interval to obtain the movement speed of the abnormal region, measured in pixels per second. For example, in a city street scene, the horizontal and vertical displacements of an illegal vehicle between two adjacent timestamps are a certain number of pixels, and the time interval is a certain number of seconds. The above calculation can be used to determine the vehicle's movement speed between these two timestamps.

[0127] Step S1435: performing sliding window averaging processing on the moving direction angles and moving speed values ​​of a plurality of consecutive time stamps to generate a smoothed moving direction sequence and a smoothed moving speed sequence.

[0128] In this embodiment, sliding window averaging is a commonly used data smoothing method. In urban street scenes, a window size can be set for the movement direction angle and movement speed values ​​of abnormal areas (such as pedestrians or vehicles). For example, data from several consecutive timestamps is selected as a window, and the average movement direction angle and movement speed values ​​within the window are calculated. Over time, the window slides forward, and the average values ​​within the window are recalculated after each slide. This is done to reduce data instability caused by image acquisition noise or minor fluctuations in abnormal areas. The resulting smoothed movement direction sequence and smoothed movement speed sequence can more accurately reflect the movement trends of abnormal areas.

[0129] Step S1436: Generate a motion trend attribute describing the motion direction stability and speed change regularity of the abnormal event according to the change trend of the smoothed motion direction sequence and the change trend of the smoothed motion speed sequence.

[0130] In this embodiment, by observing the smoothed movement direction sequence, if the values ​​in the movement direction sequence vary slightly, it indicates that the movement direction of the abnormal region is relatively stable; if the values ​​vary significantly, it indicates that the movement direction is unstable. Similarly, for the smoothed movement speed sequence, if the values ​​in the movement speed sequence gradually increase, it indicates that the abnormal region is accelerating; if the values ​​gradually decrease, it indicates that the abnormal region is decelerating; if the values ​​remain largely unchanged, it indicates that the abnormal region is moving at a constant speed. By integrating this information, motion trend attributes are generated, such as "stable and uniform movement to the right" and "unstable and accelerated movement to the left." These attributes can help analysts better understand the dynamics of the abnormal event.

[0131] Step S144: performing spatial distribution statistics on the time series position data, calculating the coverage area and distribution density of the abnormal area in the image, and generating the spatial coverage attribute of the abnormal event.

[0132] In urban street scenes, time series location data contains the location information of abnormal areas at different times. For spatial distribution statistics, the first step is to calculate the area covered by the abnormal area in the image. This area can be determined based on the position and size of the abnormal area candidate box. For example, for a rectangular abnormal area candidate box, the width and height are calculated using the coordinates of its upper left and lower right corners. The width and height are then multiplied together to obtain the coverage area. The coverage area is measured in square pixels. During the calculation process, ensure that the coordinate values ​​and size information used are consistent in the dimensions of pixels.

[0133] Next, the distribution density of the abnormal regions is calculated. This density can be expressed as the number of abnormal regions per unit area or as the ratio of the area covered by the abnormal regions to the total image area. For example, the total area covered by the abnormal regions is divided by the total image area to obtain a ratio, which is a representation of the distribution density. The distribution density is dimensionless because the dimensions of the area cancel each other out during the calculation.

[0134] Combining the calculated coverage area and distribution density generates the spatial coverage attribute of the abnormal event. This attribute can reflect the spatial distribution of the abnormal area in the image, such as whether it is concentrated or dispersed, and the size of the coverage area.

[0135] Step S145: combining the category attribute, severity attribute, motion trend attribute and spatial coverage attribute into abnormal event attributes, and combining the time series position data, moving direction, moving speed, coverage area and distribution density into spatiotemporal distribution feature information.

[0136] In the urban street scene, the previous steps have yielded the abnormal event's category attributes (e.g., "vehicle driving violation"), severity attribute (e.g., "moderate"), motion trend attribute (e.g., "steady rightward uniform speed movement"), and spatial coverage attribute (e.g., coverage area of ​​a certain number of square pixels, distribution density of a certain ratio). Combining these attributes creates a complete abnormal event attribute. This attribute comprehensively describes the abnormal event's characteristics, facilitating subsequent processing and decision-making.

[0137] At the same time, the time series location data (recording the location of the abnormal area at different times), movement direction (obtained from the previous trajectory analysis), movement speed (also obtained from trajectory analysis), coverage area (obtained from spatial distribution statistics), and distribution density (obtained from spatial distribution statistics) are combined to form spatiotemporal distribution characteristics. This spatiotemporal distribution characteristics describe the distribution of abnormal events from both temporal and spatial dimensions, helping analysts understand the dynamic changes and impact of abnormal events.

[0138] Step S150: generating a scene warning instruction including event location coordinates based on the abnormal event attributes and the spatiotemporal distribution feature information, and sending the scene warning instruction to a target warning terminal to trigger a response operation.

[0139] In the security scenario of urban streets, based on the abnormal event attributes and spatiotemporal distribution characteristics information obtained above, it is necessary to generate a scene warning instruction containing the event location coordinates.

[0140] Step S151: extracting the category attribute from the abnormal event attribute, matching it with a predefined response strategy mapping table, and determining the type of response operation that needs to be triggered.

[0141] A predefined response strategy mapping table, established during the system design phase, associates different abnormal event categories with corresponding response action types. In a city street scenario, if the abnormal event category attribute is "pedestrian violation," a corresponding entry in the mapping table might identify the response action type as "notify nearby security personnel to investigate." This mapping ensures appropriate response measures are taken for different abnormal event types, improving the efficiency and accuracy of the security system.

[0142] Step S152: extracting the severity attribute from the abnormal event attribute, matching it with a predefined warning level mapping table, and determining the level type of the warning instruction to be generated.

[0143] A predefined warning level mapping table, also created during system design, associates the severity attribute of an abnormal event with the level type of the warning instruction. For example, in a city street scenario, if the severity attribute of an abnormal event is "moderate," a lookup in the mapping table might determine that the warning instruction level type is "Level 2." Different warning level types correspond to different response intensities and processing procedures, ensuring that appropriate warning measures are taken based on the severity of the abnormal event.

[0144] Step S153: extracting the motion trend attribute from the abnormal event attributes, and obtaining descriptive information of the motion direction stability and speed change law of the abnormal event.

[0145] In urban street scenes, the motion trend attribute of an abnormal event includes information describing the stability of the movement direction and the pattern of speed changes. For example, a motion trend attribute of "stable and uniform movement to the right" indicates that the movement direction of an abnormal area (such as an illegal vehicle) is relatively stable and moving to the right at a uniform speed. This information is crucial for predicting the future location and development trend of abnormal events.

[0146] Step S154: extracting the spatial coverage attribute from the abnormal event attributes to obtain the combined features of the coverage area and distribution density of the abnormal event.

[0147] The coverage area and distribution density of the spatial coverage attribute reflect the spatial distribution of abnormal events within the image. In urban street scenes, if the coverage area of ​​abnormal events is large and the distribution density is high, it indicates that the abnormal events have a wide impact and may require more extensive response measures. For example, if multiple abnormal areas are concentrated in a certain area of ​​the street, with a large coverage area and high distribution density, it may be necessary to notify more security personnel to handle the situation in that area.

[0148] Step S155: extracting the time series position data in the spatiotemporal distribution feature information, and establishing a corresponding relationship between the abnormal area position and the timestamp as a historical trajectory record.

[0149] Time series location data records the location of abnormal areas at different times. In urban street scenarios, the location information of each abnormal area is associated with the corresponding timestamp to form a historical trajectory record. This historical trajectory record can be stored in a list or database format to facilitate subsequent query and analysis. For example, the location of a violating vehicle at different times can be recorded. By reviewing the historical trajectory record, the vehicle's travel path and range can be understood.

[0150] Step S156: Combined with the description information of the movement direction stability and speed change law of the abnormal event, a trend enhancement analysis is performed on the historical trajectory record to establish a dynamic prediction model of the abnormal area position changing over time.

[0151] In urban street scenes, information describing the stability of the movement direction and speed changes of abnormal events can be used to conduct in-depth analysis of historical trajectory records. For example, if the movement direction of an abnormal area (such as a pedestrian or vehicle) is stable and the speed is uniform, the approximate location of the abnormal area at a certain point in the future can be predicted based on the historical trajectory records and movement speed. A dynamic prediction model can be established through linear interpolation or machine learning algorithms. This dynamic prediction model can predict the location of the abnormal area in subsequent time-stamped image units based on the current location information and movement trends.

[0152] Step S1561: extracting the abnormal area position coordinates corresponding to consecutive timestamps from the historical trajectory record to form a position coordinate sequence.

[0153] In urban street scenarios, historical trajectory records contain the location information of abnormal areas at different times. Arranging the location coordinates corresponding to these consecutive timestamps in chronological order forms a location coordinate sequence. Each element in this location coordinate sequence is a coordinate point, containing horizontal and vertical coordinate values, and the dimension is pixel.

[0154] Step S1562: extracting a directional stability parameter and a speed change law parameter from the motion trend attribute, wherein the directional stability parameter reflects the degree of fluctuation of the moving direction, and the speed change law parameter reflects the increase and decrease pattern of the moving speed.

[0155] Directional stability parameters can be expressed as the variance or standard deviation of the smoothed movement direction sequence. Smaller variances or standard deviations indicate more stable movement directions. Velocity variation parameters can be derived by analyzing the smoothed movement velocity sequence, for example, to determine whether the velocity is uniform, accelerating, or decelerating. The dimensions of these parameters must be consistent with the corresponding physical quantities to ensure the accuracy of subsequent calculations.

[0156] Step S1563: performing differential processing on the position coordinate sequence to obtain a position change sequence of adjacent time stamps.

[0157] Differential processing involves calculating the difference between adjacent position coordinates to obtain the horizontal and vertical position changes. For example, for two adjacent coordinate points in a position coordinate sequence, the horizontal and vertical differences between their coordinates are calculated to form a position change sequence. The position change is measured in pixels, the same as the position coordinates.

[0158] Step S1564: performing correlation analysis on the position variation sequence, the directional stability parameter, and the speed variation law parameter to determine the corresponding relationship between the position variation and the motion trend attribute.

[0159] This correlation analysis allows us to understand how the position change varies with the motion trend attributes. For example, if the directional stability parameter indicates a stable direction of movement, the position change is likely to be relatively stable in that direction. If the speed change parameter indicates an acceleration pattern, the position change is likely to increase gradually.

[0160] Step S1565: constructing a dynamic prediction model based on the corresponding relationship, wherein the dynamic prediction model is used to calculate the position coordinate prediction value of the subsequent time stamp according to the position coordinate, directional stability parameter and speed change law parameter of the current time stamp.

[0161] Dynamic prediction models can be constructed using methods such as regression analysis and neural networks. During the construction process, ensure that the input parameters (position coordinates at the current timestamp, directional stability parameters, and speed change law parameters) and the output position coordinate prediction values ​​are consistent in dimensions, both in pixels.

[0162] Step S1566: The parameters of the dynamic prediction model are calibrated in real time by a sliding window method so that the deviation between the predicted value and the actual position coordinates meets the preset consistency requirement.

[0163] The sliding window method selects a set number of consecutive timestamps as a window and adjusts the parameters of the dynamic prediction model within this window. By continuously sliding the window and updating the model parameters in real time, the deviation between the predicted value and the actual location coordinates remains within a preset range, improving the model's prediction accuracy.

[0164] Step S157: Based on the combined features of the coverage area and distribution density of the abnormal event and in combination with the position coordinates in the historical trajectory record, the current impact range boundary and dynamic expansion trend of the abnormal area in the image are calculated.

[0165] In urban street scenes, the coverage area and distribution density of abnormal events reflect the spatial distribution of the abnormal area, while the location coordinates in the historical trajectory record record the movement path of the abnormal area. Based on this information, the current impact range boundary of the abnormal area in the image can be calculated. For example, by integrating the location and coverage area of ​​the abnormal area at different times, a minimum bounding rectangle or polygon containing all the abnormal areas can be determined. The boundary of this bounding rectangle is the current impact range boundary.

[0166] Furthermore, based on historical trajectory records and the movement trends of abnormal events, the dynamic expansion trend of the abnormal area can be predicted. If the abnormal area is accelerating and its coverage area is gradually increasing, its impact range may expand rapidly; if the abnormal area is decelerating and its coverage area is stable, its impact range may expand more slowly. Through a comprehensive analysis of these factors, the current impact range boundaries and dynamic expansion trends of the abnormal area can be accurately calculated.

[0167] For example, step S1571: extracting the coverage area and distribution density of the current timestamp from the spatial coverage attribute, the coverage area is the total number of pixels in the abnormal area, and the distribution density is the density of abnormal pixels per unit area.

[0168] In urban street scenes, the spatial coverage attribute includes the coverage area and distribution density of abnormal areas at different times. Extracting the coverage area and distribution density corresponding to the current timestamp ensures that the latest data is used for analysis. The coverage area is measured in square pixels, while the distribution density is a dimensionless value.

[0169] Step S1572: extracting the abnormal area position coordinates of the current timestamp from the historical trajectory record, and determining the geometric center coordinates of the abnormal area.

[0170] Based on the historical trajectory records, find the coordinates of the abnormal area corresponding to the current timestamp. Determine the geometric center coordinates of the abnormal area by calculating the average of these coordinates or using other appropriate methods. The dimension of the geometric center coordinates is pixels, which is the same as the dimension of the position coordinates.

[0171] Step S1573: Generate an initial impact range boundary based on the geometric center according to the coverage area and the geometric center coordinates. The initial impact range boundary is the minimum circumscribed rectangle that includes all pixels in the abnormal area.

[0172] Based on the geometric center coordinates and the size of the coverage area, a minimum bounding rectangle is generated. For example, the side length of the rectangle is calculated based on the coverage area, and then the coordinates of the four vertices of the rectangle are determined with the geometric center as the center to obtain the initial impact range boundary.

[0173] Step S1574: extracting the coverage area and distribution density of the previous timestamp from the historical trajectory record, and calculating the time change rate of the coverage area and the time change rate of the distribution density, wherein the time change rate reflects the expansion speed or contraction speed of the abnormal area.

[0174] By comparing the coverage area and distribution density between the previous and current timestamps, the difference is calculated and then divided by the time interval to obtain the time rate of change. The dimensions of the time rate of change are related to the original physical quantity. The time rate of change of coverage area is measured in pixels squared per second, while the time rate of change of distribution density is the proportional change per second.

[0175] Step S1575: Dynamically adjust the boundary of the initial impact range based on the time change rate of the coverage area and the time change rate of the distribution density to generate a dynamic expansion trend description reflecting the expansion trend or contraction trend of the abnormal area.

[0176] If the time rate of change of the coverage area is positive, it indicates that the anomaly area is expanding; if it is negative, it indicates that the anomaly area is shrinking. Based on these change rates, the initial impact range boundaries are adjusted accordingly, such as expanding or reducing the size of the boundaries, thus generating a dynamic expansion trend description.

[0177] Step S1576: Smoothing the dynamic expansion trend description to generate the current impact range boundary and dynamic expansion trend information including a boundary coordinate sequence and a trend direction vector.

[0178] Smoothing can reduce instability caused by data fluctuations or minor changes. By smoothing the dynamic expansion trend, a more accurate boundary coordinate sequence and trend direction vector are obtained. The boundary coordinate sequence records the coordinates of each point on the boundary of the impact range, and the trend direction vector indicates the direction and speed of the anomaly area's expansion. This information can more clearly demonstrate the current impact range and dynamic expansion trend of the anomaly area.

[0179] Step S158: The position of the abnormal area predicted by the dynamic prediction model in the subsequent time-stamp image unit is used as a predicted trajectory point.

[0180] In urban street scenarios, a dynamic prediction model has been established based on information such as the historical trajectory and movement trends of abnormal events. This dynamic prediction model can predict the location of the abnormal area in subsequent time-stamped image units. These predicted locations are known as predicted trajectory points. For example, based on the current location and movement trend of the abnormal area, the predicted location of the abnormal area at several future time stamps can be predicted. This location information is presented as coordinate points, each containing horizontal and vertical coordinate values, all measured in pixels. Predicted trajectory points can help analysts understand the future direction of the abnormal area in advance. The accuracy of the predicted trajectory points depends on the reliability of the dynamic prediction model, which in turn is closely related to factors such as the integrity of the historical trajectory data and the accuracy of the movement trend analysis. In practical applications, new abnormal area location data can be continuously collected to update and optimize the dynamic prediction model to improve the accuracy of the predicted trajectory points.

[0181] Step S159: combining the position coordinates in the historical trajectory record, the abnormal area position coordinates in the current timestamp image unit, and the predicted trajectory point coordinates into an event location coordinate chain with temporal continuity.

[0182] In urban security scenarios, historical trajectory records detail the location coordinates of abnormal areas at different past timestamps. These coordinates reflect the historical movement path of the abnormal area. The abnormal area location coordinates in the image unit at the current timestamp clearly define the abnormal area's current location. The predicted trajectory point coordinates, derived from a dynamic prediction model, represent the possible locations of the abnormal area at future timestamps. Combining these three coordinates in chronological order forms a temporally continuous event location coordinate chain.

[0183] For example, if a vehicle is detected driving illegally on a city street, its historical trajectory records the vehicle's location coordinates at various times, starting from the moment it enters the monitoring area. Image analysis determines the vehicle's current location at the current timestamp. Combined with a dynamic prediction model, this predicts the vehicle's likely locations at the next several timestamps. Arranging these location coordinates at different times creates a complete coordinate chain, clearly demonstrating the vehicle's movement trajectory from the past to the present and into the future, demonstrating temporal continuity.

[0184] The event location coordinate chain can help security personnel fully understand the development process and trends of abnormal events and prepare for response in advance. For example, the coordinate chain can determine whether an illegal vehicle will enter a critical area or collide with other vehicles or pedestrians, allowing timely action.

[0185] Step S1510: Combine the response operation type, warning level type, current impact range boundary, dynamic expansion trend and event location coordinate chain into a scene warning instruction, and the scene warning instruction includes coordinate information for identifying the event location, boundary information for indicating the impact range, dynamic information for describing the expansion trend and operation information for triggering a response.

[0186] In urban security scenarios, the various information determined in the previous steps is integrated to form scenario-based warning instructions. Response action types are specific response measures determined by the type of abnormal event, such as notifying nearby security personnel to investigate or dispatching traffic police for traffic diversion. Warning levels are categorized based on the severity of the abnormal event. Different warning levels correspond to different response intensities and processing procedures. For example, a level one warning might require a full emergency response, while a level two warning might require less severe measures.

[0187] The current impact range boundary is derived by analyzing and processing the spatial coverage properties of the anomaly region, clearly defining the area affected by the anomaly at the current moment. The dynamic expansion trend reflects the potential changes in the anomaly region over time, including whether it will expand, shrink, or remain stable. The event location coordinate chain records the location of the anomaly region from the past to the present and into the future, demonstrating the temporal continuity and spatial trajectory of the anomaly event.

[0188] Combining this information to form a scenario-based warning instruction plays a number of important roles. Coordinate information used to identify the location of the incident allows relevant personnel to quickly and accurately locate the site of the abnormal event; boundary information used to indicate the scope of impact helps them understand the scale of the abnormal event's impact; dynamic information used to describe expansion trends helps predict the development of the abnormal event; and operational information used to trigger a response clarifies the specific actions that should be taken. For example, if an abnormal incident occurs where vehicles gather illegally on a city street, a scenario-based warning instruction can clearly inform security personnel of the location of the vehicle gathering, the current scope of impact, the possible future expansion direction, and the response measures that need to be taken, such as evacuating the crowd and directing traffic.

[0189] Step S1511: Send the scene warning instruction to the target warning terminal to trigger a response operation.

[0190] In the urban security system, target warning terminals are key devices for receiving and processing scene warning instructions. These target warning terminals can be computer terminals installed in the security monitoring center, mobile terminal devices carried by security personnel (such as smart phones, tablets, etc.), or other smart devices connected to the security system.

[0191] Once a scenario warning command is generated, it is sent to the target warning terminal. To prevent errors or loss during transmission, a data verification and retransmission mechanism can be implemented. For example, before sending the command, the command data can be encrypted and encoded, and a checksum can be added. Upon receiving the command, the target warning terminal can verify the data based on the checksum. If any errors are found, it can request a retransmission from the sender.

[0192] Once the target warning terminal successfully receives a scenario warning command, it can trigger corresponding response actions based on the information in the scenario warning command. If the scenario warning command requires notification of nearby security personnel to investigate, the terminal will automatically send a message to the security personnel's mobile terminal, including detailed information such as the location, type, and severity of the abnormal event. At the same time, an audible and visual alarm may be triggered to alert security personnel to their immediate attention. If the scenario warning command requires the dispatch of traffic police to divert traffic, the terminal will interact with the traffic command center's system to arrange for police to go to the appropriate streets to carry out traffic control.

[0193] Furthermore, the target warning terminal records and stores the scene warning instructions for subsequent query and analysis. These records can be used to evaluate the effectiveness of abnormal event handling, summarize lessons learned, and optimize the security system's response strategy. For example, by analyzing the handling records of multiple similar abnormal events, it is possible to identify which response actions were effective and which require improvement, thereby continuously improving the efficiency and level of urban security.

[0194] It's worth noting that training large-scale visual models requires a large amount of labeled image data. This data should cover both normal and abnormal conditions across a wide range of urban security scenarios to ensure the model learns comprehensive feature information. The training process involves forward propagation and backpropagation, continuously adjusting model parameters to ensure that the model's output matches the labeled data as closely as possible. Training parameters include the learning rate, batch size, and number of training rounds, and these parameters need to be optimized based on the specific dataset and model architecture.

[0195] For collected image data, especially those involving privacy-sensitive data, appropriate privacy protection and anti-leakage technologies must be implemented. During the data collection phase, images can be anonymized to remove personally identifiable information such as facial features and license plate numbers. During data storage and transmission, encryption technology is used to prevent unauthorized access and tampering. Furthermore, strict access rights are set to ensure that only authorized personnel can access and process this data, ensuring data security and privacy.

[0196] Figure 2 The following diagram illustrates exemplary hardware and software components of a visual large model-based image analysis and early warning system 100 that can implement the concepts of the present application, as provided in some embodiments of the present application. For example, the processor 120 can be used in the visual large model-based image analysis and early warning system 100 to perform the functions described in the present application.

[0197] The image analysis and early warning system 100 based on a large visual model can be a general-purpose server or a special-purpose server, both of which can be used to implement the image analysis and early warning method based on a large visual model of the present application. Although only one server is shown in this application, for convenience, the functions described in this application can be implemented in a distributed manner on multiple similar platforms to balance the processing load.

[0198] For example, the image analysis and early warning system 100 based on the visual large model may include a network port 110 connected to the network, one or more processors 120 for executing program instructions, a communication bus 130, and storage media 140 in different forms, such as a disk, ROM, or RAM, or any combination thereof. Exemplarily, the image analysis and early warning system 100 based on the visual large model may also include program instructions stored in ROM, RAM, or other types of non-transitory storage media, or any combination thereof. The method of the present application can be implemented according to these program instructions. The image analysis and early warning system 100 based on the visual large model also includes an I / O interface 150 between the computer and other input and output devices.

[0199] For ease of explanation, only one processor is described in the image analysis and early warning system 100 based on the visual big model. However, it should be noted that the image analysis and early warning system 100 based on the visual big model in the present application may also include multiple processors, so the steps performed by one processor described in the present application may also be performed jointly or individually by multiple processors. For example, if the processor of the image analysis and early warning system 100 based on the visual big model executes step A and step B, it should be understood that step A and step B may also be executed jointly by two different processors or individually in one processor. For example, the first processor executes step A, the second processor executes step B, or the first processor and the second processor execute steps A and B together.

[0200] In addition, an embodiment of the present invention further provides a readable storage medium, in which computer-executable instructions are preset. When a processor executes the computer-executable instructions, the above-mentioned image analysis and early warning method based on the visual large model is implemented.

[0201] It should be noted that in order to simplify the description of the present invention and thus help understand one or more embodiments of the invention, in the foregoing description of the embodiments of the present invention, multiple features are sometimes combined into one embodiment, figure or description thereof.

Claims

1. An image analysis and early warning method based on a large visual model, characterized in that: The method comprises: Acquire an original image data set of a scene to be analyzed, wherein the original image data set includes a plurality of image units that are continuously acquired and marked with time stamps; Performing a preprocessing operation on the original image data set to obtain a preprocessed image set, wherein the preprocessed image set includes enhanced target area features and stable background environment features; Calling a pre-trained visual large model to perform image analysis processing on the pre-processed image set to generate an image analysis result including abnormal area location information and abnormal type identification; Determine, based on the image analysis results, the attributes of the abnormal events present in the scene to be analyzed and the spatiotemporal distribution feature information of the abnormal events in the image; A scene warning instruction including event location coordinates is generated based on the abnormal event attributes and the spatiotemporal distribution feature information, and the scene warning instruction is sent to a target warning terminal to trigger a response operation.

2. The image analysis and early warning method based on visual large model according to claim 1 is characterized in that: The preprocessing operation is performed on the original image data set to obtain a preprocessed image set, wherein the preprocessed image set contains enhanced target area features and stable background environment features, including: Performing noise suppression on the original image data set, reducing random noise interference in the image unit by an adaptive filtering algorithm, and obtaining a noise-suppressed image unit; Performing brightness equalization processing on the noise-suppressed image units, adjusting the brightness distribution of the image units based on global histogram statistics so that the brightness mean values ​​of different image units remain consistent, thereby obtaining brightness-equalized image units; Performing target region enhancement processing on the brightness-equalized image unit, identifying the target region boundary in the image unit through an edge detection algorithm, and performing nonlinear enhancement on the pixel value of the target region based on the boundary information to obtain an image unit with enhanced target region features; Performing background stabilization processing on the image unit with enhanced target area features, extracting background area pixel values ​​of continuous image units, eliminating random fluctuations of the background area through a time series smoothing algorithm, and obtaining image units with stable background environment features; The image unit with enhanced target area features and the image unit with stable background environment features are subjected to feature fusion processing to generate a preprocessed image set, wherein the preprocessed image set includes enhanced target area features and stable background environment features.

3. The image analysis and early warning method based on visual large model according to claim 2 is characterized in that: The performing noise suppression processing on the original image data set, reducing random noise interference in the image unit by an adaptive filtering algorithm, and obtaining the noise-suppressed image unit, comprises: Calculating statistical features of local neighborhood pixel values ​​for each pixel point in a single image unit, wherein the statistical features include a mean and a variance of the neighborhood pixel values; Determine a noise intensity estimate for each pixel point based on the mean and variance of the local neighborhood pixel values; Dynamically adjusting the size of the filter window and the filter weight based on the noise intensity estimate, wherein a larger filter window and a smaller center pixel weight are used for an area with a larger noise intensity estimate, and a smaller filter window and a larger center pixel weight are used for an area with a smaller noise intensity estimate; Performing a weighted average calculation on the neighborhood pixel values ​​of each pixel point to generate a filtered pixel value, wherein the weighted average calculation uses a dynamically adjusted filter window and filter weight; The filtered pixel values ​​of all pixels are combined into a single noise-suppressed image unit, and the above processing is performed on each image unit in the original image data set to obtain a noise-suppressed image unit set.

4. The image analysis and early warning method based on visual large model according to claim 1 is characterized in that: The calling of the pre-trained visual large model to perform image analysis processing on the pre-processed image set to generate an image analysis result containing abnormal area positioning information and abnormal type identification, including: Inputting the preprocessed image set into the feature extraction module of the visual macro model, performing hierarchical feature extraction on the pixel values ​​of the image units through a multi-layer convolutional neural network, and generating a multi-scale image feature set including basic texture features, structural contour features and contextual association features; Inputting the multi-scale image feature set into the anomaly detection module of the visual large model, performing weighted aggregation on the multi-scale image features through the attention mechanism, highlighting the feature response of the abnormal area, and generating a set of abnormal area candidate frames; Performing boundary correction processing on the abnormal region candidate frame set, adjusting the position and size of the candidate frame through regression prediction, so that the candidate frame fits the boundary of the actual abnormal region, and obtaining a corrected abnormal region candidate frame set; Input the corrected abnormal region candidate frame set into the classification module of the visual large model, extract local features within the candidate frame and perform similarity comparison with pre-stored normal feature templates to generate the abnormality type confidence corresponding to each candidate frame; The abnormal region candidate boxes whose confidence exceeds a preset threshold are selected as the final abnormal region, the position coordinates of the final abnormal region and the corresponding abnormal type are recorded, and an image analysis result including abnormal region positioning information and abnormal type identification is generated.

5. The image analysis and early warning method based on visual large model according to claim 4 is characterized in that: The pre-processed image set is input into the feature extraction module of the visual large model, and the pixel values ​​of the image units are subjected to hierarchical feature extraction through a multi-layer convolutional neural network to generate a multi-scale image feature set including basic texture features, structural contour features and contextual association features, including: In a shallow convolutional layer of the multi-layer convolutional neural network, a convolution kernel of a first size is used to perform local feature extraction on pixel values ​​of an image unit to generate a basic texture feature map reflecting fine-grained texture changes in the image unit; In a middle convolutional layer of the multi-layer convolutional neural network, a convolution kernel of a second size is used to perform feature fusion processing on the basic texture feature map, extract edge contour and shape structure information of the target object in the image unit, and generate a structural contour feature map, wherein the second size is larger than the first size; In the deep convolutional layer of the multi-layer convolutional neural network, global pooling operations and cross-layer connections are used to integrate contextual information of the structural contour feature map, extract spatial associations and semantic relationships between different regions in the image unit, and generate a contextual association feature map; Performing channel dimension splicing processing on the basic texture feature map, the structural contour feature map, and the context association feature map to generate a fused feature map with multi-scale feature representation; The fused feature map is normalized to eliminate the difference in response intensity between features at different levels, and to generate a multi-scale image feature set including basic texture features, structural contour features, and contextual association features.

6. The image analysis and early warning method based on visual large model according to claim 1 is characterized in that: The determining, based on the image analysis result, the attributes of the abnormal event existing in the scene to be analyzed and the spatiotemporal distribution feature information of the abnormal event in the image picture includes: Parsing the abnormality type identifier in the image analysis result, matching it with a predefined abnormal event attribute mapping table, and determining the category attribute and severity attribute of the abnormal event; Extracting the abnormal area positioning information of each image unit in the image analysis result, combining it with the time stamp mark of the image unit, establishing a corresponding relationship between the abnormal area position and time, and obtaining time series position data of the abnormal area; Performing trajectory analysis on the time series position data, calculating the moving direction and moving speed of the abnormal area in the continuous image units, and generating a motion trend attribute of the abnormal event; Performing spatial distribution statistics on the time series position data, calculating the coverage area and distribution density of the abnormal area in the image, and generating spatial coverage attributes of the abnormal event; The category attribute, severity attribute, movement trend attribute and space coverage attribute are combined into abnormal event attributes, and the time series position data, moving direction, moving speed, coverage area and distribution density are combined into spatiotemporal distribution feature information.

7. The image analysis and early warning method based on visual large model according to claim 6 is characterized in that: The performing trajectory analysis on the time series position data, calculating the moving direction and moving speed of the abnormal area in the continuous image units, and generating the motion trend attribute of the abnormal event includes: Converting each abnormal area position in the time series position data into a coordinate point in an image coordinate system, wherein the coordinate point includes a horizontal coordinate value and a vertical coordinate value; Calculate the difference between the coordinates of the abnormal area in adjacent time stamp image units to obtain the horizontal displacement and the vertical displacement; Calculating the moving direction angle of the abnormal area according to the horizontal displacement and the vertical displacement; Calculating the time interval between adjacent time stamp image units, and obtaining the moving speed value of the abnormal area according to the ratio of the modulus length of the horizontal displacement and the vertical displacement to the time interval; Perform sliding window averaging on the moving direction angle and moving speed values ​​of multiple consecutive time stamps to generate a smoothed moving direction sequence and a smoothed moving speed sequence; According to the changing trend of the smoothed moving direction sequence and the changing trend of the smoothed moving speed sequence, a motion trend attribute describing the stability of the moving direction and the speed changing law of the abnormal event is generated.

8. The image analysis and early warning method based on visual large model according to claim 1 is characterized in that: The generating of a scene warning instruction including event location coordinates based on the abnormal event attributes and the spatiotemporal distribution feature information includes: Extract the category attribute from the abnormal event attribute, match it with the predefined response strategy mapping table, and determine the type of response operation that needs to be triggered; Extracting the severity attribute from the abnormal event attribute, matching it with a predefined warning level mapping table, and determining the level type of the warning instruction to be generated; Extracting the motion trend attribute from the abnormal event attributes to obtain descriptive information of the motion direction stability and speed change law of the abnormal event; Extracting the spatial coverage attribute from the abnormal event attributes to obtain the combined features of the coverage area and distribution density of the abnormal event; Extracting time series position data from the spatiotemporal distribution feature information, and establishing a correspondence between the abnormal area position and the timestamp as a historical trajectory record; Combined with the description information of the movement direction stability and speed change law of the abnormal event, the historical trajectory record is subjected to trend enhancement analysis to establish a dynamic prediction model of the abnormal area position change over time; Based on the combined features of the coverage area and distribution density of the abnormal event, combined with the location coordinates in the historical trajectory record, the current impact range boundary and dynamic expansion trend of the abnormal area in the image are calculated; The position of the abnormal area predicted by the dynamic prediction model in the subsequent time-stamp image unit is used as a predicted trajectory point; Combining the position coordinates in the historical trajectory record, the abnormal area position coordinates in the current timestamp image unit, and the predicted trajectory point coordinates into an event location coordinate chain containing time continuity; The response operation type, warning level type, current impact range boundary, dynamic expansion trend and event location coordinate chain are combined into a scene warning instruction, which includes coordinate information for identifying the event location, boundary information for indicating the impact range, dynamic information for describing the expansion trend and operation information for triggering a response.

9. The image analysis and early warning method based on visual large model according to claim 8 is characterized in that: The method combines the description information of the movement direction stability and speed change law of the abnormal event, performs trend enhancement analysis on the historical trajectory record, and establishes a dynamic prediction model of the abnormal area position changing over time, including: Extracting the abnormal area position coordinates corresponding to consecutive timestamps from the historical trajectory record to form a position coordinate sequence; Extracting a directional stability parameter and a speed change law parameter from the motion trend attribute, wherein the directional stability parameter reflects the degree of fluctuation of the moving direction, and the speed change law parameter reflects the increase and decrease pattern of the moving speed; Performing differential processing on the position coordinate sequence to obtain a position change sequence of adjacent time stamps; Correlation analysis is performed on the position change sequence with the directional stability parameter and the speed change law parameter to determine the corresponding relationship between the position change and the motion trend attribute; Building a dynamic prediction model based on the corresponding relationship, the dynamic prediction model is used to calculate the position coordinate prediction value of the subsequent time stamp according to the position coordinate, directional stability parameter and speed change law parameter of the current time stamp; The parameters of the dynamic prediction model are calibrated in real time using a sliding window method so that the deviation between the predicted value and the actual position coordinates meets the preset consistency requirements.

10. An image analysis and early warning system based on a large visual model, characterized in that: It includes a processor and a memory, the memory is connected to the processor, the memory is used to store programs, instructions or codes, and the processor is used to execute the programs, instructions or codes in the memory to implement the image analysis and early warning method based on the visual large model as described in any one of claims 1 to 9.

Citation Information

Patent Citations

  • Method for detecting and locating emergent abnormal event of group

    CN107506734A

  • Intelligent machine vision detection method and system based on image processing and storage medium

    CN119205719A

  • Image enhancement method and device, electronic equipment and storage medium

    CN119904374A

  • Construction personnel dangerous area border crossing detection method and system based on neural network

    CN120014549A

  • Invader detection method and device applied to perimeter security system

    CN120088456A

Cited By

  • Computer information monitoring method, system and device, medium and program product

    CN121301138A

  • Fan abnormity monitoring method, device and equipment based on image recognition and medium

    CN121788911A

  • Fan abnormality monitoring method, device and equipment based on image recognition and medium

    CN121788911B