Intelligent monitoring-oriented monitoring video efficient compression method and system
Through visual recognition and macroblock size adjustment technology, key target areas in surveillance videos are identified and compression rates are optimized, solving the problem of key target blur in existing technologies and improving the clarity and compression efficiency of surveillance videos.
Patent Information
- Application Number
- CN202511013704.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-23
- Publication Date
- 2025-10-14
- Estimated Expiration
- 2045-07-23
AI Technical Summary
Existing video compression methods fail to differentiate between key target areas when processing surveillance videos, resulting in blurring of key targets and loss of details after compression, affecting subsequent event analysis and identification.
Through visual recognition technology, key target areas in surveillance videos are identified, the target change probability of each pixel is analyzed, and the macroblock size is adjusted according to the probability to perform video compression to reduce the compression rate of key target areas and increase the compression rate of background areas.
It improves the image clarity of key targets, reduces the amount of background data, and improves the compression accuracy and overall compression rate of surveillance videos. It is suitable for intelligent monitoring systems in complex scenarios.
Smart Images

Figure CN120547360B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of video compression, and in particular to a method and system for efficiently compressing surveillance videos for intelligent monitoring. Background Art
[0002] With the rapid development of intelligent technology, surveillance video systems have been widely used in various fields, including urban security, traffic management, industrial production, and public spaces. The massive amount of surveillance data not only places higher demands on storage resources, but also poses new challenges to video transmission efficiency and quality. To improve the storage efficiency and transmission performance of video data, video compression technology has become a core tool. Traditional H.264 / AVC and subsequent video coding standards have made significant progress in balancing compression efficiency and image quality and have been widely used in various surveillance scenarios. However, with the increasing intelligence of video surveillance systems, the structure of video content and the differences in key areas have gradually become apparent, placing higher demands on efficient and accurate compression technology.
[0003] Existing video compression methods primarily rely on intra- and inter-frame motion vector analysis when processing surveillance video, considering only pixel changes and motion compensation. These methods fail to differentiate between key target areas within the surveillance scene that hold security or business value. For example, in an urban traffic monitoring scenario, existing methods compress all image areas equally. Even areas of focus for vehicles, pedestrians, or unusual events receive the same compression ratio as the background. This can cause key targets to appear blurred and details to be lost, hindering subsequent event analysis and identification.
[0004] Therefore, in the video compression process, how to reduce the compression of key targets while ensuring the overall compression rate of the surveillance video, thereby enhancing the clarity of the key targets in the compressed video, is a problem that needs to be solved urgently in current technology. Summary of the Invention
[0005] In order to solve the technical problem of insufficient clarity of key objects in compressed surveillance videos, the present invention aims to provide a surveillance video efficient compression method and system for intelligent monitoring. The technical solutions adopted are as follows:
[0006] In a first aspect, an embodiment of the present invention provides a method for efficiently compressing surveillance videos for intelligent monitoring, the method comprising:
[0007] Performing visual recognition on the first surveillance video to obtain visual recognition information, wherein the visual recognition information is used to identify an area in the video frame corresponding to a preset surveillance target;
[0008] Analyzing the visual recognition information to obtain a target change probability for each pixel in each video frame, wherein the target change probability indicates a probability that the corresponding pixel is one of a plurality of pixels constituting the monitored target in the corresponding video frame;
[0009] adjusting the sizes of the multiple initial macroblocks corresponding to each video frame according to the target change probability of each pixel in each video frame to obtain multiple adjusted macroblocks corresponding to each video frame, wherein the number of the multiple adjusted macroblocks corresponding to the same video frame is less than or equal to the number of the corresponding multiple initial macroblocks, and the target change probability is negatively correlated with the size of the adjusted macroblock covering the corresponding pixel;
[0010] According to the multiple adjustment macroblocks corresponding to each of the video frames, video compression is performed on the multiple video frames to obtain a second monitoring video.
[0011] Furthermore, performing visual recognition on the first surveillance video to obtain visual recognition information includes:
[0012] Taking the monitoring target as a detection target, performing target detection on each of the plurality of video frames included in the first monitoring video to obtain a plurality of detection information, wherein the plurality of detection information corresponds to the plurality of video frames in a one-to-one manner, and the detection information is used to indicate n detection boxes in the corresponding video frames, where n is a non-negative integer;
[0013] Perform target tracking based on the multiple detection information to obtain multiple tracking information, wherein the multiple tracking information corresponds one-to-one to multiple monitoring moments corresponding to the first surveillance video, and the tracking information is used to indicate m target frames in the video frame at the corresponding monitoring moments, where the target frames are the detection frames of the indicated target that appear multiple times consecutively in the first surveillance video, and m is a non-negative integer;
[0014] Obtaining a plurality of target change information based on the plurality of tracking information, wherein the plurality of target change information corresponds one-to-one to the plurality of monitoring moments, and the target change information includes: a position change sequence of a center point of each target frame in m target frames included in a video frame at the corresponding monitoring moment;
[0015] In the multiple target frames indicated by the multiple tracking information, the visual recognition information is generated based on the position change sequence of the center point of each target frame, and the area corresponding to the preset monitoring target includes the center point of the corresponding target frame, and the pixel points whose correlation between the position change sequence and the position change sequence of the center point of the corresponding target frame is greater than the correlation threshold.
[0016] Furthermore, in the plurality of target frames indicated by the plurality of tracking information, generating the visual recognition information according to a position change sequence of a center point of each target frame includes:
[0017] Performing local feature extraction on each of the plurality of target frames indicated by the plurality of tracking information to obtain a plurality of feature information corresponding one-to-one to the plurality of target frames, the feature information including a plurality of feature points associated with the corresponding target frames, the plurality of feature points being used to represent an outline of the target indicated by the corresponding target frame;
[0018] Correlation calculations are performed on the position change sequences of the plurality of feature points included in each piece of feature information and the position change sequence of the corresponding center point, respectively, to obtain a plurality of feature-related information corresponding one-to-one to the plurality of feature information, the feature-related information including: correlation of the position change sequence between each feature point and the corresponding center point among the plurality of feature points included in the corresponding feature information;
[0019] The visual recognition information is generated based on the multiple feature-related information and the multiple feature information, wherein the area corresponding to the preset monitoring target includes the feature points whose correlation with the position change sequence between the corresponding center point is greater than the correlation threshold.
[0020] Furthermore, analyzing the visual recognition information to obtain the target change probability of each pixel in each video frame includes:
[0021] Obtaining, based on the plurality of feature information, a plurality of boundary information corresponding one-to-one to the plurality of target frames, wherein the boundary information includes a plurality of boundary points corresponding one-to-one to the corresponding plurality of feature points, the boundary points being intersections of a straight line connecting the corresponding feature point and the corresponding center point and a boundary of the corresponding target frame;
[0022] Correlation calculation is performed on each of the plurality of boundary information items, wherein the plurality of position change sequences of the plurality of boundary points included in each boundary information item and the position change sequences of the corresponding feature points are respectively performed to obtain a plurality of boundary-related information items corresponding one-to-one to the plurality of boundary information items, wherein the boundary-related information items include: correlation of the position change sequences between each boundary point and the corresponding feature point in the plurality of boundary points included in the corresponding boundary information item;
[0023] Calculations are performed based on the multiple feature-related information and the multiple boundary-related information to determine multiple feature attention information, the multiple feature attention information and the multiple target frames having one-to-one correspondence, the feature attention information including multiple attention levels of the multiple feature points associated with the corresponding target frames, the attention levels being the ratios of the corresponding boundary-related values to the corresponding feature-related values, the boundary-related values being used to represent the correlation of the position change sequence between the corresponding boundary points and the corresponding feature points, and the feature-related values being used to represent the correlation of the position change sequence between the corresponding feature points and the corresponding center points;
[0024] The plurality of feature attention information is analyzed to obtain a target change probability of each pixel in each of the video frames.
[0025] Furthermore, analyzing the plurality of feature attention information to obtain the target change probability of each pixel in each of the video frames includes:
[0026] Analyzing the plurality of feature attention information to obtain a plurality of change probability values corresponding to each pixel point in each of the video frames, wherein the plurality of change probability values correspond one-to-one to the plurality of feature points indicated by the corresponding feature attention information, and the change probability values are jointly determined based on the corresponding attention degree and the corresponding point distance, where the point distance is the distance between the corresponding feature point and the corresponding pixel point;
[0027] For each pixel point in each of the video frames, the maximum value of the multiple change probability values corresponding to the pixel point is determined as the target change probability value of the pixel point to obtain the target change probability of each pixel point in each of the video frames.
[0028] Furthermore, the sizes of the multiple initial macroblocks corresponding to each video frame are adjusted according to the target change probability of each pixel in each video frame to obtain the multiple adjusted macroblocks corresponding to each video frame, including:
[0029] According to the target change probability of each pixel point in each of the video frames, each initial macroblock corresponding to each of the video frames is iteratively divided into one or more adjusted macroblocks to obtain multiple adjusted macroblocks corresponding to each of the video frames, wherein the size of the initial macroblock is greater than or equal to the size of the adjusted macroblock iteratively divided from the initial macroblock, and when the size of the adjusted macroblock is greater than a preset minimum size and the size of the adjusted macroblock is smaller than the size of the corresponding initial macroblock, the average of multiple target change probabilities of multiple pixel points covered by the adjusted macroblock in the corresponding video frame is less than a preset probability threshold.
[0030] Furthermore, the extracting local features of the multiple target frames indicated by the multiple tracking information to obtain multiple feature information corresponding to the multiple target frames includes:
[0031] Merging the target frames indicated by the tracking information to obtain a plurality of merged frames corresponding to the target frames, wherein the merged frames are unions of the corresponding target frames and their associated frames, and the associated frames and the corresponding target frames indicate the same target;
[0032] Local features are extracted for each of the multiple merged frames to obtain multiple feature information corresponding to the multiple target frames.
[0033] In a second aspect, another embodiment of the present invention provides a surveillance video efficient compression system for intelligent monitoring, the system comprising:
[0034] a visual recognition module, configured to perform visual recognition on the first surveillance video to obtain visual recognition information, wherein the visual recognition information is used to identify an area in the video frame corresponding to a preset surveillance target;
[0035] an information analysis module, configured to analyze the visual recognition information to obtain a target change probability for each pixel in each video frame, wherein the target change probability indicates a probability that the corresponding pixel is one of a plurality of pixels constituting the monitored target in the corresponding video frame;
[0036] a macroblock adjustment module, configured to adjust the sizes of the multiple initial macroblocks corresponding to each video frame according to a target change probability of each pixel in each video frame, thereby obtaining multiple adjusted macroblocks corresponding to each video frame, wherein the number of the multiple adjusted macroblocks corresponding to the same video frame is less than or equal to the number of the corresponding multiple initial macroblocks, and the target change probability is negatively correlated with the size of the adjusted macroblock covering the corresponding pixel;
[0037] The video compression module is used to perform video compression on the multiple video frames according to the multiple adjustment macroblocks corresponding to each of the video frames to obtain a second monitoring video.
[0038] In a third aspect, another embodiment of the present invention further provides an electronic device, comprising a processor, a memory, and a computer program stored in the memory and executable on the processor, wherein the computer program implements the steps of the method described in the first aspect when executed by the processor.
[0039] In a fourth aspect, another embodiment of the present invention further provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the method described in the first aspect are implemented.
[0040] The present invention has the following beneficial effects:
[0041] The present invention first performs visual recognition on a first surveillance video to identify the area corresponding to a preset surveillance target in each video frame of the first surveillance video, and then determines the probability of each pixel in each video frame constituting the corresponding surveillance target by analyzing the information obtained from the visual recognition processing, so as to adjust the size of multiple initial macroblocks corresponding to each video frame accordingly. By limiting the aforementioned probability and the size of the macroblock covering the corresponding similar points to be negatively correlated, the compression amplitude of the pixels constituting the surveillance target in each video frame is reduced as much as possible, thereby reducing the compression rate of the area containing the key target focused on by the surveillance video and correspondingly increasing the compression rate of the background area not containing the key target. This can enable the image area corresponding to the key target to retain more detailed information during the compression process, so that the image clarity of the key target in the compressed second surveillance video is improved, and at the same time, the data volume of the background part of the non-key target in the second surveillance video is reduced, so as to enhance the clarity of the key target in the compressed video while reducing the data volume of the background part in the compressed video. This can improve the compression accuracy and overall compression rate of the surveillance video, making it suitable for use in intelligent surveillance systems in various complex scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] In order to more clearly illustrate the technical solutions and advantages of the embodiments of the present invention or the prior art, the following briefly introduces the drawings required for use in the embodiments or the prior art descriptions. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0043] Figure 1 A schematic flow chart of an efficient compression method for surveillance video for intelligent monitoring provided by one embodiment of the present invention;
[0044] Figure 2 A schematic diagram of the structure of a surveillance video efficient compression system for intelligent monitoring provided by one embodiment of the present invention;
[0045] Figure 3 The present invention provides a schematic structural diagram of an electronic device according to an embodiment of the present invention. DETAILED DESCRIPTION
[0046] To further illustrate the technical means and effectiveness of the present invention in achieving its intended purpose, the following, in conjunction with the accompanying drawings and preferred embodiments, describes in detail the specific implementation, structure, features, and effectiveness of a high-efficiency surveillance video compression method and system for intelligent monitoring proposed by the present invention. In the following description, different references to "one embodiment" or "another embodiment" do not necessarily refer to the same embodiment. Furthermore, specific features, structures, or characteristics of one or more embodiments may be combined in any suitable manner.
[0047] Unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention belongs.
[0048] The specific scheme of the efficient compression method and system for surveillance video for intelligent monitoring provided by the present invention is described in detail below with reference to the accompanying drawings.
[0049] This paper proposes a high-efficiency compression method for surveillance video for intelligent monitoring. Figure 1 , which shows a schematic flow chart of a method for efficiently compressing surveillance videos for intelligent monitoring provided by one embodiment of the present invention, the method comprising the following steps:
[0050] Step S1: Perform visual recognition on a first surveillance video to obtain visual recognition information.
[0051] The visual recognition information is used to identify an area in a video frame corresponding to a preset monitoring target.
[0052] The above-mentioned first surveillance video can be understood as the original surveillance video obtained by a camera deployed in the surveillance scene, wherein the original surveillance video can be obtained by using the Real Time Streaming Protocol (RTSP) or by reading from the local storage of the camera.
[0053] Exemplarily, the monitoring scene may be a shopping mall monitoring scene, an elevator monitoring scene, a road monitoring scene, etc.
[0054] In the application, after obtaining the first surveillance video, the frame sequence of the first surveillance video can be time synchronized and frame rate balanced to ensure consistent intervals between frames, eliminate time discontinuity problems caused by frame loss, repeated frames, etc., and thereby ensure the smooth execution of subsequent visual recognition and video compression operations.
[0055] In addition, in order to improve the accuracy of subsequent target detection, before performing visual recognition on the first surveillance video, you can also choose to perform data preprocessing on each video frame in the first surveillance video to ensure that the image quality and clarity of each video frame meet the quality requirements and clarity requirements of the input image required by the visual recognition algorithm, wherein the data preprocessing includes: color space adjustment (such as adjusting the RGB image to a grayscale image, or adjusting the RGB image to an HSV image), brightness normalization (such as global brightness normalization or histogram stretching), noise suppression (such as Gaussian filtering or median filtering) and image enhancement (such as histogram equalization or sharpening). At least one of the following.
[0056] The aforementioned video frames that have undergone data preprocessing and their metadata information (such as camera number, timestamp, frame index) will be stored in a cache area to provide stable, reliable and standardized input data for the subsequent visual recognition process and video compression process. The setting of the cache area can realize the decoupling of the preprocessing process and the subsequent processing process (i.e., the visual recognition process and the video compression process) of the video frame, so as to balance the processing efficiency difference between the preprocessing process and the subsequent processing process by dynamically adjusting the data read and write rate of the cache area, thereby ensuring the smooth execution of the efficient compression method of surveillance video for intelligent monitoring described in the present invention.
[0057] The above-mentioned camera number is used to uniquely identify the camera that captures the corresponding video frame among the multiple cameras in the aforementioned monitoring scene, the above-mentioned timestamp is used to indicate the acquisition time or shooting time of the corresponding video frame, and the above-mentioned frame index is used to indicate the order and position of the corresponding video frame in the multiple video frames included in the first monitoring video.
[0058] The aforementioned visual recognition can be understood as: based on computer vision technology, a process of identifying video frames that may contain preset monitoring targets from each video frame of the first monitoring video (or each video frame that has undergone data preprocessing), and marking the video frames that may contain preset monitoring targets (such as simple selection in the form of a detection frame, or accurately marking the outline of the monitoring target through feature extraction and related calculations).
[0059] Exemplarily, target detection can be performed on each video frame in the first surveillance video based on HOG (Histogram of Oriented Gradients) combined with SVM (Support Vector Machine), DPM (Deformable Part Model), YOLO (You Only Look Once), SSD (Single Shot MultiBox Detector), Anchor-Free algorithm or Transformer model to obtain the aforementioned visual recognition information.
[0060] For example, the monitoring target may be a pedestrian, an animal, a vehicle, smoke, water stains, etc.
[0061] In one example, the visual recognition information may include multiple detection frame positions in each video frame (such as the center point coordinates of the corresponding rectangular detection frame and the size of the corresponding rectangular detection frame, or the lower left endpoint coordinates of the corresponding rectangular detection frame and the upper right endpoint coordinates of the corresponding rectangular detection frame, etc.), each detection frame includes the confidence of the monitoring target (the value is in the range of 0-1, the larger the value, the higher the confidence, and the higher the probability that the corresponding detection frame includes the monitoring target), and the category of the monitoring target indicated by each detection frame (when there are multiple categories of monitoring targets).
[0062] Step S2: Analyze the visual recognition information to obtain the target change probability of each pixel in each video frame.
[0063] The target change probability is used to indicate the probability that the corresponding pixel point is one of the pixel points constituting the monitoring target in the corresponding video frame.
[0064] In one example, for the multiple pixels included in each of the video frames, the target change probability of the pixels located in the area corresponding to the preset monitoring target can be determined as 1, while the target change probability of the pixels located in the background area can be determined as 0.
[0065] In another example, for each of the multiple pixel points included in the video frame, the distance between each pixel point and the center point of the detection frame closest to it in the corresponding video frame can be calculated, and the value after normalization (such as maximum and minimum value normalization, Z-score standardization, etc.) of the distance can be determined as the target change probability of the pixel point.
[0066] Step S3: adjusting the sizes of the multiple initial macroblocks corresponding to each video frame according to the target change probability of each pixel in each video frame to obtain multiple adjusted macroblocks corresponding to each video frame.
[0067] The number of the multiple adjusted macroblocks corresponding to the same video frame is less than or equal to the number of the corresponding multiple initial macroblocks, and the target change probability is negatively correlated with the size of the adjusted macroblock covering the corresponding pixel point.
[0068] The negative correlation between the target change probability and the size of the adjusted macroblock covering the corresponding pixel point should be understood as follows: the greater the target change probability, the greater the probability of the smaller the size of the adjusted macroblock covering the corresponding pixel point.
[0069] In the present invention, resizing the initial macroblock may be understood as a process of splitting a macroblock of a larger size (e.g., 16×16) into multiple macroblocks of smaller sizes (e.g., 8×8, 4×4), that is, adjusting the size of the macroblock to be smaller than or equal to the size of the corresponding initial macroblock.
[0070] Exemplarily, among the multiple target change probabilities of the multiple pixel points covered by each initial macroblock corresponding to each of the video frames, the proportion of the number of pixel points whose target change probability exceeds a set probability threshold (such as 0.7 or 0.8, etc.) can be counted, and when the proportion of the number exceeds the set proportion threshold (such as 0.65), the macroblock is equally split into 4 smaller macroblocks, and the above-mentioned judgment is continued on the small macroblocks after splitting to determine whether to continue splitting the small macroblocks after splitting, and so on, until the size of the macroblock is adjusted to the set minimum size (such as 1×1), or the proportion of the number corresponding to the macroblock does not exceed the said proportion threshold.
[0071] Step S4: performing video compression on the multiple video frames according to the multiple adjustment macroblocks corresponding to each of the video frames to obtain a second monitoring video.
[0072] The aforementioned video compression process can be understood as: using the H.264 video compression algorithm, according to the multiple adjusted macroblocks corresponding to each of the video frames, through the targeted configuration of macroblocks of different sizes, the video compression encoding process is performed on each video frame in the first surveillance video.
[0073] After the above video compression processing, the compression rate of the area corresponding to the preset monitoring target in the video frame will be lower than the compression rate of the background area (ie, the area not including the monitoring target) in the video frame.
[0074] The compression rate in the present invention should be understood as: the ratio of the data volume of the area (such as the area corresponding to the preset monitoring target or the background area) in the second monitoring video to the data volume of the area in the first monitoring video.
[0075] The present invention first performs visual recognition on a first surveillance video to identify the area corresponding to a preset surveillance target in each video frame of the first surveillance video, and then determines the probability of each pixel in each video frame constituting the corresponding surveillance target by analyzing the information obtained from the visual recognition processing, so as to adjust the size of multiple initial macroblocks corresponding to each video frame accordingly. By limiting the aforementioned probability and the size of the macroblock covering the corresponding similar points to be negatively correlated, the compression amplitude of the pixels constituting the surveillance target in each video frame is reduced as much as possible, thereby reducing the compression rate of the area containing the key target focused on by the surveillance video and correspondingly increasing the compression rate of the background area not containing the key target. This can enable the image area corresponding to the key target to retain more detailed information during the compression process, so that the image clarity of the key target in the compressed second surveillance video is improved, and at the same time, the data volume of the background part of the non-key target in the second surveillance video is reduced, so as to enhance the clarity of the key target in the compressed video while reducing the data volume of the background part in the compressed video. This can improve the compression accuracy and overall compression rate of the surveillance video, making it suitable for use in intelligent surveillance systems in various complex scenarios.
[0076] In one embodiment, performing visual recognition on the first surveillance video to obtain visual recognition information includes:
[0077] Taking the monitoring target as a detection target, performing target detection on each of the plurality of video frames included in the first monitoring video to obtain a plurality of detection information, wherein the plurality of detection information corresponds to the plurality of video frames in a one-to-one manner, and the detection information is used to indicate n detection boxes in the corresponding video frames, where n is a non-negative integer;
[0078] Perform target tracking based on the multiple detection information to obtain multiple tracking information, wherein the multiple tracking information corresponds one-to-one to multiple monitoring moments corresponding to the first surveillance video, and the tracking information is used to indicate m target frames in the video frame at the corresponding monitoring moments, where the target frames are the detection frames indicating the tracked target that appear multiple times consecutively in the first surveillance video, and m is a non-negative integer;
[0079] Obtaining a plurality of target change information based on the plurality of tracking information, wherein the plurality of target change information corresponds one-to-one to the plurality of monitoring moments, and the target change information includes: a position change sequence of a center point of each target frame in m target frames included in a video frame at the corresponding monitoring moment;
[0080] In the multiple target frames indicated by the multiple tracking information, the visual recognition information is generated based on the position change sequence of the center point of each target frame, and the area corresponding to the preset monitoring target includes the center point of the corresponding target frame, and the pixel points whose correlation between the position change sequence and the position change sequence of the center point of the corresponding target frame is greater than a correlation threshold (such as 0.8).
[0081] In this embodiment, the above target detection operation can be completed based on deep learning target detection technology, such as completing the above target detection operation based on the YOLO algorithm.
[0082] For example, the aforementioned timestamp and frame index can be used to perform frame-by-frame analysis on multiple video frames in the first surveillance video to identify whether each video frame includes a detection target through the YOLO model, and when the video frame includes a detection target, the position of the detection target included in the video frame is marked by the anchor box (which can be understood as the aforementioned detection box, usually a rectangle), and the classification corresponding to the selected detection target is shown (when there are multiple categories of detection targets). At this time, the model input of the YOLO model is the video frame and its metadata information after data preprocessing, and the corresponding model output is the position information of the target detected in the video frame (the coordinate information of the anchor box), the category label (such as used to indicate that the detected target is a person, vehicle or other target) and the confidence score.
[0083] In the application, the appropriate version of the YOLO model can be flexibly selected to perform the above target detection action according to the actual monitoring area size and the characteristics of the detection target, and the present invention is not limited to this.
[0084] Among them, the multiple monitoring moments corresponding to the first monitoring video can be understood as multiple timestamps of multiple video frames.
[0085] It should be understood that the target box indicating the target appears multiple times in a row in the first surveillance video, which can be understood as: the target box indicating the target is detected in multiple consecutive video frames in the first surveillance video, and the number of video frames of the multiple consecutive video frames is greater than a preset number threshold (such as 180, the video frame rate is 60 frames per second), or the duration of the multiple consecutive video frames is greater than a preset duration threshold (such as 3 seconds).
[0086] For example, the target tracking process may be:
[0087] At time t among the multiple monitoring moments corresponding to the first monitoring video, the area corresponding to the i-th detection frame in the video frame whose timestamp matches time t is , you can remember The center point is .
[0088] Taking time t as the starting time, calculate The intersection over union (IoU) between all detection frames in the video frame at time t+1 is calculated, and the detection frame (at time t+1) with the largest IoU and exceeding the preset intersection over union threshold (such as 0.9) is considered to be the same as the detection frame at time t+1. Indicates the detection box of the same object;
[0089] Calculate the intersection over union (IoU) between the detection frame (at time t+1) and all the detection frames in the video frame at time t+2, and consider the detection frame (at time t+2) with the largest IoU that exceeds the preset IoU threshold (such as 0.9) as the detection frame with the largest IoU. Indicates the detection box of the same object;
[0090] The same logic is repeated until no matching image is recognized in the new video frame. Indicates the detection box of the same target. If the previously counted If the number of detection frames indicating the same target reaches a preset number threshold, or the duration reaches a preset duration threshold, it can be determined A target box belonging to the above definition, statistics and with Indicates the position change of the center point of the detection box of the same target, so that Center point Position change sequence;
[0091] It should be understood that no When indicating the detection frame of the same target, if the previously counted If the number of detection frames indicating the same target does not reach the preset number threshold, and the duration does not reach the preset duration threshold, then It will not be determined as the aforementioned target frame.
[0092] The position change sequence of the center point of the target frame can be used to represent the overall change trend of the target indicated by the target frame (such as a pedestrian, vehicle, etc.) in consecutive video frames. Therefore, the pixel points whose correlation between the position change sequence and the position change sequence of the center point of the corresponding target frame is greater than the correlation threshold can be understood as the pixel points used to represent the target indicated by the target frame, that is, the pixel points located within the area corresponding to the preset monitoring target. Based on the center point of the corresponding target frame and the pixel points whose correlation between the position change sequence and the position change sequence of the center point of the corresponding target frame is greater than the correlation threshold, the area corresponding to the preset monitoring target can be more accurately located in the corresponding video frame, thereby reducing the probability of redundant points unrelated to the target indicated by the target frame being classified as the area corresponding to the preset monitoring target, thereby making the subsequent compression of the area corresponding to the preset monitoring target and the background area (the area not corresponding to the preset monitoring target) more accurate, thereby further improving the compression accuracy and overall compression rate of the monitoring video.
[0093] For example, the correlation between the position change sequence of the pixel points and the position change sequence of the center point of the corresponding target frame can be calculated based on the Pearson correlation coefficient between the two.
[0094] For example, the correlation between the position change sequence of the pixel point and the position change sequence of the center point of the corresponding target box The mathematical representation of can be:
[0095]
[0096] In the above formula, represents the Pearson correlation coefficient between two series; Indicates the first The position change sequence of pixel points in the x-coordinate component; Represents the position change sequence of the center point of the i-th target box in the video frame at time t on the x-coordinate component, Indicates the first The position change sequence of pixel points on the y coordinate component; Represents the position change sequence of the center point of the i-th target box in the video frame at time t on the y-coordinate component. x and y can be understood as two coordinate components used to indicate the point position in the coordinate system of the corresponding video frame (for example, when the coordinate system is a plane rectangular coordinate system, x can be understood as the horizontal axis coordinate of the corresponding point in the plane linear coordinate system, and y can be understood as the vertical axis coordinate of the corresponding point in the plane rectangular coordinate system).
[0097] In one embodiment, generating the visual recognition information according to a position change sequence of a center point of each target frame in the plurality of target frames indicated by the plurality of tracking information includes:
[0098] Performing local feature extraction on each of the plurality of target frames indicated by the plurality of tracking information to obtain a plurality of feature information corresponding one-to-one to the plurality of target frames, the feature information including a plurality of feature points associated with the corresponding target frames, the plurality of feature points being used to represent an outline of the target indicated by the corresponding target frame;
[0099] Correlation calculations are performed on the position change sequences of the plurality of feature points included in each piece of feature information and the position change sequence of the corresponding center point, respectively, to obtain a plurality of feature-related information corresponding one-to-one to the plurality of feature information, the feature-related information including: correlation of the position change sequence between each feature point and the corresponding center point among the plurality of feature points included in the corresponding feature information;
[0100] The visual recognition information is generated based on the multiple feature-related information and the multiple feature information, wherein the area corresponding to the preset monitoring target includes the feature points whose correlation with the position change sequence between the corresponding center point is greater than the correlation threshold.
[0101] In this embodiment, local features are extracted for multiple target frames respectively to obtain multiple feature points representing the outline of the target indicated by each target frame, and then the correlation of the position change sequence of the multiple feature points with the center point of the corresponding target frame is calculated one by one, and combined with the setting of the correlation threshold, several feature points whose correlation with the position change sequence between the corresponding center point is greater than the correlation threshold are identified to obtain a more accurate outline representation of the target indicated by the target frame, that is, to obtain a more accurate area corresponding to the preset monitoring target.
[0102] In the application, the aforementioned local feature extraction operation can be performed based on the SIFT (Scale Invariant Feature Transform) algorithm. The correlation between the position change sequence of the feature points and the position change sequence of the center point of the corresponding target frame can be calculated based on the Pearson correlation coefficient between the two. The specific calculation process is shown in the above example. To avoid repetition, the calculation will not be described again.
[0103] In one example, the process of generating the visual recognition information based on the plurality of feature-related information may be:
[0104] For any feature-related information, multiple feature-related values (correlation of the position change sequence between the corresponding feature points and the corresponding center point) of multiple feature points in the current feature-related information are compared with the correlation threshold respectively, so as to select several feature points whose feature-related values are greater than the correlation threshold from the current feature-related information, and then use the selected several feature points as the outline of the target indicated by the current feature-related information to obtain a more accurate target outline, that is, to obtain a more accurate area corresponding to the preset monitoring target (the area surrounded by the target outline).
[0105] In one embodiment, extracting local features from the multiple target frames indicated by the multiple tracking information to obtain multiple feature information corresponding to the multiple target frames includes:
[0106] Merging the target frames indicated by the tracking information to obtain a plurality of merged frames corresponding to the target frames, wherein the merged frames are unions of the corresponding target frames and their associated frames, and the associated frames and the corresponding target frames indicate the same target;
[0107] Local features are extracted for each of the multiple merged frames to obtain multiple feature information corresponding to the multiple target frames.
[0108] In this embodiment, by taking the union of the target frame and multiple associated frames that indicate the same target as the target frame, the corresponding target frame can fully include the image portion of the indicated target, thereby avoiding feature omissions caused by detection errors. Local feature extraction is performed based on this, and multiple feature points that more comprehensively indicate the outline of the target can be obtained.
[0109] In one embodiment, analyzing the visual recognition information to obtain the target change probability of each pixel in each video frame includes:
[0110] Obtaining, based on the plurality of feature information, a plurality of boundary information corresponding one-to-one to the plurality of target frames, wherein the boundary information includes a plurality of boundary points corresponding one-to-one to the corresponding plurality of feature points, the boundary points being intersections of a straight line connecting the corresponding feature point and the corresponding center point and a boundary of the corresponding target frame;
[0111] Correlation calculation is performed on each of the plurality of boundary information items, wherein the plurality of position change sequences of the plurality of boundary points included in each boundary information item and the position change sequences of the corresponding feature points are respectively performed to obtain a plurality of boundary-related information items corresponding one-to-one to the plurality of boundary information items, wherein the boundary-related information items include: correlation of the position change sequences between each boundary point and the corresponding feature point in the plurality of boundary points included in the corresponding boundary information item;
[0112] Calculations are performed based on the multiple feature-related information and the multiple boundary-related information to determine multiple feature attention information, the multiple feature attention information and the multiple target frames having one-to-one correspondence, the feature attention information including multiple attention levels of the multiple feature points associated with the corresponding target frames, the attention levels being the ratios of the corresponding boundary-related values to the corresponding feature-related values, the boundary-related values being used to represent the correlation of the position change sequence between the corresponding boundary points and the corresponding feature points, and the feature-related values being used to represent the correlation of the position change sequence between the corresponding feature points and the corresponding center points;
[0113] The plurality of feature attention information is analyzed to obtain a target change probability of each pixel in each of the video frames.
[0114] In this embodiment, based on the setting of boundary points, and combined with the correlation calculation results of the position change sequence between the boundary points and the corresponding feature points, as well as the correlation calculation results of the position change sequence between the feature points corresponding to the boundary points and the corresponding center points, the ratio of the two is determined as the attention degree of the corresponding feature points, so as to measure the severity of the position change of the feature points. This can reduce the problem of missed recognition caused by the drastic change in the position of pixel points when facing complex moving targets, ensure the comprehensiveness and accuracy of the extracted target contour, and thus ensure the recognition accuracy of the area corresponding to the preset monitoring target under complex motion conditions.
[0115] It should be noted that the boundary point in the present invention is specifically: one or more intersection points of a straight line connecting the corresponding feature point and the corresponding center point and the boundary of the corresponding target frame, the intersection point that is closest to the corresponding feature point.
[0116] Among them, the introduction of boundary points and boundary-related values can suppress the interference introduced by complex motion when the target of focus undergoes complex motion such as rotation, so as to accurately identify the edge contours of complex moving targets, thereby reducing the risk of missing feature points that were originally part of the target, and thus avoiding the problem of over-compression caused by this, so that the target can still maintain sufficient integrity and clarity in the compressed video picture.
[0117] In the application, to avoid abnormal calculation of attention (the feature correlation value is 0), you can set:
[0118] The attention degree = (corresponding boundary correlation value + 0.1) The corresponding feature correlation value, where the value of 0.1 can be adaptively adjusted according to actual needs.
[0119] Furthermore, analyzing the plurality of feature attention information to obtain the target change probability of each pixel in each of the video frames includes:
[0120] Analyzing the plurality of feature attention information to obtain a plurality of change probability values corresponding to each pixel point in each of the video frames, wherein the plurality of change probability values correspond one-to-one to the plurality of feature points indicated by the corresponding feature attention information, and the change probability values are jointly determined based on the corresponding attention degree and the corresponding point distance, where the point distance is the distance between the corresponding feature point and the corresponding pixel point;
[0121] For each pixel in each video frame, the maximum value of the multiple change probability values corresponding to the pixel is determined as the target change probability value for the pixel, thereby obtaining the target change probability for each pixel in each video frame. In one example, the corresponding change probability value can be directly determined as the ratio of the corresponding attention level to the corresponding point distance.
[0122] In another example, the product of the corresponding attention and the corresponding reference standard deviation can be calculated first, and then the ratio of the product and the corresponding point distance can be determined as the corresponding change probability value. By introducing the reference standard deviation, the point distance can be normalized, thereby reducing the impact of extreme point distances on the calculated change probability value, so as to obtain a more accurate change probability value, wherein the reference standard deviation can be understood as the standard deviation of the position change sequence of the corresponding center point.
[0123] In this embodiment, the attention level of the current feature point and the point distance between the current feature point and the current pixel point are summarized to comprehensively determine the probability that the current pixel point changes position with the current feature point (represented by the aforementioned change probability value), and the maximum change probability value of the current pixel point is selected to determine the probability that the current pixel point is part of a complex moving target (represented by the aforementioned target change probability). Based on this, the size of the macroblock covering each pixel point in the current video frame is adaptively determined to minimize the size of the macroblock covering multiple pixel points indicating complex moving targets (and multiple pixel points indicating complex moving targets need to be covered by a larger number of macroblocks), thereby improving the compression quality of complex moving targets.
[0124] In one embodiment, adjusting the sizes of the multiple initial macroblocks corresponding to each video frame according to the target change probability of each pixel in each video frame to obtain the multiple adjusted macroblocks corresponding to each video frame includes:
[0125] According to the target change probability of each pixel point in each of the video frames, each initial macroblock corresponding to each of the video frames is iteratively divided into one or more adjusted macroblocks to obtain multiple adjusted macroblocks corresponding to each of the video frames, wherein the size of the initial macroblock is greater than or equal to the size of the adjusted macroblock iteratively divided from the initial macroblock, and when the size of the adjusted macroblock is greater than a preset minimum size and the size of the adjusted macroblock is smaller than the size of the corresponding initial macroblock, the average of multiple target change probabilities of multiple pixel points covered by the adjusted macroblock in the corresponding video frame is less than a preset probability threshold.
[0126] Exemplarily, the process of iteratively dividing each initial macroblock corresponding to each video frame into one or more adjusted macroblocks according to the target change probability of each pixel in each video frame to obtain multiple adjusted macroblocks corresponding to each video frame may be:
[0127] Suppose the size of an initial macroblock is 16×16. The mean of the target change probabilities of the 256 pixels covered by the initial macroblock is calculated. If the mean is less than or equal to the aforementioned probability threshold (e.g., 0.85), the initial macroblock is determined as the adjusted macroblock. In this case, the size of the adjusted macroblock is the same as that of the initial macroblock.
[0128] If the mean is greater than the aforementioned probability threshold, the size of the initial macroblock is evenly divided to obtain four 8×8 macroblocks. Subsequently, the mean of the target change probability of the 64 pixels covered by each 8×8 macroblock is calculated. If the four means are all less than or equal to the aforementioned probability threshold, the four 8×8 macroblocks are all determined as the adjusted macroblocks. At this time, the size of the adjusted macroblock is smaller than the size of the initial macroblock.
[0129] If, when calculating the mean of the target change probabilities of the 64 pixels covered by each 8×8 macroblock, it is found that the mean corresponding to any 8×8 macroblock is still greater than the aforementioned probability threshold, then the 8×8 macroblock is further divided into four 4×4 macroblocks, and the mean of the target change probabilities of the 16 pixels covered by each 4×4 macroblock is again calculated and compared with the aforementioned probability threshold.
[0130] The same process is repeated until the macroblock size reaches the set minimum size (such as 2×2 or 1×1).
[0131] In this embodiment, the mean of the target change probabilities of the multiple pixel points covered by the macroblock is calculated, and combined with the setting of the probability threshold to determine whether to reduce the size of the macroblock, so as to achieve adaptive adjustment of the macroblock size, thereby reducing the size of the macroblock covering the multiple pixel points indicating complex moving targets as much as possible, thereby improving the compression quality of the complex moving targets.
[0132] This invention proposes a surveillance video high-efficiency compression system for intelligent monitoring. Figure 2 , which shows a schematic structural diagram of a surveillance video efficient compression system 200 for intelligent monitoring provided by one embodiment of the present invention, the system comprising:
[0133] A visual recognition module 201 is configured to perform visual recognition on the first surveillance video to obtain visual recognition information, wherein the visual recognition information is used to identify an area in the video frame corresponding to a preset surveillance target;
[0134] An information analysis module 202 is configured to analyze the visual recognition information to obtain a target change probability for each pixel in each video frame, wherein the target change probability indicates a probability that the corresponding pixel is one of a plurality of pixels constituting the monitored target in the corresponding video frame;
[0135] The macroblock adjustment module 203 is configured to adjust the sizes of the multiple initial macroblocks corresponding to each video frame according to the target change probability of each pixel in each video frame, thereby obtaining multiple adjusted macroblocks corresponding to each video frame, wherein the number of the multiple adjusted macroblocks corresponding to the same video frame is less than or equal to the number of the corresponding multiple initial macroblocks, and the target change probability is negatively correlated with the size of the adjusted macroblock covering the corresponding pixel.
[0136] The video compression module 204 is configured to perform video compression on the multiple video frames according to the multiple adjustment macroblocks corresponding to each of the video frames to obtain a second monitoring video.
[0137] Furthermore, the visual recognition module 201 includes:
[0138] a target detection submodule, configured to perform target detection on each of the plurality of video frames included in the first surveillance video, using the surveillance target as a detection target, to obtain a plurality of detection information, wherein the plurality of detection information corresponds one-to-one to the plurality of video frames, and the detection information is used to indicate n detection boxes in the corresponding video frames, where n is a non-negative integer;
[0139] a target tracking submodule, configured to track a target based on the plurality of detection information to obtain a plurality of tracking information, wherein the plurality of tracking information corresponds one-to-one to a plurality of monitoring moments corresponding to the first surveillance video, and the tracking information is configured to indicate m target frames in the video frame at the corresponding monitoring moments, wherein the target frames are the detection frames where the indicated target appears multiple times consecutively in the first surveillance video, and m is a non-negative integer;
[0140] an information acquisition submodule, configured to obtain a plurality of target change information based on the plurality of tracking information, wherein the plurality of target change information corresponds one-to-one to the plurality of monitoring moments, and the target change information includes: a position change sequence of a center point of each target frame in m target frames included in a video frame at the corresponding monitoring moment;
[0141] An information generation submodule is used to generate the visual recognition information according to the position change sequence of the center point of each target frame in the multiple target frames indicated by the multiple tracking information, wherein the area corresponding to the preset monitoring target includes the center point of the corresponding target frame, and the pixel points whose correlation between the position change sequence and the position change sequence of the center point of the corresponding target frame is greater than a correlation threshold.
[0142] Furthermore, the information generation submodule includes:
[0143] a feature extraction unit, configured to perform local feature extraction on each of the plurality of target frames indicated by the plurality of tracking information, to obtain a plurality of feature information corresponding one-to-one to the plurality of target frames, the feature information including a plurality of feature points associated with the corresponding target frames, the plurality of feature points being used to represent the outline of the target indicated by the corresponding target frame;
[0144] a feature correlation calculation unit, configured to perform correlation calculations on the plurality of position change sequences of the plurality of feature points included in each piece of feature information and the position change sequence of the corresponding center point, respectively, to obtain a plurality of feature-related information corresponding one-to-one to the plurality of feature information, the feature-related information including: a correlation between the position change sequence between each feature point and the corresponding center point among the plurality of feature points included in the corresponding feature information;
[0145] An identification information generation unit is used to generate the visual identification information based on the multiple feature-related information and the multiple feature information, wherein the area corresponding to the preset monitoring target includes the feature points whose correlation with the position change sequence between the corresponding center point is greater than the correlation threshold.
[0146] Furthermore, the information analysis module 202 includes:
[0147] a boundary acquisition unit, configured to obtain, based on the plurality of feature information, a plurality of boundary information corresponding one-to-one to the plurality of target frames, wherein the boundary information includes a plurality of boundary points corresponding one-to-one to the corresponding plurality of feature points, and the boundary points are intersection points of a straight line connecting the corresponding feature point and the corresponding center point and a boundary of the corresponding target frame;
[0148] a boundary correlation calculation unit, configured to perform correlation calculations on the plurality of position change sequences of the plurality of boundary points included in each of the plurality of boundary information and the position change sequences of the corresponding feature points, respectively, to obtain a plurality of boundary-related information corresponding one-to-one to the plurality of boundary information, the boundary-related information comprising: a correlation of the position change sequences between each of the plurality of boundary points included in the corresponding boundary information and the corresponding feature point;
[0149] an attention degree calculation unit, configured to perform calculations based on the multiple feature-related information and the multiple boundary-related information to determine multiple feature attention information, wherein the multiple feature attention information corresponds to the multiple target frames one-to-one, and the feature attention information includes multiple attention degrees of multiple feature points associated with the corresponding target frames, wherein the attention degrees are the ratios of the corresponding boundary-related values and the corresponding feature-related values, wherein the boundary-related values are used to represent the correlation of the position change sequence between the corresponding boundary points and the corresponding feature points, and the feature-related values are used to represent the correlation of the position change sequence between the corresponding feature points and the corresponding center points;
[0150] The probability analysis unit is used to analyze the plurality of feature attention information to obtain the target change probability of each pixel point in each of the video frames.
[0151] Furthermore, the probability analysis unit is specifically used to:
[0152] Analyzing the plurality of feature attention information to obtain a plurality of change probability values corresponding to each pixel point in each of the video frames, wherein the plurality of change probability values correspond one-to-one to the plurality of feature points indicated by the corresponding feature attention information, and the change probability values are jointly determined based on the corresponding attention degree and the corresponding point distance, where the point distance is the distance between the corresponding feature point and the corresponding pixel point;
[0153] For each pixel point in each of the video frames, the maximum value of the multiple change probability values corresponding to the pixel point is determined as the target change probability value of the pixel point to obtain the target change probability of each pixel point in each of the video frames.
[0154] Furthermore, the macroblock adjustment module 203 is specifically configured to:
[0155] According to the target change probability of each pixel point in each of the video frames, each initial macroblock corresponding to each of the video frames is iteratively divided into one or more adjusted macroblocks to obtain multiple adjusted macroblocks corresponding to each of the video frames, wherein the size of the initial macroblock is greater than or equal to the size of the adjusted macroblock iteratively divided from the initial macroblock, and when the size of the adjusted macroblock is greater than a preset minimum size and the size of the adjusted macroblock is smaller than the size of the corresponding initial macroblock, the average of multiple target change probabilities of multiple pixel points covered by the adjusted macroblock in the corresponding video frame is less than a preset probability threshold.
[0156] Furthermore, the feature extraction unit is specifically used to:
[0157] Merging the target frames indicated by the tracking information to obtain a plurality of merged frames corresponding to the target frames, wherein the merged frames are unions of the corresponding target frames and their associated frames, and the associated frames and the corresponding target frames indicate the same target;
[0158] Local features are extracted for each of the multiple merged frames to obtain multiple feature information corresponding to the multiple target frames.
[0159] It should be noted that the system provided in the above embodiment is merely an example of the division of the above functional modules. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the computer device can be divided into different functional modules to complete all or part of the functions described above. In addition, the above embodiment provides a high-efficiency surveillance video compression system for intelligent monitoring and an embodiment of a high-efficiency surveillance video compression method for intelligent monitoring, which are based on the same concept. The specific implementation process is detailed in the method embodiment and will not be repeated here.
[0160] The embodiment of the present invention also provides an electronic device. Figure 3 , the electronic device may include a processor 301, a memory 302, and a program 3021 stored in the memory 302 and executable on the processor 301.
[0161] When the program 3021 is executed by the processor 301, it can achieve Figure 1 Any steps in the corresponding method embodiments and achieving the same beneficial effects will not be repeated here.
[0162] Those skilled in the art will appreciate that all or part of the steps of implementing the above-described embodiment method may be accomplished through hardware associated with program instructions, and the program may be stored in a readable medium.
[0163] The embodiment of the present application further provides a readable storage medium, wherein the readable storage medium stores a computer program, and the computer program can realize the above-mentioned method when executed by a processor. Figure 1 Any step in the corresponding method embodiment can be achieved, and the same technical effects can be achieved, to avoid repetition, which will not be described here.
[0164] The computer readable storage medium of the embodiment of the present application can adopt any combination of one or more computer readable media. The computer readable medium can be a computer readable signal medium or a computer readable storage medium. The computer readable storage medium may, for example, be an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or apparatus, or any combination thereof. More specific examples (non-exhaustive list) of the computer readable storage medium include: an electrical connection having one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In this document, the computer readable storage medium can be any tangible medium containing or storing a program that can be used by or in connection with an instruction execution system, device or apparatus.
[0165] The computer readable signal medium can include a data signal propagating in a baseband or as part of a carrier wave propagating through a transmission medium, in which the computer readable program code is embodied. Such a propagating data signal can take many forms, including but not limited to electro-magnetic, optical, or any suitable combination thereof. The computer readable signal medium can also be any computer readable medium that is not a computer readable storage medium and that can communicate, propagate or transport program for use by or in connection with an instruction execution system, apparatus or device.
[0166] The program code contained on the storage medium can be transmitted in any suitable medium, including but not limited to wireless, wire, optical cable, RF, etc., or any suitable combination thereof.
[0167] Computer program code for performing the operations of the present invention may be written in one or more programming languages, or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, C++, and conventional procedural programming languages such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or terminal. In cases involving a remote computer, the remote computer may be connected to the user's computer via any type of network, such as a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0168] An embodiment of the present invention further provides a computer program product. When the computer program product is run on a computer, the computer is caused to execute the above-mentioned related steps to implement an efficient compression method for surveillance video for intelligent monitoring provided in the above embodiment.
[0169] It should be noted that the order in which the embodiments of the present invention are described above is for illustrative purposes only and does not necessarily represent the superiority or inferiority of the embodiments. The processes depicted in the accompanying drawings do not necessarily require the specific order or sequential order shown to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0170] The various embodiments in this specification are described in a progressive manner, and the same or similar parts between the various embodiments can be referred to each other. Each embodiment focuses on the differences from other embodiments.
Claims
1. An efficient compression method for surveillance video for intelligent monitoring, characterized in that: The method comprises: Performing visual recognition on the first surveillance video to obtain visual recognition information, wherein the visual recognition information is used to identify an area in the video frame corresponding to a preset surveillance target; Analyzing the visual recognition information to obtain a target change probability for each pixel in each video frame, wherein the target change probability indicates a probability that the corresponding pixel is one of a plurality of pixels constituting the monitored target in the corresponding video frame; adjusting the sizes of the multiple initial macroblocks corresponding to each video frame according to the target change probability of each pixel in each video frame to obtain multiple adjusted macroblocks corresponding to each video frame, wherein the number of the multiple adjusted macroblocks corresponding to the same video frame is less than or equal to the number of the corresponding multiple initial macroblocks, and the target change probability is negatively correlated with the size of the adjusted macroblock covering the corresponding pixel; Performing video compression on the multiple video frames according to the multiple adjusted macroblocks corresponding to each of the video frames to obtain a second surveillance video, wherein a compression rate of an area corresponding to a preset surveillance target in the video frame of the second surveillance video is less than a compression rate of a background area, and the background area is other areas in the video frame except the area corresponding to the preset surveillance target; wherein performing visual recognition on the first surveillance video to obtain visual recognition information includes: Taking the monitoring target as a detection target, performing target detection on each of the plurality of video frames included in the first monitoring video to obtain a plurality of detection information, wherein the plurality of detection information corresponds to the plurality of video frames in a one-to-one manner, and the detection information is used to indicate n detection boxes in the corresponding video frames, where n is a non-negative integer; Perform target tracking based on the multiple detection information to obtain multiple tracking information, wherein the multiple tracking information corresponds one-to-one to the multiple monitoring moments corresponding to the first surveillance video, and the tracking information is used to indicate m target frames in the video frame at the corresponding monitoring moments, where the target frames are the detection frames of the tracked target that appear multiple times consecutively in the first surveillance video, and m is a non-negative integer; Obtaining a plurality of target change information based on the plurality of tracking information, wherein the plurality of target change information corresponds one-to-one to the plurality of monitoring moments, and the target change information includes: a position change sequence of a center point of each target frame in m target frames included in a video frame at the corresponding monitoring moment; generating the visual recognition information based on a position change sequence of a center point of each target frame in the plurality of target frames indicated by the plurality of tracking information, wherein the region corresponding to the preset monitoring target includes the center point of the corresponding target frame and pixels for which a correlation between the position change sequence and the position change sequence of the center point of the corresponding target frame is greater than a correlation threshold; The step of generating the visual recognition information according to a position change sequence of a center point of each target frame in the plurality of target frames indicated by the plurality of tracking information includes: Performing local feature extraction on each of the plurality of target frames indicated by the plurality of tracking information to obtain a plurality of feature information corresponding one-to-one to the plurality of target frames, the feature information including a plurality of feature points associated with the corresponding target frames, the plurality of feature points being used to represent an outline of the target indicated by the corresponding target frame; Correlation calculations are performed on the position change sequences of the plurality of feature points included in each piece of feature information and the position change sequence of the corresponding center point, respectively, to obtain a plurality of feature-related information corresponding one-to-one to the plurality of feature information, the feature-related information including: correlation of the position change sequence between each feature point and the corresponding center point among the plurality of feature points included in the corresponding feature information; generating the visual recognition information based on the plurality of feature-related information and the plurality of feature information, wherein the area corresponding to the preset monitoring target includes the feature points whose correlation of the position change sequence with the corresponding center point is greater than the correlation threshold; The analyzing the visual recognition information to obtain the target change probability of each pixel in each video frame includes: Obtaining, based on the plurality of feature information, a plurality of boundary information corresponding one-to-one to the plurality of target frames, wherein the boundary information includes a plurality of boundary points corresponding one-to-one to the corresponding plurality of feature points, the boundary points being intersections of a straight line connecting the corresponding feature point and the corresponding center point and a boundary of the corresponding target frame; Correlation calculation is performed on each of the plurality of boundary information items, wherein the plurality of position change sequences of the plurality of boundary points included in each boundary information item and the position change sequences of the corresponding feature points are respectively performed to obtain a plurality of boundary-related information items corresponding one-to-one to the plurality of boundary information items, wherein the boundary-related information items include: correlation of the position change sequences between each boundary point and the corresponding feature point in the plurality of boundary points included in the corresponding boundary information item; Calculations are performed based on the multiple feature-related information and the multiple boundary-related information to determine multiple feature attention information, the multiple feature attention information and the multiple target frames having one-to-one correspondence, the feature attention information including multiple attention levels of the multiple feature points associated with the corresponding target frames, the attention levels being the ratios of the corresponding boundary-related values to the corresponding feature-related values, the boundary-related values being used to represent the correlation of the position change sequence between the corresponding boundary points and the corresponding feature points, and the feature-related values being used to represent the correlation of the position change sequence between the corresponding feature points and the corresponding center points; Analyze the plurality of feature attention information to obtain a target change probability for each pixel in each of the video frames; The step of analyzing the plurality of feature attention information to obtain a target change probability for each pixel in each of the video frames includes: Analyzing the plurality of feature attention information to obtain a plurality of change probability values corresponding to each pixel point in each of the video frames, wherein the plurality of change probability values correspond one-to-one to the plurality of feature points indicated by the corresponding feature attention information, and the change probability values are jointly determined based on the corresponding attention degree and the corresponding point distance, where the point distance is the distance between the corresponding feature point and the corresponding pixel point; For each pixel point in each of the video frames, the maximum value of the multiple change probability values corresponding to the pixel point is determined as the target change probability value of the pixel point to obtain the target change probability of each pixel point in each of the video frames.
2. The method for efficiently compressing surveillance video for intelligent monitoring according to claim 1, characterized in that: The step of adjusting the sizes of the multiple initial macroblocks corresponding to each video frame according to the target change probability of each pixel in each video frame to obtain the multiple adjusted macroblocks corresponding to each video frame includes: According to the target change probability of each pixel point in each of the video frames, each initial macroblock corresponding to each of the video frames is iteratively divided into one or more adjusted macroblocks to obtain multiple adjusted macroblocks corresponding to each of the video frames, wherein the size of the initial macroblock is greater than or equal to the size of the adjusted macroblock iteratively divided from the initial macroblock, and when the size of the adjusted macroblock is greater than a preset minimum size and the size of the adjusted macroblock is smaller than the size of the corresponding initial macroblock, the average of multiple target change probabilities of multiple pixel points covered by the compressed macroblock in the corresponding video frame is less than a preset probability threshold.
3. The method for efficiently compressing surveillance videos for intelligent monitoring according to claim 1, characterized in that: The extracting local features of the multiple target frames indicated by the multiple tracking information to obtain multiple feature information corresponding to the multiple target frames includes: Merging the target frames indicated by the tracking information to obtain a plurality of merged frames corresponding to the target frames, wherein the merged frames are unions of the corresponding target frames and their associated frames, and the associated frames and the corresponding target frames indicate the same target; Local features are extracted for each of the multiple merged frames to obtain multiple feature information corresponding to the multiple target frames.
4. An efficient surveillance video compression system for intelligent monitoring, characterized in that: The system comprises: a visual recognition module, configured to perform visual recognition on the first surveillance video to obtain visual recognition information, wherein the visual recognition information is used to identify an area in the video frame corresponding to a preset surveillance target; an information analysis module, configured to analyze the visual recognition information to obtain a target change probability for each pixel in each video frame, wherein the target change probability indicates a probability that the corresponding pixel is one of a plurality of pixels constituting the monitored target in the corresponding video frame; a macroblock adjustment module, configured to adjust the sizes of the multiple initial macroblocks corresponding to each video frame according to a target change probability of each pixel in each video frame, thereby obtaining multiple adjusted macroblocks corresponding to each video frame, wherein the number of the multiple adjusted macroblocks corresponding to the same video frame is less than or equal to the number of the corresponding multiple initial macroblocks, and the target change probability is negatively correlated with the size of the adjusted macroblock covering the corresponding pixel; a video compression module, configured to perform video compression on each of the plurality of video frames according to the plurality of adjusted macroblocks corresponding to each of the video frames to obtain a second surveillance video, wherein a compression rate of an area corresponding to a preset surveillance target in the video frame of the second surveillance video is less than a compression rate of a background area, the background area being an area other than the area corresponding to the preset surveillance target in the video frame; Wherein, the visual recognition module includes: a target detection submodule, configured to perform target detection on each of the plurality of video frames included in the first surveillance video, using the surveillance target as a detection target, to obtain a plurality of detection information, wherein the plurality of detection information corresponds one-to-one to the plurality of video frames, and the detection information is used to indicate n detection boxes in the corresponding video frames, where n is a non-negative integer; a target tracking submodule, configured to track a target based on the plurality of detection information to obtain a plurality of tracking information, wherein the plurality of tracking information corresponds one-to-one to a plurality of monitoring moments corresponding to the first surveillance video, and the tracking information is configured to indicate m target frames in the video frame at the corresponding monitoring moments, wherein the target frames are the detection frames where the indicated target appears multiple times consecutively in the first surveillance video, and m is a non-negative integer; an information acquisition submodule, configured to obtain a plurality of target change information based on the plurality of tracking information, wherein the plurality of target change information corresponds one-to-one to the plurality of monitoring moments, and the target change information includes: a position change sequence of a center point of each target frame in m target frames included in a video frame at the corresponding monitoring moment; an information generation submodule, configured to generate the visual recognition information based on a position change sequence of a center point of each target frame in the plurality of target frames indicated by the plurality of tracking information, wherein the area corresponding to the preset monitoring target includes the center point of the corresponding target frame and pixels for which the correlation between the position change sequence and the position change sequence of the center point of the corresponding target frame is greater than a correlation threshold; Wherein, the information generation submodule includes: a feature extraction unit, configured to perform local feature extraction on each of the plurality of target frames indicated by the plurality of tracking information, to obtain a plurality of feature information corresponding one-to-one to the plurality of target frames, the feature information including a plurality of feature points associated with the corresponding target frames, the plurality of feature points being used to represent the outline of the target indicated by the corresponding target frame; a feature correlation calculation unit, configured to perform correlation calculations on the plurality of position change sequences of the plurality of feature points included in each piece of feature information and the position change sequence of the corresponding center point, respectively, to obtain a plurality of feature-related information corresponding one-to-one to the plurality of feature information, the feature-related information including: a correlation between the position change sequence between each feature point and the corresponding center point among the plurality of feature points included in the corresponding feature information; an identification information generating unit, configured to generate the visual identification information based on the plurality of feature-related information and the plurality of feature information, wherein the area corresponding to the preset monitoring target includes the feature points whose correlation of the position change sequence with the corresponding center point is greater than the correlation threshold; Wherein, the information analysis module includes: a boundary acquisition unit, configured to obtain, based on the plurality of feature information, a plurality of boundary information corresponding one-to-one to the plurality of target frames, wherein the boundary information includes a plurality of boundary points corresponding one-to-one to the corresponding plurality of feature points, and the boundary points are intersection points of a straight line connecting the corresponding feature point and the corresponding center point and a boundary of the corresponding target frame; a boundary correlation calculation unit, configured to perform correlation calculations on the plurality of position change sequences of the plurality of boundary points included in each of the plurality of boundary information and the position change sequences of the corresponding feature points, respectively, to obtain a plurality of boundary-related information corresponding one-to-one to the plurality of boundary information, the boundary-related information comprising: a correlation of the position change sequences between each of the plurality of boundary points included in the corresponding boundary information and the corresponding feature point; an attention degree calculation unit, configured to perform calculations based on the multiple feature-related information and the multiple boundary-related information to determine multiple feature attention information, wherein the multiple feature attention information corresponds to the multiple target frames one-to-one, and the feature attention information includes multiple attention degrees of multiple feature points associated with the corresponding target frames, wherein the attention degrees are the ratios of the corresponding boundary-related values and the corresponding feature-related values, wherein the boundary-related values are used to represent the correlation of the position change sequence between the corresponding boundary points and the corresponding feature points, and the feature-related values are used to represent the correlation of the position change sequence between the corresponding feature points and the corresponding center points; A probability analysis unit, configured to analyze the plurality of feature attention information to obtain a target change probability for each pixel in each of the video frames; Wherein, the probability analysis unit is specifically used to: Analyzing the plurality of feature attention information to obtain a plurality of change probability values corresponding to each pixel point in each of the video frames, wherein the plurality of change probability values correspond one-to-one to the plurality of feature points indicated by the corresponding feature attention information, and the change probability values are jointly determined based on the corresponding attention degree and the corresponding point distance, where the point distance is the distance between the corresponding feature point and the corresponding pixel point; For each pixel point in each of the video frames, the maximum value of the multiple change probability values corresponding to the pixel point is determined as the target change probability value of the pixel point to obtain the target change probability of each pixel point in each of the video frames.
5. An electronic device, characterized in that: The method comprises a processor, a memory, and a computer program stored in the memory and executable on the processor, wherein when the computer program is executed by the processor, the method implements the steps of the efficient compression method for surveillance video for intelligent monitoring as described in any one of claims 1 to 3.
6. A readable storage medium, characterized in that: The readable storage medium stores a computer program, which, when executed by a processor, implements the steps of the efficient compression method for surveillance video for intelligent monitoring according to any one of claims 1 to 3.
Citation Information
Patent Citations
Method, apparatus and system for encoding and decoding video data
AU2016203314A1
Video information processing method based on video information processing model
CN119094812A