Mass video monitoring high-compression storage system and device based on artificial intelligence

By introducing a high-compression storage system based on artificial intelligence into the video surveillance system, dynamically adjusting encoding parameters and optimizing encoding tools, the problems of insufficient compression rate, high encoding complexity and poor video quality in massive monitoring video storage are solved, and efficient video data compression and storage are achieved.

CN120201194AInactive Publication Date: 2025-06-24ANHUI TONGSHUO INFORMATION TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510318297.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-18
Publication Date
2025-06-24
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The existing video surveillance system has problems such as insufficient compression rate, high encoding complexity and poor video quality in massive surveillance video storage, which is difficult to meet the needs of high-compression storage.

Method used

Adopting a video surveillance high-compression storage system based on artificial intelligence, through video data acquisition, data redundancy analysis, adaptive compression coding, quality evaluation and adjustment modules, we dynamically adjust the encoding parameters, optimize the encoding tools and processing flow, and improve the compression rate and video quality.

Benefits of technology

It significantly improves the compression rate, which is more than 10 times higher than traditional methods, saves storage space, and reduces encoding complexity, which is suitable for real-time encoding needs and ensures video quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120201194A_ABST
    Figure CN120201194A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of data compression and storage, and discloses a mass video monitoring high-compression storage system and device based on artificial intelligence, and the system comprises a video data acquisition module, a video data processing module, a video data processing module and a video data storage module, obtaining video frame data from the original video stream by using a video frame collector; the data redundancy analysis module is used for analyzing the correlation coefficient between the video frames and the data redundancy of the video frames; the self-adaptive compression coding module is used for self-adaptively adjusting compression coding parameters according to the redundancy; the quality evaluation and adjustment module is used for carrying out quantitative evaluation on the quality of the video subjected to self-adaptive compression coding, and optimizing the video quality by adjusting quantization parameters when the quality does not reach the standard; and the video output module outputs the compressed video file, the compression quality report, the compression efficiency statistics and the storage space saving report. According to the invention, the video quality is ensured, the high compression rate of the video is realized, and the storage space is saved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of data compression and storage technology, and more specifically, to a high-compression storage system and device for massive video surveillance based on artificial intelligence. Background Art

[0002] With the widespread application of video surveillance systems, surveillance video data has shown an explosive growth trend. Traditional video surveillance systems usually use general video coding standards such as H.264 and H.265 to compress and store surveillance videos. However, these general coding methods are mainly aimed at compressing ordinary multimedia videos. They lack targeted optimization design for the specific field of surveillance videos and are difficult to meet the high compression storage requirements of massive surveillance videos. Specifically, the following problems exist in the prior art: insufficient compression rate. The general video coding standard does not fully mine the data redundancy of typical features of surveillance videos such as static backgrounds and low activity, resulting in a low compression rate and unable to effectively alleviate the storage pressure brought by massive surveillance videos. High coding complexity. In order to pursue high compression performance, new generation coding standards such as H.265 have introduced a large number of coding tools with high computational complexity. However, for surveillance video coding with high real-time requirements, the excessively high coding complexity limits its practical application. The video quality is not ideal. The general coding method lacks the ability to understand and adaptively process the content of surveillance videos. It adopts a unified coding configuration, which makes it difficult to balance the compression efficiency and video quality of different surveillance scenes. It is easy to introduce visual artifacts, affecting the availability of surveillance videos.

[0003] Therefore, there is an urgent need for a high-compression storage system for surveillance videos that can fully utilize the data redundancy characteristics of surveillance videos, significantly improve the compression rate while ensuring video quality, and reduce encoding complexity to meet the storage needs of massive surveillance video data. Summary of the invention

[0004] The present invention provides a high-compression storage system for massive video surveillance based on artificial intelligence, comprising:

[0005] The video data acquisition module uses a video frame collector to obtain video frame data from the original video stream; and reads data frame by frame through a video processing function;

[0006] A data redundancy analysis module is used to analyze the correlation coefficient between video frames and the data redundancy of video frames, where the data redundancy includes time redundancy and space redundancy;

[0007] An adaptive compression coding module adaptively adjusts compression coding parameters according to redundancy; and compresses and codes video frames using the adjusted parameters;

[0008] A quality assessment and adjustment module quantifies and assesses the quality of the video after adaptive compression encoding, and optimizes the video quality by adjusting quantization parameters when the quality does not meet the standard;

[0009] A video output module outputs the compressed video file, compression quality report, compression efficiency statistics, and storage space savings report.

[0010] In a preferred embodiment, the steps of obtaining the original surveillance video stream include: extracting data frame by frame in chronological order from the video stream through a video frame collector to obtain the current frame and the previous frame, and caching them in the frame buffer.

[0011] In a preferred embodiment, the steps of analyzing the inter-frame correlation of the video include: calculating the pixel mean value of the current frame and the previous frame to obtain the inter-frame correlation coefficient.

[0012] In a preferred embodiment, the steps of analyzing the data redundancy include:

[0013] Estimating the conditional entropy based on the inter-frame correlation coefficient and calculating the temporal redundancy;

[0014] Dividing the current frame into sub-blocks, analyzing the relationship between the information entropy of the sub-blocks and the adjacent regions, and calculating the spatial redundancy.

[0015] In a preferred embodiment, the steps of adaptively adjusting the compression encoding parameters include:

[0016] Calculating the target bit rate according to the original video bit rate and the desired compression ratio;

[0017] Adjusting the quantization parameters based on the temporal redundancy and the spatial redundancy;

[0018] Generating the GOP structure according to the temporal redundancy.

[0019] In a preferred embodiment, the steps of compressing and encoding the video frames with the adjusted parameters include:

[0020] Selecting an encoder and setting the quantization parameters, GOP structure, and target bit rate parameters;

[0021] Performing encoding and calculating the actual output bit rate;

[0022] Adjusting the quantization parameters through a bit rate control loop iteration to make the actual bit rate close to the target bit rate.

[0023] In a preferred embodiment, the steps of evaluating the quality of the compressed video include:

[0024] Decoding the compressed video frames;

[0025] Calculate the peak signal-to-noise ratio (PSNR) and structural similarity (SSIM) metrics between the compressed frame and the original frame;

[0026] If the PSNR and SSIM do not reach the preset threshold, adjust the quantization parameters and return to the compression encoding step to execute again until the quality requirement is met or the maximum number of iterations is reached.

[0027] In a preferred embodiment, the steps of outputting the compressed video file and the related report include:

[0028] Output the compressed video file after quality evaluation and adjustment, and the storage space is reduced to less than 1 / 10 of the original;

[0029] Generate a compressed quality report including the PSNR and SSIM metrics;

[0030] Generate a compressed efficiency statistical report including the compression ratio and processing time;

[0031] Generate a storage space savings report reflecting the percentage of storage space saved.

[0032] A high-compression storage device for mass video surveillance based on artificial intelligence, including a memory and one or more processors. Executable code is stored in the memory. When the one or more processors execute the executable code, it is used to implement the modules in a high-compression storage system for mass video surveillance based on artificial intelligence as described above.

[0033] A storage medium stores non-transitory computer-readable instructions. When the non-transitory computer-readable instructions are executed by a computer, they can execute the modules in a high-compression storage system for mass video surveillance based on artificial intelligence as described above.

[0034] The beneficial effects of the present invention are as follows:

[0035] Significant improvement in compression ratio: Utilize artificial intelligence technology to adaptively analyze the spatio-temporal redundancy characteristics of surveillance videos, and dynamically adjust the encoding parameters accordingly, making the compression algorithm more suitable for the data characteristics of surveillance videos. The compression ratio is increased by more than 10 times compared with traditional general encoding methods, and storage space can be saved;

[0036] Reduction in encoding complexity: Optimize the encoding tools and processing flow according to the content characteristics of surveillance videos, reduce unnecessary encoding calculations, reduce the encoding complexity while increasing the compression ratio, and are more suitable for the real-time encoding requirements of surveillance videos. BRIEF DESCRIPTION OF THE DRAWINGS

[0037] Figure 1 It is a module diagram of a high-compression storage system for mass video surveillance based on artificial intelligence of the present invention;

[0038] Figure 2 These are examples of the parameters and results of each step of video processing in the present invention. Detailed implementation manners

[0039] The subject matter described herein will now be discussed with reference to example embodiments. It should be understood that discussing these embodiments is only to enable those skilled in the art to better understand and thus implement the subject matter described herein. Without departing from the scope of protection of the content of this specification, changes can be made to the functions and arrangements of the elements discussed. Each example can omit, substitute, or add various processes or components as needed. Additionally, the features described in some examples can also be combined in other examples.

[0040] At least one embodiment of the present invention discloses a high-compression storage system for mass video surveillance based on artificial intelligence, as Figure 1 shown, including:

[0041] A video data acquisition module that uses a video frame collector to obtain frame data from the original video stream and reads data from the video stream frame by frame in chronological order through a video processing function;

[0042] In one embodiment of the present invention, the calculation formula of the video processing function is as follows:

[0043] ;

[0044] Where represents the video frame obtained from the video stream at time ; represents the video processing function;

[0045] Perform cache processing on the current and previous frames:

[0046] ;

[0047] Wherein, , respectively represent the video frames obtained from the video stream at times and ;

[0048] To analyze the similarity between video frames, calculate the correlation between two adjacent frames and ;

[0049] In one embodiment of the present invention, the calculation formula for the similarity between video frames is as follows:

[0050] ;

[0051] Wherein, represents frame and frame The correlation coefficient between Representation frame Pixels in Representation frame Pixels in Representation frame The pixel mean of ; Representation frame The pixel mean of ; Indicates frame Each pixel in With frame The corresponding pixel in The sum of Indicates frame All pixels in The sum of Indicates frame All pixels in The sum of

[0052] Data redundancy analysis module, which calculates the temporal redundancy of video sequences and the spatial redundancy of video images;

[0053] In one embodiment of the present invention, temporal redundancy reflects the degree of information repetition of a video sequence in the temporal dimension; the calculation formula is as follows:

[0054] ;

[0055] in, Indicates time redundancy; Representation frame Information entropy of Indicates that in a known frame conditional frame The conditional entropy of

[0056] ;

[0057] in, represents conditional entropy; Representation frame Information entropy of represents the inter-frame correlation coefficient; represents the absolute value of the correlation coefficient;

[0058] Spatial redundancy focuses on the repetition of information in the internal space of the image; the calculation formula is as follows:

[0059] ;

[0060] in Indicates spatial redundancy; Representing images Located in sub-blocks of locations; Represents a sub-block adjacent areas; represents the information entropy of the sub-block; represents the conditional entropy of the sub-block under the condition of known adjacent regions;

[0061] Redundancy analysis report generation:

[0062] ;

[0063] in, Indicates redundant reporting; , They represent temporal redundancy and spatial redundancy respectively.

[0064] Adaptive compression coding module, which calculates the target bit rate, adjusts the quantization parameters and generates the GOP structure according to the redundancy analysis results of the video data;

[0065] In one embodiment of the present invention, the target bit rate calculation is to achieve the goal of reducing the surveillance video storage space to more than 1 / 10 of the original, and the calculation formula is as follows:

[0066] ;

[0067] in represents the target bit rate; Indicates the bit rate of the original video; Indicates the desired compression ratio, requiring ;

[0068] Quantization parameter adjustment is a key step in compression coding. To control the compression ratio and video quality; the quantization parameter adjustment calculation formula is as follows:

[0069] ;

[0070] in, Indicates the final quantization parameter used; represents the basic quantization parameter; Indicates the adjustment amount of the quantization parameter; Indicates time redundancy; Indicates spatial redundancy; It is the adjustment amount of the basic quantization parameter according to the temporal redundancy and spatial redundancy. The adjustment formula is implemented as: ,The higher the redundancy, the more appropriate the QP can be reduced to improve the compression ratio;

[0071] Generation of GOP Structure: The Group of Pictures (GOP) structure is an important organization method in video coding. It divides consecutive video frames into groups, and each group contains different types of frames (such as I-frames, P-frames, B-frames, etc.). This structure helps improve coding efficiency and random access performance; generating the GOP structure based on temporal redundancy:

[0072] ;

[0073] Among them, represents the generated Group of Pictures; represents the function for generating the Group of Pictures; represents the temporal redundancy;

[0074] The calculation formula for the GOP length is as follows:

[0075] ;

[0076] Among them, represents the length of the generated GOP structure; represents the maximum allowed GOP length; represents the temporal redundancy;

[0077] Encoder Selection: According to the characteristics of video data and compression requirements, the H.265 encoder is selected;

[0078] Encoding Parameter Setting: Set various parameters of the encoder:

[0079] Quantization Parameter: Use the .

[0080] GOP Structure: Use the .

[0081] Target Bitrate: Set to the .

[0082] Video Coding Execution:

[0083] ;

[0084] Actual Bitrate Calculation:

[0085] ;

[0086] Bitrate Control Loop:

[0087] When the actual bitrate is greater than the target bitrate , the quantization parameter is increased by 1, and then the updated quantization parameter is used and the GOP encodes the current video frame to obtain the encoded video frame , and then recalculates the actual bitrate of the encoded video frame ; repeat this process until the actual bitrate is not greater than the target bitrate ;

[0088] The quality evaluation and adjustment module quantitatively evaluates the quality of the video after adaptive compression encoding, and optimizes the video quality by adjusting quantization parameters when the quality does not meet the standard;

[0089] Decode the compressed frame:

[0090] ;

[0091] wherein represents the decoded image; represents the encoded video frame; represents the decoding function;

[0092] PSNR calculation:

[0093] The peak signal-to-noise ratio is a commonly used metric to measure the quality of a video, and its calculation formula is as follows:

[0094] ;

[0095] wherein represents the peak signal-to-noise ratio, with the unit of dB; represents the logarithm to the base 10; represents the mean squared error; represents the maximum possible value of the pixel value (255 for an 8-bit image);

[0096] ;

[0097] wherein represents the mean squared error; represents the width of the image (number of pixels); represents the height of the image (number of pixels); represents the pixel value of the original frame at the coordinate ; represents the pixel value of the decoded frame at the coordinate ; represents the double summation over all pixel positions;

[0098] Structural Similarity Index (SSIM) evaluation: The Structural Similarity Index evaluates the video quality from the perspective of image structure information; its calculation formula is as follows:

[0099] ;

[0100] Among them, represents the structural similarity index; represents the original image mean value; represents the decoded image mean value; represents the standard of the original image; represents the standard deviation of the decoded image; represents the covariance of the two images; respectively represent the first and second small constants to avoid a zero denominator; , where is the dynamic range of pixel values (usually 255), ; , where .

[0101] Quality constraint check:

[0102] ;

[0103] Among them, represents the quality constraint check, and its value is or ; represents the peak signal-to-noise ratio; represents the structural similarity index; represents the logical AND operation. When and only when both conditions are satisfied, the value of the entire expression is true;

[0104] Parameter adjustment when the quality does not meet the standard:

[0105] When the quality check fails ( ), the system executes the parameter adjustment process:

[0106] First, adjust the quantization parameter QP according to the current PSNR and SSIM values to obtain a new QP value;

[0107] Then, use the adjusted parameter to return to the compression encoding step and re-execute the encoding to obtain a new encoded frame ;

[0108] Next, decode the newly encoded frame to obtain ;

[0109] Subsequently, recalculate the PSNR and SSIM values based on the original frame and the decoded frame ; This process will be executed in a loop until the quality constraint conditions are met ( And ) or reach the preset maximum number of attempts.

[0110] Quality assessment report generation:

[0111] ;

[0112] Among them, represents the quality assessment report; represents the peak signal-to-noise ratio; represents the structural similarity index; represents the quality constraint check.

[0113] Output module, outputting the compressed video file, compression quality report, compression efficiency statistics, and storage space savings report;

[0114] Calculate the actual compression ratio:

[0115] ;

[0116] Among them, represents the actual compression ratio; represents the bit rate of the original video; represents the actual bit rate of the compressed video;

[0117] Store the compressed frames:

[0118] The storage rules of the storage system stipulate that the encoded video frames need to determine the corresponding storage location according to their specific identification information (such as time stamp, video stream number, etc.); this identification information helps to quickly locate and retrieve video frames in the storage system. According to this rule, is stored in the corresponding location;

[0119] Calculate the saved space:

[0120] ;

[0121] Among them, represents the percentage of the saved storage space; represents the actual compression ratio; represents the ratio of the saved space;

[0122] Compression efficiency report generation:

[0123] ;

[0124] Among them, represents the compression efficiency report; represents the actual compression ratio; Indicates the percentage of storage space saved; Indicates the processing time;

[0125] Storage report generation:

[0126] Obtain the storage location information of the encoded video frames from the storage system by querying the metadata index of the storage system in; The data index records the mapping relationship between the identification information of each video frame (such as timestamp, video stream number, etc.) and the actual storage location;

[0127] Determine the stored file size. For the encoded video frames , obtain its size by reading the attributes of the corresponding file in the storage system;

[0128] Organize the obtained relevant information such as storage location and file size into a report format;

[0129] Through the function Generate a storage report from the organized information ; The storage report contains key information such as storage location and file size, which is convenient for users to understand the storage situation of video frames in the storage system;

[0130] In an embodiment of the present invention, for the problems of tight storage space and low retrieval efficiency, a high-compression storage system for massive video surveillance based on artificial intelligence further includes:

[0131] A semantic extraction module that uses a video semantic extraction module to analyze the input video frames ;

[0132] The semantic extraction module performs feature extraction on the video frames based on a three-layer attention network to obtain feature vectors. Through the attention mechanism formula:

[0133] ;

[0134] Where represents the query vector; represents the key vector; represents the value vector; represents the transpose of the key vector ; represents the dimension of the key vector, with a value of 64; represents the square root of the key vector dimension, used to scale the dot product; represents the softmax function.

[0135] Then, these weights are used to identify key objects and events in the video; for example, three key objects including the people, cars, and traffic lights in the frame are identified, as well as the vehicle stop event; finally, a semantic feature vector is generated.

[0136] ;

[0137] Each feature is attached with confidence and location information; for example, the confidence of the people feature is 0.9, and the location is in the upper left area of the center of the frame; the confidence of the car feature is 0.95, located in the right lane of the frame; the confidence of the traffic light feature is 0.88, at the intersection in the upper part of the frame; the confidence of the vehicle stop event is 0.92, involving the car in the right lane of the frame.

[0138] The adaptive semantic encoding module calculates the saliency score for each feature based on the extracted semantic features.

[0139] Saliency score calculation function is defined as:

[0140] ;

[0141] Among them, 0.5, 0.3, and 0.2 are weight coefficients, representing the contribution ratio of each factor to the saliency score; these weight coefficients are optimized through a large number of video retrieval experiments and user experience tests to ensure the best compression effect while guaranteeing the retrieval quality. represents the confidence of feature , ranging from 0 to 1;

[0142] is the importance score calculated according to the feature location, ranging from 0 to 1, and the calculation formula is:

[0143] ;

[0144] Among them is the center coordinate of feature , is the center coordinate of the frame, is the maximum distance from the center to the edge of the frame;

[0145] is the importance score of the feature in the video scene, ranging from 0 to 1, determined based on a predefined scene-feature importance mapping table;

[0146] For the four features in this example, the calculation process is as follows:

[0147] For example, for the people feature :

[0148] Confidence

[0149] Position importance: The feature is located in the left - center area of the screen, with coordinates approximately , and the center of the screen is . Calculated as:

[0150] ;

[0151] Scene importance: In the cross - road scene, the character feature is relatively important.

[0152] The saliency score is calculated as follows:

[0153] ;

[0154] Saliency threshold is used to distinguish semantic significant regions and non - significant regions in video frames. This threshold is determined by analyzing a large amount of video data and combining the requirements for retaining important information in the actual application scenario. In this example, when is set to , it can better balance the retention of important information and the compression of data volume. That is, when the saliency score of the feature , the corresponding region is classified as a semantic significant region, such as the region containing people, cars, and vehicle - stop events; when , the corresponding region is a non - significant region, that is, other background regions;

[0155] For the encoding parameters, the high - quality encoding parameter and the low - quality encoding parameter are set based on the concept of quantization parameter in the H.265 encoding standard. The quantization parameter QP determines the quantization degree of data during the video encoding process. The smaller the QP value, the finer the quantization, the higher the encoding quality, but the bitrate will also increase; conversely, the larger the QP value, the coarser the quantization, the lower the encoding quality, and the bitrate will decrease. In this example, the high - quality encoding parameter is set for the significant region because the significant region contains important semantic information and requires a higher encoding quality to ensure the accurate transmission of information; the low - quality encoding parameter is set for the non - significant region. Since the non - significant region is mostly background information, appropriately reducing the encoding quality has little impact on the overall semantic understanding of the video but can effectively reduce the bitrate.

[0156] Region - quantization parameter mapping is a mapping table that associates each macro - block in the video frame with the corresponding quantization parameter. Its generation process assigns the corresponding quantization parameter according to the region type (semantic significant region or non - significant region) where each macro - block is located. That is, if the macro - block is located in the semantic significant region, then ; if it is in a non-significant area, assign Through this mapping, the H.265 encoder can use different quantization parameters for different regions for regional adaptive encoding, thereby effectively reducing the overall bit rate while ensuring the quality of important information. The final encoded frame ;

[0157] Semantic indexing building module, based on the extracted semantic features , generating a set of keywords including people, cars, traffic lights, and vehicle stops;

[0158] Assign a unique identifier to the current video frame, such as video "vid_20241001_103045_cam3";

[0159] Then update the inverted index table and add the identifier of the current video to the video list corresponding to each keyword. In other words, perform similar operations for each keyword. If a new keyword is encountered, create a new index entry;

[0160] Finally, the encoded semantic frame Store it in the storage system and record its storage location and metadata, thus completing the construction of the semantic index;

[0161] The semantically driven video retrieval module, when the user starts a query, such as the user query and constructed semantic index ;

[0162] First, the user's query keywords are processed and standardized into semantic tags used by the system: car and vehicle stop. Then, the semantic index is used to find videos containing these keywords:

[0163] ;

[0164] ;

[0165] Perform a union operation:

[0166]

[0167] ;

[0168] Sort the results by semantic relevance, which is calculated based on the number and weight of matched keywords, and finally return a sorted video list;

[0169] The sorted results may be:

[0170] vid_20241001_103045_cam3 (Relevance: 0.95) - Matched keywords: [Automobile, Vehicle stopped]

[0171] vid_20241001_104523_cam1 (Relevance: 0.82) - Matched keywords: [Vehicle stopped]

[0172] vid_20241001_103145_cam2 (Relevance: 0.78) - Matched keywords: [Automobile]

[0173] vid_20241001_105612_cam3 (Relevance: 0.75) - Matched keywords: [Automobile]

[0174] vid_20241001_110023_cam2 (Relevance: 0.72) - Matched keywords: [Vehicle stopped]

[0175] It can be seen that the videos containing both the keywords "automobile" and "vehicle stopped" are ranked at the top, followed by the ranking according to the relevance calculated based on the confidence and weight of each keyword;

[0176] In an embodiment of the present invention, semantic relevance is an indicator for measuring the degree of semantic matching between the retrieval result and the user's query, and the calculation formula is as follows:

[0177] ;

[0178] where is the set of keywords in the user's query; is the weight of the keyword based on the importance of this keyword in video semantic understanding; is the confidence of this keyword in the video.

[0179] In an embodiment of the present invention, an application example of the aforementioned high - compression storage system for massive video surveillance based on artificial intelligence is provided:

[0180] In a real - world scenario, assume that multiple cameras are deployed in a certain surveillance area. Taking the CAM001 camera as an example, video acquisition starts at 08:00:00 on October 1, 2024, to obtain the original video stream, with a resolution set to 1920×1080, a frame rate of 30fps, and a bit rate of 10000kbps.

[0181] Processing process: The video frame collector reads the video stream data frame by frame in chronological order. First, obtain the current frame , and then obtain the previous frame from the frame buffer , and calculate the average values of the pixel values of the two frames as and ;

[0182] Calculate the value for each pixel position, and finally obtain the inter-frame correlation coefficient , indicating a high similarity between the two frames.

[0183] Output: Current frame data , previous frame data , inter-frame correlation coefficient and frame buffer .

[0184] Data redundancy analysis:

[0185] Input: Current frame data output from the previous step , previous frame data , inter-frame correlation coefficient and frame buffer .

[0186] Processing procedure: First, calculate the information entropy of the current frame to obtain . Then, calculate the conditional entropy using the inter-frame correlation coefficient:

[0187] ;

[0188] Substitute these values into the temporal redundancy formula:

[0189] ;

[0190] However, considering the actual coding efficiency and algorithm stability, adjust this value to 0.6. Then divide the current frame into multiple 16×16 pixel sub-blocks, calculate the information entropy relationship between each sub-block and its adjacent region, and finally obtain the spatial redundancy .

[0191] Output: Frame data and redundancy analysis report ;

[0192] Adaptive compression coding:

[0193] Input: Frame data and redundancy analysis report output from the previous step ; ;

[0194] Processing procedure: First, determine the desired compression ratio , and then calculate the target bit rate based on the original video bit rate :

[0195] ;

[0196] Then, adjust the quantization parameter based on the redundancy analysis result and set the basic quantization parameter , adjustment amount calculation:

[0197] ;

[0198] Therefore, the quantization parameter ; rounded to 28;

[0199] Generate the GOP structure according to the temporal redundancy and set the maximum GOP length , then frames;

[0200] Output: frame data , quantization parameter , GOP structure (length 18 frames) and target bitrate .

[0201] Compression coding:

[0202] Input: frame data output from the previous step , quantization parameter , GOP structure (length 18 frames) and target bitrate .

[0203] Processing procedure: Select the H.265 encoder, set the encoding parameters QP = 28, GOP length 18 frames, and target bitrate 1000 kbps. Perform encoding operations on the frame data :

[0204] ;

[0205] After encoding is completed, calculate the actual bitrate , which is lower than the target bitrate 1000 kbps. Therefore, there is no need to execute the bitrate control loop. If the actual bitrate exceeds the target value, the QP value will be increased and re-encoding will be performed until the bitrate requirement is met.

[0206] Output: encoded frame data , original frame data , actual bitrate and the finally used quantization parameter .

[0207] Quality assessment and adjustment:

[0208] Input: encoded frame data output from the previous step , original frame data , actual bitrate and quantization parameter .

[0209] Processing process: First, decode the encoded frame:

[0210] ;

[0211] Then calculate the mean square error between the original frame and the decoded frame:

[0212] ;

[0213] Substitute the MSE into the PSNR formula:

[0214] ;

[0215] Next, calculate the SSIM value. First, calculate the means , , standard deviations , , and covariance of the original frame and the decoded frame. Set constants , , and substitute them into the SSIM formula to calculate . Verify the quality constraints: and , both meet the requirements, and the quality inspection passes.

[0216] Output: Encoded frame data , quality assessment report and actual bitrate .

[0217] Result output:

[0218] Input: Encoded frame data output from the previous step , quality assessment report

[0219] ;

[0220] and actual bitrate .

[0221] Processing process: First, calculate the actual compression ratio:

[0222] ;

[0223] Write the encoded frame data to the distributed storage system. The final compressed size of the entire video is 100MB (original size 1000MB). Calculate the percentage of storage space saved:

[0224] ;

[0225] Generate a compression efficiency report:

[0226] ;

[0227] Generate a storage report to record information such as storage location and file size.

[0228] Output the final product:

[0229] The compressed video file (100MB, 1 / 10 of the original size)

[0230] Compression quality report ;

[0231] Compression efficiency statistics:

[0232]

[0233] Storage space savings report (900MB of storage space saved);

[0234] Examples of parameters and results for each step of video processing are as Figure 2 shown.

[0235] The above describes the embodiments of the present invention. However, these embodiments are not limited to the above specific implementation manners. The above specific implementation manners are merely illustrative and not restrictive. Under the inspiration of this embodiment, those of ordinary skill in the art can also make more forms of equivalent embodiments, all of which fall within the protection scope of this embodiment.

Claims

1. A high-compression storage system for massive video surveillance based on artificial intelligence, characterized in that: include: The video data acquisition module uses a video frame collector to obtain video frame data from the original video stream; and reads data frame by frame through a video processing function; A data redundancy analysis module is used to analyze the correlation coefficient between video frames and the data redundancy of video frames, where the data redundancy includes time redundancy and space redundancy; An adaptive compression coding module adaptively adjusts compression coding parameters according to redundancy; and compresses and codes video frames using the adjusted parameters; The quality assessment and adjustment module quantitatively assesses the quality of the video after adaptive compression encoding, and optimizes the video quality by adjusting the quantization parameters when the quality does not meet the standard; Video output module, outputs compressed video files, compression quality reports, compression efficiency statistics and storage space saving reports.

2. According to the artificial intelligence-based massive video surveillance high-compression storage system of claim 1, it is characterized in that: The steps of obtaining the original monitoring video stream include: extracting data frame by frame from the video stream in chronological order through a video frame collector, obtaining the current frame and the previous frame, and caching them in a frame buffer.

3. The artificial intelligence-based massive video surveillance high-compression storage system according to claim 1 is characterized in that: The step of analyzing the correlation between video frames includes: calculating the mean value of pixels of the current frame and the previous frame to obtain the correlation coefficient between frames.

4. The artificial intelligence-based massive video surveillance high-compression storage system according to claim 1 is characterized in that: The steps to analyze data redundancy include: The conditional entropy is estimated based on the inter-frame correlation coefficient to calculate the temporal redundancy; The current frame is divided into sub-blocks, the information entropy relationship between the sub-blocks and adjacent areas is analyzed, and the spatial redundancy is calculated.

5. The artificial intelligence-based massive video surveillance high-compression storage system according to claim 1 is characterized in that: The steps of adaptively adjusting compression encoding parameters include: Calculate the target bit rate according to the original video bit rate and the expected compression ratio; Adjusting quantization parameters based on temporal redundancy and spatial redundancy; The GOP structure is generated based on temporal redundancy.

6. The artificial intelligence-based high-compression storage system for massive video surveillance according to claim 1, characterized in that: The steps of compressing and encoding the video frame using the adjusted parameters include: Select the encoder and set the quantization parameters, GOP structure, and target bit rate parameters; Perform encoding and calculate the actual output bit rate; The quantization parameters are adjusted iteratively through the bit rate control loop to make the actual bit rate close to the target bit rate.

7. The artificial intelligence-based high-compression storage system for massive video surveillance according to claim 1 is characterized in that: The steps to assess the quality of compressed video include: Decode compressed video frames; Calculate the peak signal-to-noise ratio (PSNR) and structural similarity (SSIM) index between the compressed frame and the original frame; If the PSNR and SSIM do not reach the preset threshold, the quantization parameters are adjusted and the compression encoding step is returned to be re-executed until the quality requirement is met or the maximum number of iterations is reached.

8. The artificial intelligence-based massive video surveillance high-compression storage system according to claim 1, characterized in that: The steps to output compressed video files and related reports include: Output compressed video files after quality assessment and adjustment, and the storage space is reduced to less than 1 / 10 of the original; Generate compression quality report including PSNR and SSIM indicators; Generate compression efficiency statistics report including compression ratio and processing time; Generates a storage space savings report showing the storage space savings percentage.

9. A high-compression storage device for massive video surveillance based on artificial intelligence, characterized in that: It includes a memory and one or more processors, wherein the memory stores executable code, and when the one or more processors execute the executable code, they are used to implement a module in a massive video surveillance high-compression storage system based on artificial intelligence as described in any one of claims 1-8.

10. A storage medium storing non-transitory computer-readable instructions, characterized in that: When the non-temporary computer-readable instructions are executed by a computer, they can execute the module in the artificial intelligence-based massive video surveillance high-compression storage system described in any one of claims 1-8.