Video compression storage method and system for low activity
By introducing human body recognition model and cache queue technology into the video surveillance system, the problem of wasted storage space in video surveillance areas is solved, and efficient video compression storage is achieved.
Patent Information
- Application Number
- CN202510755339.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-06
- Publication Date
- 2025-08-15
Smart Images

Figure CN120499332A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of video processing, and in particular to a video compression storage method and system for low-activity videos. Background Art
[0002] Video surveillance systems are security technology systems that use cameras and related technologies to monitor, record, and manage specific areas in real time. They are widely used in fields such as power, industry, and healthcare. Typically, a video surveillance system built on closed-circuit television systems consists of cameras, transmission lines, routing equipment, storage devices, and back-end monitoring equipment. With the advancement of computer technology, monitoring servers are often added, combined with artificial intelligence programs to detect video content.
[0003] For example, patent document CN201410553488.0 discloses a railway station integrated monitoring system and design method, which includes a monitoring center, a communication network, an environment and equipment monitoring system BAS, and an electric power monitoring system PSCADA or an integrated information system. The monitoring center is connected to the electromechanical equipment monitoring system BAS and the electric power monitoring system PSCADA via the communication network, and the monitoring center is connected to the integrated information system via a router. The system has comprehensive integrated monitoring functions. Under normal operating conditions, it can complete electromechanical equipment and environmental monitoring, circuit automation, energy-saving control and management, and can exchange information with various subsystems and realize linkage. When a disaster occurs, it can automatically enter the disaster operation mode, integrate on-site alarm information, coordinate the work of various related systems, complete the alarm linkage function, prevent or reduce the harm caused by the disaster, and provide a decision-making basis for disaster prevention and mitigation emergency dispatch and command.
[0004] For another example, patent document CN201410017212.0 discloses a security monitoring system, in which a monitoring server is connected to a local area network; a control room computer is connected to the local area network and is also connected to the Internet; an alarm host is connected to a 110 alarm linkage, a fire alarm linkage, and various alarms installed at monitoring points; a video monitoring host is connected to a video camera and access control system installed at each monitoring point; a gas monitoring host is connected to a gas sensor installed at each gas monitoring point; an electric power monitoring host is connected to an electric power monitor installed at each electric power monitoring point; and a user terminal and a remote control terminal are connected to the Internet. This invented system can monitor various environments and can sound an alarm in the event of an abnormality, thereby jointly maintaining the safety of a factory or home environment. It can be used to monitor personnel, as well as the environment such as electricity and gas, to jointly maintain a safe and reliable environment. It can also be accessed through the network for easy monitoring and management.
[0005] However, during the actual implementation process, the inventors found that when this type of security monitoring system is used in relatively fixed scenes, such as computer rooms, most of its time is spent on recording fixed scenes, and most of the video content it collects is fixed scenes. As the video clarity increases, it will occupy a large amount of storage space. Summary of the Invention
[0006] In view of the above problems existing in the prior art, a video compression storage method for low activity is provided;
[0007] On the other hand, a system for implementing the video compression storage method is also provided.
[0008] The specific technical solutions are as follows:
[0009] A video compression storage method for low activity, comprising:
[0010] Step S1: obtaining a real-time video stream in a corresponding scene, performing human body recognition on video frames in the real-time video stream, and determining whether a human body appears in the current video frame;
[0011] If yes, the real-time video stream is used as the video stream to be stored, and then the process goes to step S3; if no, the process goes to step S2:
[0012] Step S2: adding the video frame to a cache queue, and compressing the video frame according to the cache queue to form the video stream to be stored;
[0013] In the cache queue, for the multiple consecutive video frames, one of the video frames is used to replace the other video frames;
[0014] Step S3: compress and store the video stream to be stored.
[0015] On the other hand, in the step S1, the video frame is detected using a human body recognition model;
[0016] The human body recognition model includes:
[0017] An input layer, wherein the input layer obtains the video frame as an input image;
[0018] A feature extraction layer, the feature extraction layer is connected to the input layer;
[0019] The feature extraction layer extracts image features from the input image through a series of convolutional networks and pooling layers to form a feature map;
[0020] An SSD detection layer, the SSD detection layer is connected to the feature extraction layer;
[0021] The SSD detection layer uses multiple convolution kernels of different scales to predict the target of the feature map and generate a label box and a label category;
[0022] A non-maximum suppression layer, the non-maximum suppression layer being connected to the SSD detection layer;
[0023] The non-maximum suppression layer filters the repeated annotation boxes and retains the best annotation box;
[0024] an output layer, the output layer being connected to the non-maximum suppression layer;
[0025] The output layer generates a detection result corresponding to whether a human body exists according to the annotation box.
[0026] On the other hand, the step S1 includes:
[0027] Step S11: obtaining the real-time video stream, and obtaining the video frame from the real-time video stream;
[0028] Step S12: inputting the video frame into the human body recognition model to obtain the detection result;
[0029] Step S13: determining whether a human body appears in the current video frame according to the detection result;
[0030] If yes, take the real-time video stream as the video stream to be stored, and then go to step S3;
[0031] If not, go to step S2.
[0032] On the other hand, the step S2 includes:
[0033] Step S21: receiving the video frame and determining whether the buffer queue has reached a preset length;
[0034] If yes, go to step S22;
[0035] If not, go to step S23
[0036] When the video frame is received for the first time, creating the cache queue based on the video frame;
[0037] Step S22: creating a new cache queue based on the currently received video frame, and clearing the original cache queue, and then turning to step S23;
[0038] Step S23: adding the video frame to the cache queue, and using the first frame of the video frame in the cache queue to add it to the video stream to be stored.
[0039] On the other hand, in step S3, compression processing is performed on the static video frames in the video stream to be stored.
[0040] On the other hand, step S3 includes:
[0041] Step S31: caching the video stream to be stored, and determining a static video frame among the video frames to be stored;
[0042] Step S32: Recording a static time period for the static video frame and performing interception.
[0043] A video compression storage system for low activity, used to implement the above-mentioned video compression storage method;
[0044] The video compression storage system comprises:
[0045] A human body recognition module, which obtains a real-time video stream in a corresponding scene, performs human body recognition on video frames in the real-time video stream, and determines whether a human body appears in the current video frame;
[0046] a cache compression module, the cache compression module being connected to the human body recognition module;
[0047] The cache compression module adds the video frame to a cache queue when a human body is detected, and compresses the video frame according to the cache queue to form the video stream to be stored;
[0048] In the cache queue, for the multiple consecutive video frames, one of the video frames is used to replace the other video frames;
[0049] a storage module, the storage module being connected to the cache compression module;
[0050] The storage module compresses and stores the video stream to be stored.
[0051] On the other hand, the human body recognition module includes:
[0052] A video acquisition module, wherein the video acquisition module acquires the real-time video stream and obtains the video frame from the real-time video stream;
[0053] A detection module, the detection module is connected to the video acquisition module;
[0054] The detection module inputs the video frame into the human body recognition model to obtain the detection result;
[0055] a judgment module, the judgment module being connected to the detection module;
[0056] The judgment module judges whether a human body appears in the current video frame according to the detection result.
[0057] On the other hand, the cache compression module includes:
[0058] a queue determination module, the queue determination module receiving the video frame and determining whether the cache queue has reached a preset length;
[0059] When the video frame is received for the first time, creating the cache queue based on the video frame;
[0060] A queue clearing module, the queue clearing module is connected to the queue judging module;
[0061] The queue clearing module creates a new cache queue based on the currently received video frame and clears the original cache queue, and then turns to step S23;
[0062] An adding module, the adding module is connected to the queue clearing module;
[0063] The adding module adds the video frame to the cache queue, and uses the first frame of the video frame in the cache queue to add it to the video stream to be stored.
[0064] On the other hand, the storage module includes:
[0065] a cache module, wherein the cache module caches the video stream to be stored and determines a static video frame among the video frames to be stored;
[0066] An interception module, the interception module being connected to the cache module;
[0067] The interception module records the static time period of the static video frame and intercepts it.
[0068] The above technical solution has the following advantages or beneficial effects:
[0069] In order to address the problem that video surveillance systems in existing technologies have a large amount of blank content when monitoring low-activity areas, in this solution, before storing the video stream, it is pre-determined whether there is a human body in the video. If no human body appears, the video frame is added to a cache queue and replaced with a static video frame. When stored in the back end, the static video frame can be efficiently compressed based on the compression algorithm, thereby reducing the space requirement. BRIEF DESCRIPTION OF THE DRAWINGS
[0070] The embodiments of the present invention will be described more fully with reference to the accompanying drawings, which are provided for illustration and description only and are not intended to limit the scope of the present invention.
[0071] Figure 1 is an overall schematic diagram of an embodiment of the present invention;
[0072] Figure 2 Schematic diagram of a human body recognition model in an embodiment of the present invention;
[0073] Figure 3 This is a schematic diagram of step S1 in an embodiment of the present invention;
[0074] Figure 4 This is a schematic diagram of step S2 in an embodiment of the present invention;
[0075] Figure 5 This is a schematic diagram of step S3 in an embodiment of the present invention;
[0076] Figure 6 A schematic diagram of a system in an embodiment of the present invention;
[0077] Figure 7 Schematic diagram of a human body recognition module in an embodiment of the present invention;
[0078] Figure 8 This is a schematic diagram of a cache compression module in an embodiment of the present invention;
[0079] Figure 9 Schematic diagram of a storage module in an embodiment of the present invention. DETAILED DESCRIPTION
[0080] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making any creative efforts shall fall within the scope of protection of the present invention.
[0081] It should be noted that, in the absence of conflict, the embodiments of the present invention and the features in the embodiments may be combined with each other.
[0082] The present invention will be further described below with reference to the accompanying drawings and specific embodiments, but they are not intended to limit the present invention.
[0083] The present invention comprises:
[0084] A video compression storage method for low activity, such as Figure 1 Shown, including:
[0085] Step S1: obtaining a real-time video stream in a corresponding scene, performing human body recognition on video frames in the real-time video stream, and determining whether a human body appears in the current video frame;
[0086] If yes, the real-time video stream is used as the video stream to be stored, and then the process goes to step S3;
[0087] If not, go to step S2:
[0088] Step S2: adding the video frame to the cache queue, and compressing it according to the cache queue to form a video stream to be stored;
[0089] In the cache queue, for multiple consecutive video frames, one of the video frames is used to replace the other video frames;
[0090] Step S3: compress and store the video stream to be stored.
[0091] Specifically, in view of the problem that the video surveillance system in the existing technology has a large amount of blank content when monitoring low-activity areas, in this solution, before storing the video stream, it is judged in advance whether there is a human body in the video. If no human body appears, the video frame is added to a cache queue and replaced with a static video frame. When stored in the back end, the static video frame can be efficiently compressed based on the compression algorithm, thereby reducing the space requirement.
[0092] In the actual implementation process, the above technical solution is configured as a software embodiment in a video surveillance system. The video surveillance system is mainly composed of cameras, edge nodes, and storage servers. Multiple cameras are connected to the edge nodes, which act as hubs and perform pre-processing on the multiple video streams input by the cameras, including identifying and independently compressing static video frames, and then transmitting them back to the back-end storage server for storage. The storage server has pre-configured video segmentation and compression storage software according to corresponding requirements, which can segment and store the input video streams and compress them through related compression codes. However, when multiple video streams are input, the compression calculation process on the storage server side is relatively complex, and the segmentation and compression processes are carried out simultaneously, which results in a large load.
[0093] In one embodiment, in step S1, a human body recognition model is used to detect the video frame;
[0094] like Figure 2 As shown, the human body recognition model includes:
[0095] Input layer A1, input layer A1 obtains the video frame as input image;
[0096] Feature extraction layer A2, feature extraction layer A2 is connected to input layer A1;
[0097] The feature extraction layer A2 extracts image features from the input image through a series of convolutional networks and pooling layers to form a feature map;
[0098] SSD detection layer A3, SSD detection layer A3 is connected to feature extraction layer A2;
[0099] The SSD detection layer A3 uses multiple convolution kernels of different scales to predict the target and generate the annotation box and annotation category.
[0100] Non-maximum suppression layer A4, which is connected to the SSD detection layer A3;
[0101] The non-maximum suppression layer A4 filters the repeated annotation boxes and retains the best annotation box;
[0102] Output layer A5, output layer A5 is connected to non-maximum suppression layer A4;
[0103] The output layer A5 generates a detection result corresponding to whether a human body exists based on the annotation box.
[0104] Specifically, in order to achieve a better detection effect on the human body, in this embodiment, the above-mentioned human body recognition model is introduced for recognition.
[0105] Specifically, for the input real-time video stream, it is first split frame by frame, and the obtained video frames are input into the input layer as input images.
[0106] Subsequently, feature extraction layer A2 extracts image features from the input image through a series of convolutional networks and pooling layers to form a feature map. Feature extraction layer A2 comprises a series of convolutional layers, each of which extracts features from the image using a convolution kernel. These features are then fed into the pooling layer to flatten the features for easier feature extraction in the next layer.
[0107] After a series of convolutional network extractions, a feature map is generated. This feature map is then fed into the SSD detection layer A3, which uses multiple convolution kernels of varying scales to predict objects and generate bounding boxes and labeled categories. This primarily includes information such as the predicted object location, size, category, and confidence level from the feature map.
[0108] On this basis, the non-maximum suppression layer A4 filters out duplicate annotation boxes and retains the best ones. Specifically, during object detection, overlapping annotation boxes may appear. These are determined based on the set of pixel positions enclosed within the annotation box. When overlapping annotation boxes appear, they are identified based on their confidence levels, and those with low confidence levels are removed.
[0109] Finally, the output layer A5 generates a detection result corresponding to whether a human body exists based on the annotation box.
[0110] In actual implementation, the ResNet50-SSD model pre-trained on the VOC dataset can be selected as the initial model. This model combines the advantages of ResNet50 and SSD (Single Shot MultiBox Detector), performs well in object detection tasks, and has powerful feature extraction capabilities.
[0111] Perform transfer learning on the selected ResNet50-SSD model, freezing the weights of the base network and fine-tuning only the parameters of the SSD detection head. This can accelerate model convergence and improve training results.
[0112] During the training process, the model is trained on a server equipped with an NVIDIA GPU, using the SGD (Stochastic Gradient Descent) optimizer and setting appropriate learning rates and decay strategies to accelerate model convergence and improve training efficiency.
[0113] At the same time, the model is evaluated using the test set, and the model's accuracy, recall rate, and average precision are recorded to ensure that the model maintains high accuracy while meeting real-time processing speed requirements.
[0114] After training is complete, save the model in a format suitable for deep learning inference frameworks, such as ONNX format, for subsequent deployment and integration.
[0115] In one embodiment, Figure 3 As shown, step S1 includes:
[0116] Step S11: obtaining a real-time video stream, and obtaining a video frame from the real-time video stream;
[0117] Step S12: inputting the video frame into the human body recognition model to obtain the detection result;
[0118] Step S13: judging whether a human body appears in the current video frame according to the detection result;
[0119] If yes, the real-time video stream is used as the video stream to be stored, and then the process goes to step S3;
[0120] If not, go to step S2.
[0121] Specifically, based on the aforementioned human recognition model, this embodiment can obtain a real-time video stream, extract video frames from the real-time video stream, and input the video frames one by one into the human recognition model, thereby obtaining the output of the human recognition model corresponding to the video frame, that is, the detection result. This detection result can be simplified at the output layer of the model to detect whether a human body exists, which can be achieved through a classifier.
[0122] Finally, the external software program selects the subsequent judgment branch based on the detection results output by the model, that is, whether additional compression logic needs to be called.
[0123] In one embodiment, Figure 4 As shown, step S2 includes:
[0124] Step S21: receiving a video frame and determining whether the buffer queue has reached a preset length;
[0125] If yes, go to step S22;
[0126] If not, go to step S23
[0127] When a video frame is received for the first time, a cache queue is created based on the video frame;
[0128] Step S22: creating a new cache queue based on the currently received video frame, clearing the original cache queue, and then turning to step S23;
[0129] Step S23: adding a video frame to the cache queue, and adding the first video frame in the cache queue to the video stream to be stored.
[0130] Specifically, to achieve a better compression effect, in this embodiment, when a video frame is received for the first time, a cache queue is created according to the video frame.
[0131] Typically, because video streams have a constant frame rate, there are two ways to implement cache queues: frame-based and timer-based. The two methods can be used as needed. For example, if the video stream is 30 frames per second, each cache queue can be configured with a 5-minute timer, or it can count the frames and reset every 9,000 frames.
[0132] For subsequently received video frames, it is first determined whether the cache queue has reached a preset length. When the preset length is reached, a new cache queue is created based on the currently received video frame, and the original cache queue is cleared.
[0133] If the preset length is not reached, the video frame is added to the buffer queue, and the buffer queue is controlled to form a corresponding output.
[0134] Generally speaking, the cache queue stores the first video frame as a replacement video frame and adds it to the video stream to be stored. Starting from the second frame, the cache queue deletes frames in a cyclic manner using a smaller buffer. For example, a 10-frame cycle is used. When the cycle length is met, the video frame at the beginning of the cycle is deleted in a first-in, first-out order, thus achieving a shorter queue length. However, this deletion process does not affect the frame counting or timing results of the cache queue.
[0135] The cache queue also includes a continuity determination module that uses the timestamps of the incoming video frames to determine whether the input to the cache queue is continuous. Upon receiving a video frame and determining the length of the cache queue, this determination module first determines, based on the timestamp, whether it is continuous with the last video frame in the cache queue. If so, the process proceeds to determine whether the cache queue has reached a preset length and branches thereafter. If not, the process considers the video frame received for the first time and proceeds to step S22 to clear the existing cache queue.
[0136] This processing branch can ensure that the compressed video frames have the same content, avoiding the use of the original video frames for compression after the human detection results appear and disappear, thereby avoiding the problem that in actual scenarios, the impact of the user entering the scene cannot be effectively recorded.
[0137] In one embodiment, in step S3, compression processing is performed on static video frames in the video stream to be stored.
[0138] Specifically, after the pre-processing steps, in this embodiment, the storage server can perform compressed storage based on the video segmentation and compression steps in the prior art. Because existing compression technology can calculate a high correlation between adjacent static frames, compression can be achieved by predicting and encoding the differences between frames. This method enables pre-processing of videos at the pre-processed edge node, further improving the compression efficiency of the back-end storage server.
[0139] In one embodiment, Figure 5 As shown, step S3 includes:
[0140] Step S31: caching the video stream to be stored and determining a static video frame among the video frames to be stored;
[0141] Step S32: Recording a static time period for a static video frame and performing interception.
[0142] Specifically, to achieve better compression, in this embodiment, the storage server first caches the video stream to be stored, establishing a video buffer of a certain length. Static frames are then identified within the video frames to be stored. For static frames, the server reads the timestamp to record the time period during which static video occurs, for example, 10 minutes or 5 minutes. Only one image frame is recorded during this time period, thereby reducing storage space usage.
[0143] Since static frames have been identified and replaced on the front-end edge node, the subsequent matching of storage servers can avoid the problem of static frame matching failure caused by factors such as image noise, which further reduces storage efficiency.
[0144] A video compression storage system for low activity, used to implement the above-mentioned video compression storage method;
[0145] like Figure 6 As shown, the video compression storage system includes:
[0146] Human body recognition module 1, which obtains the real-time video stream in the corresponding scene, performs human body recognition on the video frames in the real-time video stream, and determines whether a human body appears in the current video frame;
[0147] Cache compression module 2, cache compression module 2 is connected to human body recognition module 1;
[0148] The cache compression module 2 adds the video frame to the cache queue when a human body is detected, and compresses the video frame according to the cache queue to form a video stream to be stored;
[0149] In the cache queue, for multiple consecutive video frames, one of the video frames is used to replace the other video frames;
[0150] Storage module 3, storage module 3 is connected to cache compression module 2;
[0151] The storage module 3 compresses and stores the video stream to be stored.
[0152] Specifically, in view of the problem that the video surveillance system in the existing technology has a large amount of blank content when monitoring low-activity areas, in this solution, before storing the video stream, it is judged in advance whether there is a human body in the video. If no human body appears, the video frame is added to a cache queue and replaced with a static video frame. When stored in the back end, the static video frame can be efficiently compressed based on the compression algorithm, thereby reducing the space requirement.
[0153] In one embodiment, Figure 7 As shown, the human body recognition module 1 includes:
[0154] The video acquisition module 11 acquires a real-time video stream and obtains video frames from the real-time video stream;
[0155] Detection module 12, detection module 12 is connected to video acquisition module 11;
[0156] The detection module 12 inputs the video frame into the human body recognition model to obtain the detection result;
[0157] A judgment module 13, the judgment module 13 is connected to the detection module 12;
[0158] The judgment module 13 judges whether a human body appears in the current video frame according to the detection result.
[0159] Specifically, based on the aforementioned human recognition model, in this embodiment, the video acquisition module 11 can acquire a real-time video stream and extract video frames from the real-time video stream. The detection module 12 inputs the video frames one by one into the human recognition model, thereby obtaining the output of the human recognition model corresponding to the video frame, i.e., the detection result. This detection result can be simplified at the output layer of the model to detect whether a human body is present, which can be achieved through a classifier.
[0160] Finally, the judgment module 13 selects a subsequent judgment branch according to the detection result output by the model, that is, whether additional compression logic needs to be called.
[0161] In one embodiment, Figure 8 As shown, the cache compression module 2 includes:
[0162] The queue determination module 21 receives the video frame and determines whether the buffer queue has reached a preset length;
[0163] When a video frame is received for the first time, a cache queue is created based on the video frame;
[0164] The queue clearing module 22 is connected to the queue judging module 21;
[0165] The queue clearing module 22 creates a new cache queue based on the currently received video frame and clears the original cache queue, and then turns to step S23;
[0166] Add module 23, add module 23 and connect queue clearing module 22;
[0167] The adding module 23 adds a video frame to the cache queue, and uses the first video frame in the cache queue to add it to the video stream to be stored.
[0168] Specifically, to achieve a better compression effect, in this embodiment, when a video frame is received for the first time, a cache queue is created according to the video frame.
[0169] Typically, because video streams have a constant frame rate, there are two ways to implement cache queues: frame-based and timer-based. The two methods can be used as needed. For example, if the video stream is 30 frames per second, each cache queue can be configured with a 5-minute timer, or it can count the frames and reset every 9,000 frames.
[0170] For subsequently received video frames, the queue determination module 21 first determines whether the cache queue has reached a preset length. When the preset length is reached, the queue clearing module 22 creates a new cache queue based on the currently received video frame and clears the original cache queue.
[0171] If the preset length is not reached, the adding module 23 adds the video frame to the buffer queue and controls the buffer queue to form a corresponding output.
[0172] Generally speaking, the cache queue stores the first video frame as a replacement video frame and adds it to the video stream to be stored. Starting from the second frame, the cache queue deletes frames in a cyclic manner using a smaller buffer. For example, a 10-frame cycle is used. When the cycle length is met, the video frame at the beginning of the cycle is deleted in a first-in, first-out order, thus achieving a shorter queue length. However, this deletion process does not affect the frame counting or timing results of the cache queue.
[0173] In one embodiment, Figure 9 As shown, the storage module 3 includes:
[0174] The cache module 31 caches the video stream to be stored and determines the static video frame in the video frames to be stored;
[0175] An interception module 32 connected to the cache module 31;
[0176] The interception module 32 records the static time period for the static video frame and intercepts it.
[0177] Specifically, to achieve better compression, in this embodiment, the cache module 31 on the storage server first caches the video stream to be stored, establishing a video cache of a certain length. It then identifies static video frames within the video frames to be stored. For static video frames, the interception module 32 reads the timestamp to record the time period during which static video occurs, for example, 10 minutes or 5 minutes. Only one image frame is recorded during this time period, thereby reducing storage space usage.
[0178] Since static frames have been identified and replaced on the front-end edge node, the subsequent matching of storage servers can avoid the problem of static frame matching failure caused by factors such as image noise, which further reduces storage efficiency.
[0179] The above are only preferred embodiments of the present invention and do not limit the implementation mode and protection scope of the present invention. For those skilled in the art, it should be aware that all solutions obtained by equivalent substitutions and obvious changes made using the description and illustrations of the present invention should be included in the protection scope of the present invention.
Claims
1. A video compression storage method for low activity, characterized in that: include: Step S1: obtaining a real-time video stream in a corresponding scene, performing human body recognition on video frames in the real-time video stream, and determining whether a human body appears in the current video frame; If yes, take the real-time video stream as the video stream to be stored, and then go to step S3; If not, go to step S2; Step S2: adding the video frame to a cache queue, and compressing the video frame according to the cache queue to form the video stream to be stored; In the cache queue, for the multiple consecutive video frames, one of the video frames is used to replace the other video frames; Step S3: compress and store the video stream to be stored.
2. The video compression storage method according to claim 1, wherein: In the step S1, the video frame is detected using a human body recognition model; The human body recognition model includes: An input layer, wherein the input layer obtains the video frame as an input image; A feature extraction layer, the feature extraction layer is connected to the input layer; The feature extraction layer extracts image features from the input image through a series of convolutional networks and pooling layers to form a feature map; An SSD detection layer, the SSD detection layer is connected to the feature extraction layer; The SSD detection layer uses multiple convolution kernels of different scales to predict the target of the feature map and generate a label box and a label category; A non-maximum suppression layer, the non-maximum suppression layer being connected to the SSD detection layer; The non-maximum suppression layer filters the repeated annotation boxes and retains the best annotation box; an output layer, the output layer being connected to the non-maximum suppression layer; The output layer generates a detection result corresponding to whether a human body exists according to the annotation box.
3. The video compression storage method according to claim 2, wherein: The step S1 comprises: Step S11: obtaining the real-time video stream, and obtaining the video frame from the real-time video stream; Step S12: inputting the video frame into the human body recognition model to obtain the detection result; Step S13: determining whether a human body appears in the current video frame according to the detection result; If yes, take the real-time video stream as the video stream to be stored, and then go to step S3; If not, go to step S2.
4. The video compression storage method according to claim 1, wherein: The step S2 comprises: Step S21: receiving the video frame and determining whether the buffer queue has reached a preset length; If yes, go to step S22; If not, go to step S23 When the video frame is received for the first time, creating the cache queue based on the video frame; Step S22: creating a new cache queue based on the currently received video frame, and clearing the original cache queue, and then turning to step S23; Step S23: adding the video frame to the cache queue, and using the first frame of the video frame in the cache queue to add it to the video stream to be stored.
5. The video compression storage method according to claim 1, wherein: In step S3, compression processing is performed on the static video frames in the video stream to be stored.
6. The video compression storage method according to claim 1, wherein: The step S3 comprises: Step S31: caching the video stream to be stored, and determining a static video frame among the video frames to be stored; Step S32: Recording a static time period for the static video frame and performing interception.
7. A video compression storage system for low activity, characterized in that: Used to implement the video compression storage method according to any one of claims 1 to 6; The video compression storage system comprises: A human body recognition module, which obtains a real-time video stream in a corresponding scene, performs human body recognition on video frames in the real-time video stream, and determines whether a human body appears in the current video frame; a cache compression module, the cache compression module being connected to the human body recognition module; The cache compression module adds the video frame to a cache queue when a human body is detected, and compresses the video frame according to the cache queue to form the video stream to be stored; In the cache queue, for the multiple consecutive video frames, one of the video frames is used to replace the other video frames; a storage module, the storage module being connected to the cache compression module; The storage module compresses and stores the video stream to be stored.
8. The video compression storage system according to claim 2, wherein: The human body recognition module includes: A video acquisition module, wherein the video acquisition module acquires the real-time video stream and obtains the video frame from the real-time video stream; A detection module, the detection module is connected to the video acquisition module; The detection module inputs the video frame into the human body recognition model to obtain the detection result; a judgment module, the judgment module being connected to the detection module; The judgment module judges whether a human body appears in the current video frame according to the detection result.
9. The video compression storage system according to claim 1, wherein: The cache compression module includes: a queue determination module, the queue determination module receiving the video frame and determining whether the cache queue has reached a preset length; When the video frame is received for the first time, creating the cache queue based on the video frame; A queue clearing module, the queue clearing module is connected to the queue judging module; The queue clearing module creates a new cache queue based on the currently received video frame and clears the original cache queue, and then turns to step S23; An adding module, the adding module is connected to the queue clearing module; The adding module adds the video frame to the cache queue, and uses the first frame of the video frame in the cache queue to add it to the video stream to be stored.
10. The video compression storage system according to claim 1, wherein: The storage module includes: a cache module, wherein the cache module caches the video stream to be stored and determines a static video frame among the video frames to be stored; An interception module, the interception module being connected to the cache module; The interception module records the static time period of the static video frame and intercepts it.
Citation Information
Patent Citations
A security monitoring system
CN103745579B
Railway station comprehensive monitoring system and design method thereof
CN104260763A