Event-based video recording storage method, device, electronic equipment and readable medium
The event-based video storage method addresses issues of false triggers and resource inefficiency by spatio-temporal noise reduction and frame reconstruction, resulting in efficient and high-quality video storage.
Patent Information
- Application Number
- CN202411196800.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-08-29
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2044-08-29
AI Technical Summary
In the prior art, high infrared monitoring error-touch rate leads to wasted storage space for redundant video recordings, complex neural network models lead to large overhead of computing power resources, and sparse event information leads to low frame rate and poor quality of video recordings.
The event information collection is obtained through the event sensor, performs time-space noise reduction processing, constructs initial video frames, and performs frame reconstruction and video encoding, and generates a complete video recording and stores.
It reduces storage of redundant video frames, reduces computing resource consumption, improves the frame rate and quality of video recordings, and optimizes the utilization of storage space.
Smart Images

Figure CN119135818B_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present disclosure relate to the field of computer technology, and more particularly to an event-based video recording and storage method, apparatus, electronic device, and readable medium. Background Art
[0002] Using a camera device to perform dynamic monitoring on a target area, recording videos of detected dynamic events, and storing the recorded videos are common technical means in the security field. Currently, when using a monitoring device to record and store video, the commonly adopted method is as follows: detecting dynamic events through a monitoring device with infrared monitoring function, then recording videos for the dynamic events and directly storing the obtained recorded videos.
[0003] However, when collecting and storing video in the above manner, the following technical problems often exist:
[0004] First, the false touch rate of infrared monitoring is relatively high (such as the heat flow generated by an outdoor air conditioner), and the recorded videos collected in the case of false touch are often redundant videos, that is, they do not contain dynamic events, thus causing waste of storage space.
[0005] Second, denoising the event information set through a neural network model often has complex requirements for the network structure of the neural network model, resulting in a large consumption of computing power resources.
[0006] Third, when the collected event information is relatively sparse, the number of frames of the constructed video is small, resulting in a low frame rate of the video and poor quality of the reconstructed video.
[0007] The above information disclosed in this background art section is only used to enhance the understanding of the background of the inventive concept, and thus, it may include information that does not form the prior art known to those of ordinary skill in the art in this country. Summary of the Invention
[0008] This content part of the present disclosure is used to briefly introduce concepts, which will be described in detail in the following detailed implementation part. This content part of the present disclosure is not intended to identify the key features or essential features of the claimed technical solution, nor is it intended to limit the scope of the claimed technical solution.
[0009] Some embodiments of the present disclosure propose an event-based video recording and storage method, apparatus, electronic device, and readable medium to solve one or more of the technical problems mentioned in the above background art section.
[0010] In a first aspect, some embodiments of the present disclosure provide an event-based video recording and storage method, which includes: obtaining an event information set, where the event information in the event information set is collected by a monitoring device including an event sensor for a target monitoring area, the event information in the event information set includes pixel position, acquisition time, and a brightness identifier indicating the brightness change situation, and the event information set corresponds to an acquisition duration; performing spatio-temporal noise reduction processing on each piece of event information included in the event information set to obtain a sequence of denoised event information groups; constructing each initial video recording frame according to the sequence of denoised event information groups to obtain a sequence of initial video recording frames; performing frame reconstruction processing on the sequence of initial video recording frames to obtain a sequence of video recording frames, where the number of video recording frames included in the sequence of video recording frames is greater than the number of initial video recording frames included in the sequence of initial video recording frames; performing video encoding processing on the sequence of initial video recording frames to generate a video recording for the target monitoring area, where the video duration of the video recording is equal to the acquisition duration; and performing storage processing on the video recording.
[0011] In a second aspect, some embodiments of the present disclosure provide an event-based video recording and storage device, which includes: an obtaining unit configured to obtain an event information set, where the event information in the event information set is collected by a monitoring device including an event sensor for a target monitoring area, the event information in the event information set includes pixel position, acquisition time, and a brightness identifier indicating the brightness change situation, and the event information set corresponds to an acquisition duration; a spatio-temporal noise reduction unit configured to perform spatio-temporal noise reduction processing on each piece of event information included in the event information set to obtain a sequence of denoised event information groups; a construction unit configured to construct each initial video recording frame according to the sequence of denoised event information groups to obtain a sequence of initial video recording frames; a frame reconstruction unit configured to perform frame reconstruction processing on the sequence of initial video recording frames to obtain a sequence of video recording frames, where the number of video recording frames included in the sequence of video recording frames is greater than the number of initial video recording frames included in the sequence of initial video recording frames; a video encoding unit configured to perform video encoding processing on the sequence of initial video recording frames to generate a video recording for the target monitoring area, where the video duration of the video recording is equal to the acquisition duration; and a storage unit configured to perform storage processing on the video recording.
[0012] In a third aspect, some embodiments of the present disclosure provide an electronic device, which includes: one or more processors; a storage device storing one or more programs thereon, and when the one or more programs are executed by the one or more processors, the one or more processors are caused to implement the method described in any implementation manner of the first aspect.
[0013] In a fourth aspect, some embodiments of the present disclosure provide a computer-readable medium storing a computer program, wherein when the program is executed by a processor, the method described in any implementation of the first aspect above is implemented.
[0014] The above embodiments of the present disclosure have the following beneficial effects: The event-based video recording and storage method according to some embodiments of the present disclosure can reduce the storage of redundant video frames, thereby reducing the waste of storage space. Specifically, the reason for the waste of storage space is that the false touch rate of infrared monitoring is relatively high (such as the hot air flow generated by an air conditioner outdoor unit, etc.). In the case of false touch, the video recorded is often redundant video, that is, it does not contain dynamic events, thus causing waste of storage space. Based on this, the event-based video recording and storage method according to some embodiments of the present disclosure first obtains a set of event information. Among them, the event information in the above set of event information is collected by a monitoring device including an event sensor for a target monitoring area. The event information in the above set of event information includes pixel position, acquisition time, and a brightness identifier representing the brightness change situation. The above set of event information corresponds to an acquisition duration. Thus, through the event sensor, when a brightness change event greater than a preset brightness threshold is detected in the monitoring area, event information can be collected, thereby reducing redundant video recording caused by false triggering. Then, spatio-temporal noise reduction processing is performed on each piece of event information included in the above set of event information to obtain a sequence of noise-reduced event information groups. Thus, the noise information in the set of event information can be reduced, thereby reducing the additional computing power overhead caused by processing noise event information. After that, according to the above sequence of noise-reduced event information groups, each initial video recording frame is constructed to obtain a sequence of initial video recording frames. Thus, discrete pieces of event information can be converted into continuous video frames, so that the generated video recording frames contain less redundant information. Next, frame reconstruction processing is performed on the above sequence of initial video recording frames to obtain a sequence of video recording frames. Among them, the number of video recording frames included in the above sequence of video recording frames is greater than the number of initial video recording frames included in the above sequence of initial video recording frames. Thus, by frame reconstruction processing, the number of video frames is increased to fill the time gaps between dynamic events, enhancing the continuity and smoothness of the video stream. Secondly, video encoding processing is performed on the above sequence of initial video recording frames to generate a video recording for the above target monitoring area. Among them, the video duration of the above video recording is equal to the above acquisition duration. Thus, video encoding of the sequence of video recording frames can obtain a complete video recording. Finally, storage processing is performed on the above video recording. Because a method for constructing and reconstructing video frames based on event sensors is adopted, the acquisition of redundant video recording can be reduced, and a video recording containing less redundant background information can be generated, thereby reducing the waste of storage space. Description of the Drawings
[0015] In combination with the accompanying drawings and with reference to the following specific embodiments, the above and other features, advantages and aspects of the embodiments of the present disclosure will become more apparent. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic and the elements and elements are not necessarily drawn to scale.
[0016] Figure 1 is a flowchart of some embodiments of an event-based video recording and storage method according to the present disclosure;
[0017] Figure 2 is a schematic structural diagram of some embodiments of an event-based video recording and storage device according to the present disclosure;
[0018] Figure 3 is a schematic structural diagram of an electronic device suitable for implementing some embodiments of the present disclosure. Specific Embodiments
[0019] The embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although some embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. On the contrary, these embodiments are provided to more thoroughly and completely understand the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are only for exemplary purposes and are not used to limit the protection scope of the present disclosure.
[0020] In addition, it should be noted that, for the sake of convenience of description, only parts related to the relevant invention are shown in the drawings. Without conflict, the embodiments in the present disclosure and the features in the embodiments can be combined with each other.
[0021] It should be noted that the concepts such as "first" and "second" mentioned in the present disclosure are only used to distinguish different devices, modules or units, and are not used to limit the order or interdependence relationship of the functions performed by these devices, modules or units.
[0022] It should be noted that the modifications of "one" and "multiple" mentioned in the present disclosure are illustrative rather than restrictive. Those skilled in the art should understand that, unless otherwise clearly specified in the context, it should be understood as "one or more".
[0023] The names of the messages or information exchanged between multiple devices in the embodiments of the present disclosure are only for illustrative purposes and are not used to limit the scope of these messages or information.
[0024] The present disclosure will be described in detail below with reference to the drawings and in combination with the embodiments.
[0025] Figure 1Flow 100 of some embodiments of the event-based video recording and storage method according to the present disclosure is shown. The event-based video recording and storage method includes the following steps:
[0026] Step 101, obtain an event information set.
[0027] In some embodiments, the execution subject (such as a computing device) of the event-based video recording and storage method may obtain the event information set through a wired connection method or a wireless connection method. Among them, the above execution subject may be a server. The event information in the above event information set is collected by an event sensor included in the monitoring device for a target monitoring area. The above monitoring device may be a monitoring camera. The above event sensor may be a DVS sensor (Dynamic Vision Sensor). The event information in the above event information set may characterize a pixel in the above target monitoring area whose brightness change exceeds a preset brightness change threshold. The above event information may include a pixel position, a collection time, and a brightness identifier characterizing the brightness change situation. The above pixel position may be the position coordinates of the corresponding pixel. The above collection time may be the time stamp for collecting the corresponding event information. As an example, when the above brightness identifier is "1", it may characterize that the brightness value of the corresponding pixel at the corresponding collection time is greater than the brightness value at the previous collection time. When the above brightness identifier is "-1", it may characterize that the brightness value of the corresponding pixel at the corresponding collection time is less than the brightness value at the previous collection time. The above event information set corresponds to a collection duration. The above collection duration may be the duration consumed for collecting the above event information set.
[0028] It should be noted that the above wireless connection method may include, but is not limited to, 3G / 4G / 5G connections, WiFi connections, Bluetooth connections, WiMAX connections, Zigbee connections, UWB (ultra wideband) connections, and other currently known or future-developed wireless connection methods.
[0029] Step 102, perform spatio-temporal noise reduction processing on each event information included in the event information set to obtain a sequence of denoised event information groups.
[0030] In some embodiments, the above execution subject may perform spatio-temporal noise reduction processing on each event information included in the above event information set to obtain a sequence of denoised event information groups.
[0031] In some optional implementation manners of some embodiments, the above execution subject may perform spatio-temporal noise reduction processing on each event information included in the above event information set through the following steps to obtain a sequence of denoised event information groups:
[0032] First step, according to the collection time included in the event information in the above event information set, perform a partitioning process on the above event information set to obtain each event information group. In practice, the above execution entity may partition each event information with the same collection time included in the above event information set into the same event information group to obtain each event information group.
[0033] Second step, perform a time sequence sorting process on each of the above event information groups to obtain an event information group sequence. In practice, the above execution entity may sort each of the above event information groups in the order of time represented by the collection time corresponding to each event information group, from the earliest to the latest, to obtain an event information group sequence.
[0034] Third step, select an event information from the first event information group in the above event information group sequence as signal event information and insert it into a pre-constructed signal event information queue to update the signal event information queue. Among them, the above signal event information queue may be a queue used to temporarily store event information determined to be non-noise. In practice, the above execution entity may randomly select an event information from the first event information group in the above event information group sequence as signal event information and insert it into a pre-constructed signal event information queue.
[0035] Fourth step, insert the selected event information into a pre-constructed noise event information queue as noise event information to update the noise event information queue. Among them, the length of the above signal event information queue is equal to the length of the above noise event information queue. The above noise event information queue may be a queue used to temporarily store event information determined to be noise. Here, no specific limitation is made on the queue lengths of the above signal event information queue and the above noise event information queue.
[0036] Fifth step, according to the updated signal event information queue and the updated noise event information queue, perform spatial noise reduction processing on the above event information group sequence to obtain a noise-reduced event information group sequence.
[0037] In some optional implementation manners of some embodiments, the above execution entity may perform spatial noise reduction processing on the above event information group sequence according to the updated signal event information queue and the updated noise event information queue through the following steps to obtain a noise-reduced event information group sequence:
[0038] First step, for each event information group in the above event information group sequence, perform the following spatial noise reduction steps:
[0039] First sub-step, determine the event information group as the event information group to be noise-reduced.
[0040] Second sub-step, based on the event information group to be noise-reduced, perform the following event noise reduction steps:
[0041] Sub-step 1: Select a noise reduction target event information from the group of noise reduction target event information. In practice, the above-mentioned execution entity can randomly select a noise reduction target event information from the group of noise reduction target event information.
[0042] Sub-step 2: Determine the pixel position distances between the pixel positions included in the selected noise reduction target event information and the pixel positions included in each signal event information in the signal event information queue, to obtain respective pixel position distances. As an example, the determined pixel position distance can be the Manhattan distance.
[0043] Sub-step 3: In response to determining that the respective pixel position distances meet the preset neighborhood quantity condition, perform the following steps:
[0044] Sub-step 1: In response to determining that the signal event information queue is full, delete the target signal event information from the head of the signal event information queue, and shift each signal event information in the signal event information queue to update the signal event information queue. Wherein, the target signal event information is the signal event information that was first inserted into the signal event information queue. The above-mentioned preset neighborhood quantity condition is that the quantity of pixel position distances greater than or equal to the preset pixel distance threshold among the determined respective pixel position distances is greater than or equal to the preset quantity threshold. In practice, in response to determining that the signal event information queue is full, the above-mentioned execution entity can delete the target signal event information from the head of the signal event information queue, and shift each signal event information in the signal event information queue one position towards the head to update the signal event information queue.
[0045] Sub-step 2: Insert the selected noise reduction target event information as signal event information into the tail of the signal event information queue.
[0046] Sub-step 4: In response to determining that the above-mentioned respective pixel position distances do not meet the preset neighborhood quantity condition, perform the following steps:
[0047] Sub-step 1: In response to determining that the noise event information queue is full, delete the target noise event information from the noise event information queue, and shift each noise event information in the noise event information queue to update the noise event information queue. Wherein, the above-mentioned target noise event information is the noise event information that was first inserted into the noise event information queue. In practice, the above-mentioned execution entity, in response to determining that the noise event information queue is full, deletes the target noise event information from the noise event information queue, and shifts each noise event information in the noise event information queue one position towards the head to update the noise event information queue.
[0048] Sub-step 2: Insert the selected noise reduction target event information as noise event information into the noise event information queue.
[0049] Sub-step three: Delete the event information that is the same as the selected event information to be denoised from the event information group sequence to update the event information group sequence.
[0050] Sub-step five: Delete the selected event information to be denoised from the event information group to be denoised to update the event information group to be denoised.
[0051] Sub-step six: In response to determining that the event information group to be denoised is not empty, execute the above event denoising steps again.
[0052] Second step: Determine the updated event information group sequence as the denoised event information group sequence.
[0053] The relevant content of the above embodiments is an inventive point of the embodiments of the present disclosure, which can solve Technical Problem 2: "When denoising an event information set through a neural network model, there are often complex requirements for the network structure of the neural network model, resulting in a large consumption of computing power resources." The factors that lead to a large consumption of computing power resources are usually as follows: When denoising an event information set through a neural network model, there are often complex requirements for the network structure of the neural network model. If the above factors are solved, the effect of reducing the consumption of computing power resources can be achieved. To achieve this effect, the present disclosure takes into account the spatio-temporal characteristics of event information, and denoises the event information set based on the pixel position and acquisition time included in each event information. Since usually noise event information is often isolated from signal event information (far in space), the event information set can be denoised by constructing a queue, thereby avoiding the complex calculations brought by using a neural network model, and further reducing the consumption of computing power resources.
[0054] Step 103: According to the denoised event information group sequence, construct each video frame image of the video recording to obtain a video frame sequence of the video recording.
[0055] In some embodiments, the above execution subject may construct each video frame image of the video recording according to the above denoised event information group sequence to obtain a video frame sequence of the video recording.
[0056] In some optional implementation manners of some embodiments, the above execution subject may construct each initial video frame of the video recording according to the above denoised event information group sequence to obtain an initial video frame sequence through the following steps:
[0057] First step: perform clustering processing on each denoised event information in the above-mentioned denoised event information group sequence to obtain a set of denoised event information clusters. In practice, the above-mentioned execution entity can perform clustering processing on each denoised event information in the above-mentioned denoised event information group sequence through a preset clustering algorithm to obtain a set of denoised event information clusters. Thus, each denoised event information with similar spatio-temporal characteristics can be divided into the same cluster. As an example, the above-mentioned preset clustering algorithm can be the DBSCAN (Density-Based Spatial Clustering of Applications with Noise) algorithm.
[0058] Second step: determine the time window size corresponding to each denoised event information cluster in the above-mentioned set of denoised event information clusters. In practice, for each denoised event information cluster in the above-mentioned set of denoised event information clusters, first step: the above-mentioned execution entity can sort the collection times included in each denoised event information in the above-mentioned denoised event information cluster to obtain a collection time sequence. Second step: the above-mentioned execution entity can determine the time difference between the time represented by the first collection time and the time represented by the last collection time in the above-mentioned collection time sequence as the time window size.
[0059] Third step: determine the average value of the determined time window sizes as the target time window size. In practice, the above-mentioned execution entity can determine the average value of the determined time window sizes as the target time window size.
[0060] Fourth step: perform window partitioning processing on the above-mentioned denoised event information group sequence according to the above-mentioned target time window size to update the denoised event information group sequence. In practice, the above-mentioned execution entity can merge the denoised event information belonging to the same target time window size in the above-mentioned denoised event information group sequence according to the above-mentioned target time window size to update the denoised event information group sequence. For example, if the above-mentioned target time window size is 1s, the above-mentioned execution entity can merge the denoised event information groups in the denoised event information group sequence that are in the same time window.
[0061] Fifth step: generate an initial recorded video frame sequence according to the updated denoised event information group sequence and a pre-trained video frame construction model.
[0062] In some optional implementation manners of some embodiments, the above-mentioned execution entity can generate an initial recorded video frame sequence according to the updated denoised event information group sequence and a pre-trained video frame construction model through the following steps:
[0063] Step 1, for each denoised event information group in the updated denoised event information group sequence, perform the following video frame construction steps:
[0064] First sub-step, sort the above-mentioned denoised event information group to obtain a denoised event information sequence. In practice, the above-mentioned execution entity can sort each denoised time information included in the denoised event information group according to the chronological order represented by each acquisition time to obtain a denoised event information sequence.
[0065] Second sub-step, input the above-mentioned denoised event information sequence into the above-mentioned video frame construction model to generate an initial recorded video frame. Among them, the above-mentioned video frame construction model can be a temporal neural network model with the denoised event information sequence as the input and the initial video frame as the output. The above-mentioned video frame construction model can be a neural network model improved based on the U-Net model. The above-mentioned video frame construction model can include a first convolutional network, a first two-dimensional convolutional layer, a first convolutional long short-term memory network (ConvLSTM), a second two-dimensional convolutional layer, a second convolutional long short-term memory network, a residual block, a first upsampling convolutional layer, a second upsampling convolutional layer, and an image prediction layer. There is a skip connection between the output of the first convolutional layer and the input of the image prediction layer. There is a skip connection between the output of the first convolutional long short-term memory network and the second two-dimensional convolutional layer. There is a skip connection between the input and output of the residual block. The above-mentioned first convolutional network can be a convolutional neural network. The convolutional kernel sizes included in the above-mentioned first two-dimensional convolutional layer and the second two-dimensional convolutional layer can be 5. The convolutional kernels of the above-mentioned first convolutional long short-term memory network and the second convolutional long short-term memory network can be 3. The number of channels of the above-mentioned first convolutional long short-term memory network can be 3, and the convolutional kernel of the residual block can be 3. The above-mentioned first upsampling convolutional layer can include a bilinear interpolation upsampling layer and a convolutional layer with a convolutional kernel size of 5. The network structure of the second upsampling convolutional layer is the same as that of the first upsampling convolutional layer. Each of the above-mentioned convolutional layers uses a ReLU activation function and batch normalization processing. The above-mentioned image prediction layer can be a depth convolutional layer. The convolutional kernel size of the above-mentioned depth convolutional layer is 1, and a sigmoid activation function is used.
[0066] Step 2, determine the generated initial recorded video frames as an initial recorded video frame sequence. In practice, the above-mentioned execution entity can determine the generated initial recorded video frames as an initial recorded video frame sequence according to the generation order.
[0067] Step 104, perform frame reconstruction processing on the initial recorded video frame sequence to obtain a recorded video frame sequence.
[0068] In some embodiments, the above-mentioned execution entity may perform frame reconstruction processing on the above-mentioned initial video frame sequence of the video recording to obtain a video frame sequence of the video recording. Wherein, the number of video frames included in the above-mentioned video frame sequence of the video recording is greater than the number of initial video frames included in the above-mentioned initial video frame sequence of the video recording. In practice, the above-mentioned execution entity may perform frame reconstruction processing on the above-mentioned initial video frame sequence of the video recording through a video frame insertion algorithm to obtain a video frame sequence of the video recording. As an example, the above-mentioned video frame insertion algorithm may be, but is not limited to, the DAIN (Depth-Aware Video Frame Interpolation) algorithm, the DeepVO algorithm, and the FlowNet algorithm based on optical flow.
[0069] In some optional implementation manners of some embodiments, the above-mentioned execution entity may perform frame reconstruction processing on the above-mentioned initial video frame sequence of the video recording through the following steps to obtain a video frame sequence of the video recording:
[0070] In the first step, determine the target number of video frames according to the preset video frame rate and the above-mentioned acquisition duration. In practice, the above-mentioned execution entity may determine the product of the above-mentioned preset video frame rate and the above-mentioned acquisition duration as the target number of video frames.
[0071] In the second step, determine the total number of reconstructed video frames as the difference between the above-mentioned target number of video frames and the number of initial video frames in the above-mentioned initial video frame sequence of the video recording.
[0072] In the third step, use the ratio of the above-mentioned total number of reconstructed video frames to the above-mentioned acquisition duration as the number of unit reconstructed video frames.
[0073] In the fourth step, update the above-mentioned target time window size according to the above-mentioned number of unit reconstructed video frames. In practice, the above-mentioned execution entity may use the ceiling result of the ratio of the target time window size to the above-mentioned number of unit reconstructed video frames as the updated target time window size.
[0074] Step 5: According to the updated target time window size, perform a secondary window division on the updated denoised event information group sequence to obtain a divided event information group sequence. In practice, first, the above-mentioned execution entity can sort each denoised event information included in the denoised event information group sequence in the chronological order represented by each acquisition time to obtain a denoised event information sequence. Then, the above-mentioned execution entity can perform a secondary window division on the denoised event information sequence according to the updated target window size to obtain each divided event information group. For example, the updated target time window size can be 20 ms. The first denoised event information in the denoised event information sequence includes the time represented by the acquisition time as 12:00 sharp. The above-mentioned execution entity can use each denoised event information in the denoised event information sequence with the acquisition time within the range from 12:00 sharp to 12:20 ms as each divided event information and divide them into the same divided event information group to obtain a divided event information group.
[0075] Step 6: Add a pre-constructed video frame quality enhancement network to the above video frame construction model to use the added video frame construction model as a video frame reconstruction model. Among them, the above video frame quality enhancement network can be a neural network model for enhancing image quality.
[0076] Step 7: For each divided event information group in the above divided event information group sequence, perform the following video frame reconstruction steps:
[0077] The first sub-step: Sort the above divided event information group to obtain a divided event information sequence. In practice, the above-mentioned execution entity can sort each divided event information included in the above divided event information group in chronological order to obtain a divided event information sequence.
[0078] The second sub-step: Input the divided event information sequence into the video frame construction model included in the above video frame reconstruction model to obtain a reconstructed video recording video frame.
[0079] The third sub-step is to input the initial video frames of the video recording corresponding to the reconstructed video frames of the video recording and the above-mentioned reconstructed video frames into the video frame quality enhancement network included in the above-mentioned video frame reconstruction model to obtain video frames of the video recording. Among them, the above-mentioned video frames of the video recording can be reconstructed video frames of the video recording after quality enhancement. The initial video frames of the video recording corresponding to the above-mentioned reconstructed video frames of the video recording can be the initial video frames of the video recording in which each event information used is repeated with each event information used in the reconstructed video frames of the video recording. The above-mentioned video frame quality enhancement network can be a neural network model that takes the reconstructed video frames of the video recording and the initial video frames of the video recording corresponding to the reconstructed video frames of the video recording as inputs and the video frames of the video recording as outputs. In practice, since the initial video frames of the video recording corresponding to the reconstructed video frames of the video recording use more event information and there are repetitions with each event information used in constructing the reconstructed video frames of the video recording. Therefore, the quality of the reconstructed video frames of the video recording can be enhanced by the initial video frames of the video recording. The above-mentioned video frame quality enhancement model can enhance the quality of the reconstructed video frames of the video recording by the corresponding initial video frames of the video recording. As an example, the above-mentioned video frame quality enhancement model can be, but is not limited to, a double convolutional neural network model, a DCGAN network (Deep Convolutional Generative Adversarial Networks), and a U-Net neural network model.
[0080] The eighth step is to update the video frame sequence of the video recording according to the obtained video frames of the video recording and the initial video frame sequence. In practice, first, for each initial video frame in the above-mentioned initial video frame sequence of the video recording, the above-mentioned execution subject can insert the video frames of the video recording corresponding to the above-mentioned initial video frames in the order of generation before the above-mentioned initial video frames. Then, the above-mentioned execution subject can determine the initial video frame sequence after insertion as the video frame sequence of the video recording.
[0081] The relevant content of the above embodiments, as an inventive point of the embodiments of the present disclosure, can solve Technical Problem 3: "When the collected event information is relatively sparse, the number of video frames of the reconstructed video is small, resulting in a low frame rate of the reconstructed video and poor quality of the reconstructed video." The factors that lead to poor quality of the reconstructed video are often as follows: When the collected event information is relatively sparse, the number of video frames of the reconstructed video is small, resulting in a low frame rate of the reconstructed video. If the above factors are solved, the effect of improving the quality of the reconstructed video can be achieved. To achieve this effect, the present disclosure performs a secondary window division on the denoised event information group sequence to obtain the divided event information group sequence. Then, taking each divided event information group as a unit, the video frames of the reconstructed video are reconstructed. Then, the quality of the reconstructed video frames is enhanced through the corresponding initial video frames, thereby improving the picture quality of the reconstructed video frames and further improving the quality of the reconstructed video. Also, because the number of video frames of the reconstructed video is increased, the frame rate of the reconstructed video is higher, improving the quality of the reconstructed video.
[0082] Step 105: Perform video encoding processing on the video frame sequence of the recorded video to generate a recorded video for the target monitoring area.
[0083] In some embodiments, the above execution entity may perform video encoding processing on the processed video frame sequence of the recorded video to generate a recorded video for the above target monitoring area. In practice, the above execution entity may perform video encoding processing on the processed video frame sequence of the recorded video through a video encoder to generate a recorded video for the above target monitoring area. As an example, the above video encoder may be, but is not limited to, an AVC encoder, an AV1 encoder, or an FFmpeg encoder.
[0084] Step 106: Perform storage processing on the recorded video.
[0085] In some embodiments, the above execution entity may perform storage processing on the above recorded video. In practice, the above execution entity may store the above recorded video in a caching device.
[0086] In some alternative implementation manners of some embodiments, the above execution entity may perform storage processing on the above recorded video through the following steps, including:
[0087] The first step: Perform compression processing on the above recorded video to obtain the compressed recorded video. In practice, the above execution entity may perform compression processing on the above recorded video through the above video encoder to obtain the compressed recorded video.
[0088] The second step: Store the above compressed recorded video in a caching device.
[0089] The above-mentioned various embodiments of the present disclosure have the following beneficial effects: The event-based video recording and storage method according to some embodiments of the present disclosure can reduce the storage of redundant video frames, thereby reducing the waste of storage space. Specifically, the reason for the waste of storage space is that the mis-touch rate of infrared monitoring is relatively high (for example, the hot air flow generated by the outdoor unit of the air conditioner, etc.). In the case of mis-touch, the video recorded is often redundant video, that is, it does not contain dynamic events, thus causing waste of storage space. Based on this, in the event-based video recording and storage method according to some embodiments of the present disclosure, first, an event information set is obtained. Among them, the event information in the above-mentioned event information set is collected by a monitoring device including an event sensor for a target monitoring area. The event information in the above-mentioned event information set includes pixel position, acquisition time, and a brightness identifier representing the brightness change situation. The above-mentioned event information set corresponds to an acquisition duration. Thus, through the event sensor, when a brightness change event greater than a preset brightness threshold is detected in the monitoring area, event information can be collected, thereby reducing redundant recordings caused by mis-triggering. Then, spatio-temporal noise reduction processing is performed on each event information included in the above-mentioned event information set to obtain a sequence of denoised event information groups. Thus, the noise information in the event information set can be reduced, thereby reducing the additional computing power overhead caused by processing noise event information. After that, according to the above-mentioned sequence of denoised event information groups, each initial video recording frame is constructed to obtain a sequence of initial video recording frames. Thus, discrete event information can be converted into continuous video frames, so that the generated video recording frames contain less redundant information. Next, frame reconstruction processing is performed on the above-mentioned sequence of initial video recording frames to obtain a sequence of video recording frames. Among them, the number of video recording frames included in the above-mentioned sequence of video recording frames is greater than the number of initial video recording frames included in the above-mentioned sequence of initial video recording frames. Thus, by frame reconstruction processing, the number of video frames is increased to fill the time gaps between dynamic events, enhancing the continuity and smoothness of the video stream. Secondly, video encoding processing is performed on the above-mentioned sequence of initial video recording frames to generate a video recording for the above-mentioned target monitoring area. Among them, the video duration of the above-mentioned video recording is equal to the above-mentioned acquisition duration. Thus, video encoding of the sequence of video recording frames can obtain a complete video recording. Finally, storage processing is performed on the above-mentioned video recording. Because a method for constructing and reconstructing video frames based on an event sensor is adopted, the acquisition of redundant recordings can be reduced, and a video recording containing less redundant background information can be generated, thereby reducing the waste of storage space.
[0090] Further referring to Figure 2 , as an implementation of the methods shown in the above figures, the present disclosure provides some embodiments of an event-based video recording and storage device. These device embodiments correspond to Figure 1 the method embodiments shown, and the event-based video recording and storage device can be specifically applied to various electronic devices.
[0091] As Figure 2 shown, the event-based video recording and storage device 200 of some embodiments includes: an acquisition unit 201, a spatio-temporal noise reduction unit 202, a construction unit 203, a frame reconstruction unit 204, a video encoding unit 205, and a storage unit 206. Among them, the acquisition unit 201 is configured to acquire a set of event information, where the event information in the set of event information is collected by a monitoring device including event sensors for a target monitoring area, the event information in the set of event information includes pixel position, acquisition time, and a luminance identifier characterizing the luminance change situation, and the set of event information corresponds to an acquisition duration; the spatio-temporal noise reduction unit 202 is configured to perform spatio-temporal noise reduction processing on each piece of event information included in the set of event information to obtain a sequence of denoised event information groups; the construction unit 203 is configured to construct each initial video recording frame according to the sequence of denoised event information groups to obtain a sequence of initial video recording frames; the frame reconstruction unit 204 is configured to perform frame reconstruction processing on the sequence of initial video recording frames to obtain a sequence of video recording frames, where the number of video recording frames included in the sequence of video recording frames is greater than the number of initial video recording frames included in the sequence of initial video recording frames; the video encoding unit 205 is configured to perform video encoding processing on the sequence of initial video recording frames to generate a video recording for the target monitoring area, where the video duration of the video recording is equal to the acquisition duration; the storage unit 206 is configured to perform storage processing on the video recording.
[0092] It can be understood that the units described in the event-based video recording and storage device 200 correspond to the respective steps in the method described in the reference Figure 1 description. Therefore, the operations, features, and beneficial effects described above for the method also apply to the event-based video recording and storage device 200 and the units included therein, and will not be elaborated herein.
[0093] Next, refer to Figure 3 , which shows a schematic structural diagram of an electronic device 300 suitable for implementing some embodiments of the present disclosure. Figure 3 The shown electronic device is only an example and should not impose any limitations on the functions and usage scopes of the embodiments of the present disclosure.
[0094] As Figure 3As shown, the electronic device 300 may include a processing device (such as a central processing unit, a graphics processing unit, etc.) 301, which may perform various appropriate actions and processes according to the program stored in the read-only memory 302 or the program loaded from the storage device 308 into the random access memory 303. In the random access memory 303, various programs and data required for the operation of the electronic device 300 are also stored. The processing device 301, the read-only memory 302, and the random access memory 303 are connected to each other through a bus 304. The input / output interface 305 is also connected to the bus 304.
[0095] Generally, the following devices may be connected to the I / O interface 305: an input device 306 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 307 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 308 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 309. The communication device 309 may allow the electronic device 300 to communicate with other devices wirelessly or wiredly to exchange data. Although Figure 3 the electronic device 300 with various devices is shown, it should be understood that it is not required to implement or have all the shown devices. More or fewer devices may be implemented or had alternatively. Figure 3 Each block shown in may represent one device or, as needed, multiple devices.
[0096] Specifically, according to some embodiments of the present disclosure, the processes described above with reference to the flowcharts may be implemented as computer software programs. For example, some embodiments of the present disclosure include a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program contains program codes for performing the methods shown in the flowcharts. In such some embodiments, the computer program may be downloaded and installed from the network through the communication device 309, or installed from the storage device 308, or installed from the read-only memory 302. When the computer program is executed by the processing device 301, the above functions defined in the methods of some embodiments of the present disclosure are executed.
[0097] It should be noted that the computer-readable media described in some embodiments of the present disclosure may be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. A computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In some embodiments of the present disclosure, the computer-readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In some embodiments of the present disclosure, a computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium may be transmitted using any appropriate medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.
[0098] In some embodiments, the client and the server can communicate using any currently known or future-developed network protocol such as HTTP (Hyper Text Transfer Protocol), and can be interconnected with digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include local area networks ("LANs"), wide area networks ("WANs"), the Internet (e.g., the Internet), and end-to-end networks (e.g., ad hoc end-to-end networks), as well as any currently known or future-developed networks.
[0099] The above computer-readable medium may be included in the above electronic device; or it may exist independently without being assembled into the electronic device. The above computer-readable medium carries one or more programs. When the above one or more programs are executed by the electronic device, the electronic device is caused to: obtain a set of event information, wherein the event information in the set of event information is collected by a monitoring device including an event sensor for a target monitoring area, the event information in the set of event information includes pixel position, acquisition time, and a brightness identifier representing the brightness change situation, and the set of event information corresponds to an acquisition duration; perform spatio-temporal noise reduction processing on each piece of event information included in the set of event information to obtain a sequence of denoised event information groups; construct each initial video frame according to the sequence of denoised event information groups to obtain a sequence of initial video frames; perform frame reconstruction processing on the sequence of initial video frames to obtain a sequence of video frames, wherein the number of video frames included in the sequence of video frames is greater than the number of initial video frames included in the sequence of initial video frames; perform video encoding processing on the sequence of initial video frames to generate a video recording for the target monitoring area, wherein the video duration of the video recording is equal to the acquisition duration; and perform storage processing on the video recording.
[0100] Computer program code for performing the operations of some embodiments of the present disclosure may be written in one or more programming languages or combinations thereof. The programming languages include object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, executed as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., by using an Internet service provider to connect through the Internet).
[0101] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a portion of code that contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions noted in the blocks may occur in a different order than noted in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, or they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and combinations of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system that performs the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.
[0102] The units described in some embodiments of the present disclosure can be implemented in software or in hardware. The described units can also be provided in a processor. For example, it can be described as: a processor includes: an acquisition unit, a spatio-temporal denoising unit, a construction unit, a frame reconstruction unit, a video encoding unit, and a storage unit. Among them, the names of these units do not constitute a limitation to the unit itself in some cases. For example, the acquisition unit can also be described as "the unit for acquiring a set of event information".
[0103] The functions described above can be performed at least in part by one or more hardware logic components. For example, without limitation, exemplary types of hardware logic components that can be used include: field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on a chip (SOCs), complex programmable logic devices (CPLDs), and so on.
[0104] The above description is only some preferred embodiments of the present disclosure and an explanation of the applied technical principles. Those skilled in the art should understand that the scope of the invention involved in the embodiments of the present disclosure is not limited to the technical solutions formed by the specific combination of the above technical features, but should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above inventive concept. For example, technical solutions formed by mutually replacing the above features with technical features having similar functions (but not limited to) disclosed in the embodiments of the present disclosure.
Claims
1. An event-based video recording storage method, comprising: Obtaining an event information set, wherein the event information in the event information set is collected by a monitoring device including an event sensor for a target monitoring area, the event information in the event information set includes pixel position, acquisition time, and a brightness identifier characterizing the brightness change situation, and the event information set corresponds to an acquisition duration; Performing spatio-temporal noise reduction processing on each event information included in the event information set to obtain a sequence of denoised event information groups, including: dividing the event information set according to the acquisition time included in the event information in the event information set to obtain each event information group; performing a temporal sorting process on each event information group to obtain an event information group sequence; selecting an event information from the first event information group in the event information group sequence as a signal event information and inserting it into a pre-constructed signal event information queue to update the signal event information queue; inserting the selected event information as noise event information into a pre-constructed noise event information queue to update the noise event information queue; performing spatial noise reduction processing on the event information group sequence according to the updated signal event information queue and the updated noise event information queue to obtain a sequence of denoised event information groups; Constructing each initial video recording frame according to the sequence of denoised event information groups to obtain an initial video recording frame sequence; Performing frame reconstruction processing on the initial video recording frame sequence to obtain a video recording frame sequence, wherein the number of video recording frames included in the video recording frame sequence is greater than the number of initial video recording frames included in the initial video recording frame sequence; Performing video encoding processing on the video recording frame sequence to generate a video recording for the target monitoring area, wherein the video duration of the video recording is equal to the acquisition duration; Performing storage processing on the video recording.
2. The method according to claim 1, wherein The step of constructing each initial video recording frame according to the sequence of denoised event information groups to obtain an initial video recording frame sequence includes: Performing clustering processing on each denoised event information in the sequence of denoised event information groups to obtain a set of denoised event information clusters; Determining the time window size corresponding to each denoised event information cluster in the set of denoised event information clusters; Determining the average value of the determined time window sizes as the target time window size; Performing window division processing on the sequence of denoised event information groups according to the target time window size to update the sequence of denoised event information groups; Generating an initial video recording frame sequence according to the updated sequence of denoised event information groups and a pre-trained video frame construction model.
3. The method according to claim 2, wherein, The step of generating an initial video recording frame sequence according to the updated sequence of denoised event information groups and a pre-trained video frame construction model includes: For each denoised event information group in the updated sequence of denoised event information groups, performing the following video frame construction steps: Sorting the denoised event information group to obtain a denoised event information sequence; Input the denoised event information sequence into the video frame construction model to generate initial recorded video frames; Determine each of the generated initial recorded video frames as an initial recorded video frame sequence.
4. The method according to claim 1, wherein The storing process for the recorded video includes: Perform compression processing on the recorded video to obtain a compressed recorded video; Store the compressed recorded video in a buffer device.
5. An event-based recorded video storage device, comprising: An acquisition unit configured to acquire a set of event information, wherein the event information in the set of event information is collected by a monitoring device including an event sensor for a target monitoring area, the event information in the set of event information includes pixel position, acquisition time, and a luminance identifier representing the luminance change situation, and the set of event information corresponds to an acquisition duration; A spatio-temporal denoising unit configured to perform spatio-temporal denoising processing on each piece of event information included in the set of event information to obtain a sequence of denoised event information groups, including: dividing the set of event information according to the acquisition time included in the event information in the set of event information to obtain each event information group; performing a timing sorting process on each event information group to obtain a sequence of event information groups; selecting an event information from the first event information group in the sequence of event information groups as a signal event information and inserting it into a pre-constructed signal event information queue to update the signal event information queue; inserting the selected event information as noise event information into a pre-constructed noise event information queue to update the noise event information queue; performing spatial denoising processing on the sequence of event information groups according to the updated signal event information queue and the updated noise event information queue to obtain a sequence of denoised event information groups; A construction unit configured to construct each initial recorded video frame according to the sequence of denoised event information groups to obtain an initial recorded video frame sequence; A frame reconstruction unit configured to perform frame reconstruction processing on the initial recorded video frame sequence to obtain a recorded video frame sequence, wherein the number of recorded video frames included in the recorded video frame sequence is greater than the number of initial recorded video frames included in the initial recorded video frame sequence; A video encoding unit configured to perform video encoding processing on the recorded video frame sequence to generate a recorded video for the target monitoring area, wherein the video duration of the recorded video is equal to the acquisition duration; A storage unit configured to perform a storing process on the recorded video.
6. An electronic device, comprising: One or more processors; A storage device having stored thereon one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1 to 4.
7. A computer-readable medium having a computer program stored thereon, wherein, The computer program, when executed by a processor, implements the method according to any one of claims 1 to 4.
Citation Information
Patent Citations
Object recognition method and device, chip and electronic equipment
CN113408671A
Agricultural information transmission method and device, electronic equipment and storage medium
CN116320321A