A method and device for background restoration of video data
By extracting keyframes in the video and using object detection technology to separate the foreground and background, combined with the filling method of multiple background images, the problems of low video background restoration efficiency and unstable image quality in the prior art are solved, and a fast and efficient background restoration effect is achieved.
Patent Information
- Application Number
- CN202110339807.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-03-30
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2041-03-30
AI Technical Summary
The existing video background restoration methods are inefficient and slow, and the image quality has a strong dependence on the foreground movement rate, resulting in poor results when the foreground movement is slow, leaving traces.
By extracting some keyframe data of the video, using object detection to distinguish the foreground and background, and filling it with multiple background images, the foreground background on a single frame is achieved quickly and efficiently resolving the background environment.
This greatly reduces the calculation amount, significantly improves speed and efficiency, avoids the image quality problems affected by the foreground movement rate in traditional methods, and achieves high-quality background restoration.
Smart Images

Figure CN113095176B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a method and device for background restoration of video data. Background Art
[0002] Video background restoration refers to the process of removing the foreground from a video and extracting the background through certain technical means. At present, the method of video background restoration is mainly to calculate the pixels corresponding to all frames of the video pixel by pixel. However, the video content is random in a broad sense. When is the optimal value for the background image of a video when iterating frame by frame? The law of large numbers believes that the more iterations, the closer the background image is to the true value and the better the robustness. Therefore, most background restoration technologies are based on the synchronous real-time maintenance of continuous iteration of video frames. At present, the method of background restoration for a video is generally to iterate each frame to obtain the background restoration image. In addition, the stacking of algorithm calculations during iteration leads to low efficiency and slow speed of background restoration. Moreover, the quality of background restoration images processed by most traditional methods is strongly dependent on the moving rate of the foreground: that is, in a video, when the foreground object, such as a person, moves very slowly or does not move at all, the foreground occupies most of the video frames in a fixed area. At this time, the effect of processing with traditional methods is usually not good, leaving traces and resulting in low quality. Summary of the invention
[0003] In view of this, an embodiment of the present invention provides a method and device for background restoration of video data, which can extract partial frame data of the video and use target detection to distinguish the foreground and background, thereby realizing fast and effective separation of the foreground and background on a single frame, and filling in with multiple background images to efficiently obtain the restored background environment.
[0004] To achieve the above objective, according to one aspect of an embodiment of the present invention, a method for background restoration of video data is provided.
[0005] The method for restoring the background of video data according to an embodiment of the present invention includes:
[0006] Based on a sampling threshold, key frame sampling is performed on the video to be processed to obtain a key frame set; wherein the sampling threshold is less than the total number of frames of the video to be processed;
[0007] Performing foreground detection on each key frame in the key frame set to determine foreground target information in each key frame;
[0008] Based on the foreground detection, a background template and a background library are obtained; wherein the background template is obtained according to the first frame data, and the background library includes the second frame data and its foreground target information; the key frame set includes the first frame data and the second frame data;
[0009] The background template is filled with background information according to the second frame data in the background library and its foreground target information to obtain background restoration data of the video to be processed.
[0010] Optionally, based on a set sampling threshold, key frame sampling is performed on the video to be processed to obtain a key frame set, including: obtaining the video to be processed and determining the total number of frames of the video to be processed; determining a sampling interval based on the total number of frames and the sampling threshold; wherein the sampling threshold is less than the total number of frames of the video to be processed; and according to the sampling interval, key frame sampling is performed on the video to be processed to obtain a key frame set.
[0011] Optionally, based on the foreground detection, the steps of obtaining a background template and a background library include: selecting a frame data in the key frame set as the first frame data; performing foreground subtraction processing on the first frame data to obtain a background template based on the foreground target information of the first frame data; and forming a background library with the second frame data in the key frame set except the first frame data and its foreground target information.
[0012] Optionally, the step of filling the background of the background template according to the second frame data in the background library and its foreground target information includes: selecting a frame data from the second frame data in the background library as iterative frame data; verifying whether the target pixel point at the corresponding position of the iterative frame data belongs to the foreground target according to the position of the cut-out part in the background template; if not, copying the target pixel point to the background module; otherwise, selecting a new frame data from the second frame data in the background library again as iterative frame data.
[0013] Optionally, before selecting a new frame data from the second frame data of the background library as the iterative frame data again, the method further includes: determining whether all the second frame data of the background library are selected as the iterative frame data when it is determined that there are no unfilled pixels in the pixel points of the removed part of the background template;
[0014] If yes, the background template is filled according to the surrounding pixels of the unfilled pixels; otherwise, a new frame data is selected from the second frame data of the background library as the iterative frame data.
[0015] Optionally, the step of verifying whether the target pixel point at the corresponding position of the iterative frame data belongs to the foreground target includes: creating a binary matrix mask of the iterative frame data; and verifying whether the target pixel point at the corresponding position of the iterative frame data belongs to the foreground target based on the binary matrix mask.
[0016] Optionally, the step of performing foreground detection on each key frame in the key frame set and determining foreground target information in each key frame includes: performing foreground target detection on each key frame in the key frame set based on a deep learning target detection algorithm; determining position information of the foreground target in each key frame, wherein the foreground target information includes at least the position information of the foreground target.
[0017] Optionally, the position information is target frame information of the foreground object; wherein the target frame information at least includes length, width and diagonal pixel coordinates.
[0018] To achieve the above objective, according to another aspect of an embodiment of the present invention, a device for performing background restoration on video data is provided.
[0019] The device for performing background restoration on video data according to an embodiment of the present invention includes:
[0020] A frame sampling module, used for sampling key frames of the video to be processed based on a sampling threshold to obtain a key frame set; wherein the sampling threshold is less than the total number of frames of the video to be processed;
[0021] A foreground detection module, used to perform foreground detection on each key frame in the key frame set to determine foreground target information in each key frame;
[0022] A background module, used to obtain a background template and a background library based on the foreground detection; wherein the background template is obtained according to the first frame data, the background library includes the second frame data and its foreground target information; the key frame set includes the first frame data and the second frame data;
[0023] The filling module is used to fill the background template according to the second frame data in the background library and its foreground target information to obtain the background restoration data of the video to be processed.
[0024] Optionally, the frame sampling module is also used to obtain the video to be processed and determine the total number of frames of the video to be processed; determine the sampling interval based on the total number of frames and a sampling threshold; wherein the sampling threshold is less than the total number of frames of the video to be processed; and perform key frame sampling on the video to be processed according to the sampling interval to obtain a key frame set.
[0025] Optionally, the background module is also used to select a frame data in the key frame set as the first frame data; perform foreground subtraction processing on the first frame data to obtain a background template based on the foreground target information of the first frame data; and form a background library with the second frame data in the key frame set except the first frame data and its foreground target information.
[0026] Optionally, the filling module is also used to select a frame data from the second frame data of the background library as iterative frame data; based on the position of the cut-out part in the background template, verify whether the target pixel point at the corresponding position of the iterative frame data belongs to the foreground target; if not, copy the target pixel point to the background module; otherwise, select a new frame data from the second frame data of the background library again as iterative frame data.
[0027] Optionally, the filling module is further used to determine whether all the second frame data of the background library are selected as iterative frame data when it is determined that there are no unfilled pixels in the removed pixel points of the background template;
[0028] If yes, the background template is filled according to the surrounding pixels of the unfilled pixels; otherwise, a new frame data is selected from the second frame data of the background library as the iterative frame data.
[0029] Optionally, the filling module is further used to create a binary matrix mask of the iterative frame data; and verify, based on the binary matrix mask, whether a target pixel point at a corresponding position of the iterative frame data belongs to a foreground target.
[0030] Optionally, the foreground detection module performs foreground target detection on each key frame in the key frame set based on a deep learning target detection algorithm; and determines the position information of the foreground target in each key frame, wherein the foreground target information at least includes the position information of the foreground target.
[0031] Optionally, the position information is target frame information of the foreground object; wherein the target frame information at least includes length, width and diagonal pixel coordinates.
[0032] To achieve the above objective, according to another aspect of an embodiment of the present invention, an electronic device is provided.
[0033] The electronic device of an embodiment of the present invention includes: one or more processors; a storage device for storing one or more programs, when the one or more programs are executed by the one or more processors, the one or more processors implement any of the above-mentioned methods for background restoration of video data.
[0034] To achieve the above-mentioned purpose, according to another aspect of an embodiment of the present invention, a computer-readable medium is provided, on which a computer program is stored, characterized in that when the program is executed by a processor, any of the above-mentioned methods for background restoration of video data is implemented.
[0035] One embodiment of the above invention has the following advantages or beneficial effects: Unlike the existing video background restoration method which iterates frame by frame, the embodiment of the present invention only extracts some key frames, which greatly reduces the amount of calculation and significantly improves the speed and efficiency. In addition, the foreground and background are distinguished by target detection, so that the foreground and background on a single frame can be quickly and effectively separated, and multiple background images are used to fill in the gaps to efficiently obtain the restored background environment.
[0036] The further effects of the above-mentioned non-conventional optional manner will be described below in conjunction with the specific implementation manner. BRIEF DESCRIPTION OF THE DRAWINGS
[0037] The accompanying drawings are used to better understand the present invention and do not constitute an improper limitation of the present invention.
[0038] Figure 1 is a schematic diagram of the main process of a method for background restoration of video data according to an embodiment of the present invention;
[0039] Figure 2 is a schematic diagram of a method for obtaining a key frame set according to an embodiment of the present invention;
[0040] Figure 3 is a schematic diagram of a method for performing background filling on a background template according to an embodiment of the present invention;
[0041] Figure 4 is a schematic diagram of a method for performing background restoration on video data according to an embodiment of the present invention;
[0042] Figure 5 is a schematic diagram of main modules of an apparatus for performing background restoration on video data according to an embodiment of the present invention;
[0043] Figure 6 is an exemplary system architecture diagram to which embodiments of the present invention may be applied;
[0044] Figure 7 It is a schematic diagram of the structure of a computer system of a terminal device or a server suitable for implementing an embodiment of the present invention. DETAILED DESCRIPTION
[0045] The following is a description of exemplary embodiments of the present invention in conjunction with the accompanying drawings, including various details of the embodiments of the present invention to facilitate understanding, which should be considered as merely exemplary. Therefore, it should be recognized by those of ordinary skill in the art that various changes and modifications may be made to the embodiments described herein without departing from the scope and spirit of the present invention. Similarly, for clarity and conciseness, the description of well-known functions and structures is omitted in the following description.
[0046] Some technical terms in the embodiments of the present invention are explained as follows:
[0047] Static video: refers to a video with a fixed background. This type of video is usually generated by a fixed camera shooting at a fixed point. In contrast, dynamic video refers to a video with a constantly changing background. This type of video is usually shot by a fixed camera rotating or a movable camera device.
[0048] Foreground: refers to a class of objects that move in the video or have special meaning in the image (such as people, animals, vehicles and other movable objects).
[0049] Background: usually refers to the entire environment space in a video or picture excluding the foreground.
[0050] Video background restoration: refers to the process of removing the foreground of a video and extracting the background through certain technical means.
[0051] Target: In computer vision, a collection of people or objects with similar characteristics contained in a video or image (such as people, pigs, cars, traffic lights, boxes, etc.).
[0052] Object detection: One of the directions of AI computer vision, also known as object extraction, is an image segmentation based on the geometric and statistical features of the target, combining the segmentation and recognition of the target in one. Specifically, object detection is a computer vision technique that allows the identification and localization of objects in an image or video. Object detection can be understood as two parts, object localization and object classification. Localization can be understood as predicting the exact location of the object in the image (bounding box), while classification is defining which class it belongs to (human / car / dog, etc.).
[0053] Target box: The boundary that the computer uses to segment the identified target through calculation, usually in the form of a rectangle.
[0054] Figure 1 is a schematic diagram of the main process of a method for background restoration of video data according to an embodiment of the present invention. Figure 1 As shown, the method for background restoration of video data in an embodiment of the present invention mainly includes:
[0055] Step S101: based on a sampling threshold, sampling key frames of the video to be processed to obtain a key frame set; wherein the sampling threshold is less than the total number of frames of the video to be processed;
[0056] Step S102: performing foreground detection on each key frame in the key frame set to determine foreground target information in each key frame;
[0057] Step S103: Based on foreground detection, a background template and a background library are obtained; wherein the background template is obtained according to the first frame data, the background library includes the second frame data and its foreground target information; the key frame set includes the first frame data and the second frame data;
[0058] Step S104: filling the background template with the background according to the second frame data in the background library and its foreground target information to obtain background restoration data of the video to be processed.
[0059] According to the embodiments of the present invention, by extracting some key frames, the amount of calculation is greatly reduced, and the speed and efficiency are significantly improved. In addition, by using target detection to distinguish the foreground and background, the foreground and background on a single frame are quickly and effectively separated, and multiple background images are used to fill in the gaps to efficiently obtain the restored background environment.
[0060] Figure 2 is a schematic diagram of a method for obtaining a key frame set according to an embodiment of the present invention; Figure 2 As shown, for step S101, in a preferred embodiment, based on a set sampling threshold, the step of sampling key frames of the video to be processed to obtain a key frame set includes:
[0061] Step S201: Obtain a video to be processed, and determine the total number of frames of the video to be processed.
[0062] Step S202: Determine a sampling interval according to the total number of frames and a sampling threshold; wherein the sampling threshold is less than the total number of frames of the video to be processed. The sampling threshold may be extracted and preset, in which case the sampling threshold may be fixed. In another embodiment, the sampling threshold may also be determined according to the total number of frames of the video, in which case the sampling threshold may be dynamically adjusted.
[0063] Step S203: performing key frame sampling on the video to be processed according to the sampling interval to obtain a key frame set.
[0064] This embodiment adopts a uniform sampling method, firstly ensuring that the sampled frames are reasonably distributed. For example, even if the character's speed is very slow, or the background appears for a short time (a person does not move for 95 frames and leaves in the last 5 frames), as long as there is a key frame with a corresponding background pixel, the background image can be restored. Because it is a direct slice copy, there is no trace problem.
[0065] In a preferred embodiment, in the process of obtaining the background template and the background library based on foreground detection, a frame data in the key frame set is selected as the first frame data. Then, according to the foreground target information of the first frame data, the first frame data is subjected to foreground subtraction processing to obtain the background template. In addition, the second frame data in the key frame set other than the first frame data and its foreground target information are used to form a background library. According to this embodiment, a frame data can be randomly selected from the extracted frame data, and the foreground target inside can be removed to obtain the background template. The other key frames are saved together with the target detection information to form a background library for subsequent iteration.
[0066] Figure 3 is a schematic diagram of a method for filling a background template according to an embodiment of the present invention. Figure 3 As shown, for step S104, in a preferred embodiment, the step of filling the background template with the background according to the second frame data and its foreground target information in the background library includes:
[0067] Step S301: selecting a frame data from the second frame data of the background library as iterative frame data.
[0068] Step S302: According to the position of the removed part in the background template, verify whether the target pixel point at the corresponding position of the iterative frame data belongs to the foreground target. If not, proceed to step S303. Otherwise, proceed to step S301, that is, select a new frame data from the second frame data of the background library as the iterative frame data again.
[0069] Step S303: The target pixel is copied to the background module.
[0070] According to the embodiment of the present invention, in each iteration, a key frame is selected from the background library, and the corresponding position of the key frame taken out from the background library is found through the position coordinates of the background template that is cut out, and it is checked whether the pixel at the position belongs to the foreground part. Then, multiple background images are used for filling to efficiently obtain the restored background environment.
[0071] In a preferred embodiment, before selecting a new frame data from the second frame data of the background library as the iterative frame data again, it is determined whether all the second frame data of the background library are selected as the iterative frame data when it is determined that there are no unfilled pixels in the pixel points of the removed part of the background template. If so, the background template is filled according to the surrounding pixels of the unfilled pixels; otherwise, a new frame data is selected from the second frame data of the background library as the iterative frame data again. If all key frames are run, and there are still unfilled pixels, the pixel point is manually filled by taking the average value of the surrounding pixels. Preferably, the pixel value of the candidate pixel can be based on the average value of the 8 pixels around the unfilled pixel. And, it can be judged whether it is filled (whether it is 1, if it is 1, it is not filled) through the records on the mask matrix.
[0072] In a preferred embodiment, in the process of verifying whether the target pixel point at the corresponding position of the iterative frame data belongs to the foreground target, a binary matrix mask of the iterative frame data is created. According to the binary matrix mask, it is verified whether the target pixel point at the corresponding position of the iterative frame data belongs to the foreground target. The binary matrix mask K is recorded by 0 and 1 to record whether the template is the foreground (i.e., whether it is the part that is cut off). When initializing the binary matrix K, the matrix value is set to all 0, and then the part that is cut off is set to 1. When the background with the cut-off part is restored, the corresponding area of the mask matrix K corresponding to 1 is set to 0. By detecting whether there is still 1 in the mask matrix, it is also possible to quickly verify whether the entire background is filled in.
[0073] In a preferred embodiment, foreground detection is performed on each key frame in the key frame set. In the process of determining the foreground target information in each key frame, foreground target detection is performed on each key frame in the key frame set based on a deep learning target detection algorithm. Then, the position information of the foreground target in each key frame is determined, wherein the foreground target information at least includes the position information of the foreground target. The position information is the target frame information of the foreground target; wherein the target frame information at least includes the length, width and diagonal pixel coordinates. Generally speaking, the coordinates of the upper left and lower right corner pixels are used by default to remember the rectangular coordinate frame, and the position of the rectangular coordinate frame must be determined by the coordinates of the two diagonal points to be uniquely determined. For example, the diagonal pixel coordinates are: upper left and lower right.
[0074] In the prior art, the following methods are mainly used for video background restoration:
[0075] 1) Mean method, median method, sliding mean filter, single Gaussian
[0076] This type of method is to calculate the pixels corresponding to all frames of the video pixel by pixel, such as the mean method. It is believed that the background is generally an object that does not move, so it can be assumed that its pixel value is almost always the same over a long period of time. Then you can take several pictures and add up the pixel sizes of their corresponding points, and then calculate the average, which can be considered to be the required background. It can be expressed by the following formula:
[0077]
[0078] Among them, the method of processing by the mean value method has a certain effect, especially when the background is a distant view or the change is not very large, the processing effect is relatively good. However, for the moving image shaking, the effect is not very ideal, leaving traces, that is, with a blurred foreground.
[0079] 2) Inter-frame difference method
[0080] |frame(i)-frame(i-1)|>Th|frame(i)-frame(i-1)|>Th|frame(i)-frame(i-1)|>Th, the background is the previous frame. Each frame is differentially calculated with the previous frame. The extraction effect is obviously related to the speed and frame rate of the moving foreground object (the frame rate refers to the number of pictures per second). By extension, the selective background modeling based on the statistical model is actually the mixed Gaussian method.
[0081] The problem with this method is that there may be "holes" in the object. The holes are caused by a large moving object, and there is an overlapping part with very close pixels between its two frames, so this part is cut off by the difference.
[0082] 3) Mixed Gaussian method
[0083] The adaptive background difference algorithm based on the mixed Gaussian model is similar to the inter-frame difference method. It uses the mixed Gaussian distribution model to characterize the characteristics of each pixel in the image frame. When a new image frame is obtained, the mixed Gaussian distribution model is updated in time. At a certain moment, a subset of the mixed Gaussian model is selected to represent the current background. If a pixel point in the current image frame matches the background subset of the mixed Gaussian model, it is judged as the background, otherwise it is judged as the foreground point.
[0084] The problem with this method is that the quality of the background depends on the speed of the foreground object (such as a person). If it stays at a certain point for too long, it will cause interference. The background extracted by the mixed Gaussian method leaves a shadow because the person in the picture stands in the same place for a long time.
[0085] 4) Energy analysis method
[0086] The concept is slightly complicated. The continuous image sequence is regarded as a three-dimensional space consisting of two-dimensional space plus time. Then the components of each pixel on each spatiotemporal gradient are calculated, and finally these spatiotemporal gradient components are smoothed by Gaussian filtering to obtain motion energy. Since the pixels contained in the moving object basically move in one direction, the motion energy in this direction is relatively large. The motion energy method can eliminate the influence of chaotic motion and detect the real moving object. The problem with this method is that it can only roughly estimate the position of the real moving foreground object, and it is difficult to accurately extract the moving object.
[0087] 5) Optical flow method
[0088] The concept of optical flow method comes from the optical flow field. When the image of a moving object moves on the surface, it is called the optical flow field, which is a two-dimensional velocity field. The optical flow method calculates the size and direction of the movement of each pixel point based on a continuous multi-frame image sequence, which reflects the change trend of the grayscale of each pixel point on the image. The problems of this method are: the calculation is complex, often requires special hardware support, and it is difficult to meet the real-time requirements.
[0089] like Figure 4 As shown, the method for background restoration of video data in an embodiment of the present invention mainly includes:
[0090] Step S401: Video key frame sampling. In the embodiment of the present invention, the sampling threshold can be set to a constant threshold C, for example, 20, which represents the desired number of key frames. For an input video, regardless of its length, 20 frames are extracted as a set of key frames. Specifically, the sampling method adopts a balanced sampling method, first obtaining the total number of video frames T, calculating the sampling video interval: Interval = T / C, and then specifying the frame number during sampling, and extracting one frame every Interval frames.
[0091] Step S402: Use the YoloV5 target detection algorithm to identify the foreground. In an embodiment of the present invention, the YoloV5 target detection algorithm is used to identify the foreground. In other embodiments, the foreground detection algorithm can be replaced by other deep learning algorithms. Among them, the YoloV5 target detection algorithm has a smaller size than the previous generation model, and the target detection efficiency is improved. In the inference stage, Yolov5 uses the method of reducing black edges to increase the speed of inference. Modifications have been made in the letterbox function of the code datasets.py to adaptively add the least black edges to the original image. For example: For example, my 1000×800 picture is not directly scaled to 608×608, but 608 / 1000=0.608 is calculated and then scaled to 608×486, and then 608-486=122 is calculated and then np.mod(122,32) takes the remainder to get 26, and then averages 13 to fill the height of the picture, and finally 608×512.
[0092] For each key frame, the foreground target (people, cars, bicycles, motorcycles, animals, etc.) is detected by the YoloV5 target detection algorithm, and the relevant position information, that is, the target frame position (that is, the position information), is retained. In an embodiment of the present invention, the target frame position includes the length, width, upper left and lower right corner pixel coordinates.
[0093] Step S403: randomly select a key frame to remove the foreground and obtain a background template with a hole. And other key frames are used as background libraries. Specifically, randomly select a key frame in the key frame set, remove the foreground target inside, and obtain the background template M. And save the other key frames together with the foreground target detection information to form a background library.
[0094] Step S404: Fill the corresponding positions of the holes in the background template with the background library images. Select a key frame B from the background library, find the corresponding position in the key frame B through the position coordinates of the background template M that was cut out, and check whether the pixel at this position of B belongs to the foreground part. If the pixel at this position of B does not belong to the foreground part (i.e., belongs to the background part), copy the pixel to the background template M; if the pixel at this position of B belongs to the foreground part, ignore it.
[0095] Step S405: Determine whether the background template is filled. If yes, proceed to step S408, otherwise proceed to step S406. After each round of iteration, check whether all the pixels removed from the background template are filled. If not, select the next key frame from the background library for the next round of iteration, otherwise return to the filled background template.
[0096] Step S406: Determine whether there are any remaining key frames in the background library. If yes, return to step S404, otherwise proceed to step S407. If all key frames have been run, and there are still unfilled pixels, the pixel value is taken as the average value of the surrounding pixels for manual filling.
[0097] Step S407: The remaining pixels are taken as the average value of the surrounding pixels for manual filling.
[0098] Step S408: Obtain a background restoration picture with the filling completed.
[0099] The embodiment of the present invention only extracts some key frames for calculation. The amount of calculation is greatly reduced, and the speed and efficiency are significantly improved. This advantage is more reflected in the longer the video, because the extracted key frames are still fixed, so the processing time does not change much, and the number of frames of the video increases with the length of time, and the processing time increases linearly. In addition, because the quality of the background restoration image processed by most traditional methods is strongly dependent on the moving rate of the foreground: that is, in a video, when the foreground object such as a person moves very slowly, or simply does not move, resulting in the foreground occupying most of the video frames in a fixed area, the effect processed by the traditional method is usually not good, and traces will be left, resulting in low quality. The embodiment of the present invention successfully overcomes the problem of uncertain background restoration quality of the existing method. Moreover, the embodiment of the present invention adopts a uniform sampling method, first of all, to ensure that the sampled frames are reasonably distributed, even if the person's speed is very slow, or the background appears for a short time, as long as there is a corresponding background pixel in the key frame, the background picture can be restored. Because it is a direct slice copy, there is no problem of traces.
[0100] Figure 5 is a schematic diagram of main modules of an apparatus for background restoration of video data according to an embodiment of the present invention, such as Figure 5 As shown, the device 500 for performing background restoration on video data according to the embodiment of the present invention includes a frame sampling module 501 , a foreground detection module 502 , a background module 503 , and a filling module 504 .
[0101] The frame sampling module 501 is used to perform key frame sampling on the video to be processed based on a sampling threshold to obtain a key frame set; wherein the sampling threshold is less than the total number of frames of the video to be processed.
[0102] The foreground detection module 502 is used to perform foreground detection on each key frame in the key frame set to determine foreground target information in each key frame.
[0103] The background module 503 is used to obtain a background template and a background library based on foreground detection; wherein the background template is obtained according to the first frame data, the background library includes the second frame data and its foreground target information; and the key frame set includes the first frame data and the second frame data.
[0104] The filling module 504 is used to fill the background template according to the second frame data and its foreground target information in the background library to obtain the background restoration data of the video to be processed.
[0105] According to the embodiment of the present invention, since only some key frames are extracted, the amount of calculation is greatly reduced, and the speed and efficiency are significantly improved. In addition, the foreground and background are distinguished by using target detection, so that the foreground and background on a single frame can be quickly and effectively separated, and multiple background images are used to fill in the gaps, so as to efficiently obtain the restored background environment.
[0106] Preferably, the frame sampling module is also used to obtain the video to be processed and determine the total number of frames of the video to be processed; determine the sampling interval based on the total number of frames and the sampling threshold; wherein the sampling threshold is less than the total number of frames of the video to be processed; and perform key frame sampling on the video to be processed according to the sampling interval to obtain a key frame set.
[0107] Preferably, the background module is also used to select a frame data in the key frame set as the first frame data; perform foreground subtraction processing on the first frame data to obtain a background template based on the foreground target information of the first frame data; and form a background library with the second frame data in the key frame set except the first frame data and its foreground target information.
[0108] Preferably, the filling module is also used to select a frame data from the second frame data of the background library as the iterative frame data; based on the position of the cut-out part in the background template, verify whether the target pixel point at the corresponding position of the iterative frame data belongs to the foreground target; if not, copy the target pixel point to the background module; otherwise, select a new frame data from the second frame data of the background library again as the iterative frame data.
[0109] Preferably, the filling module is further used to determine whether all the second frame data of the background library are selected as iterative frame data when it is determined that there are no unfilled pixels in the removed pixel points of the background template. If yes, the background template is filled according to the surrounding pixels of the unfilled pixels; otherwise, a new frame data is selected from the second frame data of the background library as iterative frame data.
[0110] Preferably, the filling module is further used to create a binary matrix mask of the iterative frame data; and verify whether the target pixel point at the corresponding position of the iterative frame data belongs to the foreground target according to the binary matrix mask.
[0111] Preferably, the foreground detection module performs foreground target detection on each key frame in the key frame set based on a deep learning target detection algorithm; determines the position information of the foreground target in each key frame, wherein the foreground target information at least includes the position information of the foreground target.
[0112] Preferably, the position information is target frame information of the foreground target; wherein the target frame information at least includes length, width and diagonal pixel coordinates.
[0113] Different from the nature of frame-by-frame iteration of the existing methods, the embodiment of the present invention only extracts some key frames for calculation. The amount of calculation is greatly reduced, and the speed and efficiency are significantly improved. This advantage is more reflected in the longer the video, because the extracted key frames are still fixed, so the processing time does not change much, while the number of frames of the video increases with the length of time, and the processing time increases linearly. In addition, because the quality of the background restoration image processed by most traditional methods is strongly dependent on the moving rate of the foreground: that is, in a video, when the foreground object such as a person moves very slowly, or simply does not move, resulting in the foreground occupying most of the video frames in a fixed area, the effect processed by the traditional method is usually not good, and traces will be left, resulting in low quality. The embodiment of the present invention successfully overcomes the problem of uncertain background restoration quality of the existing method. Moreover, the embodiment of the present invention adopts a uniform sampling method, first of all, to ensure that the sampled frames are reasonably distributed, even if the person's speed is very slow, or the background appears for a short time, as long as there is a corresponding background pixel in the key frame, the background image can be restored. Because it is a direct slice copy, there is no problem of traces.
[0114] Figure 6 An exemplary system architecture 600 is shown in which a method for background restoration of video data or an apparatus for background restoration of video data according to an embodiment of the present invention can be applied.
[0115] like Figure 6 As shown, system architecture 600 may include terminal devices 601, 602, 603, network 604 and server 605. Network 604 is used to provide a medium for communication links between terminal devices 601, 602, 603 and server 605. Network 604 may include various connection types, such as wired, wireless communication links or optical fiber cables, etc.
[0116] Users can use terminal devices 601, 602, and 603 to interact with server 605 through network 604 to receive or send messages, etc. Various communication client applications can be installed on terminal devices 601, 602, and 603, such as shopping applications, web browser applications, search applications, instant messaging tools, email clients, social platform software, etc. (only examples).
[0117] The terminal devices 601 , 602 , and 603 may be various electronic devices having a display screen and supporting web browsing, including but not limited to smart phones, tablet computers, laptop computers, and desktop computers, etc.
[0118] Server 605 may be a server that provides various services, such as a backend management server (only an example) that provides support for shopping websites browsed by users using terminal devices 601, 602, and 603. The backend management server may analyze and process the received data such as product information query requests, and feed back the processing results to the terminal device.
[0119] It should be noted that the method for performing background restoration on video data provided in the embodiment of the present invention is generally executed by the server 605 , and accordingly, the device for performing background restoration on video data is generally disposed in the server 605 .
[0120] It should be understood that Figure 6 The number of terminal devices, networks and servers in the embodiment is only for illustration. Any number of terminal devices, networks and servers may be provided according to implementation requirements.
[0121] Reference below Figure 7 , which shows a schematic diagram of the structure of a computer system 700 of a terminal device suitable for implementing an embodiment of the present invention. Figure 7 The terminal device shown is only an example and should not bring any limitation to the functions and scope of use of the embodiments of the present invention.
[0122] like Figure 7 As shown, the computer system 700 includes a central processing unit (CPU) 701, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 702 or a program loaded from a storage part 708 into a random access memory (RAM) 703. In the RAM 703, various programs and data required for the operation of the system 700 are also stored. The CPU 701, the ROM 702, and the RAM 703 are connected to each other via a bus 704. An input / output (I / O) interface 705 is also connected to the bus 704.
[0123] The following components are connected to the I / O interface 705: an input section 706 including a keyboard, a mouse, etc.; an output section 707 including a cathode ray tube (CRT), a liquid crystal display (LCD), etc., and a speaker, etc.; a storage section 708 including a hard disk, etc.; and a communication section 709 including a network interface card such as a LAN card, a modem, etc. The communication section 709 performs communication processing via a network such as the Internet. A drive 710 is also connected to the I / O interface 705 as needed. A removable medium 711, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 710 as needed, so that a computer program read therefrom is installed into the storage section 708 as needed.
[0124] In particular, according to the embodiments disclosed in the present invention, the process described above with reference to the flowchart can be implemented as a computer software program. For example, the embodiments disclosed in the present invention include a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program includes a program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from the network through the communication part 709, and / or installed from the removable medium 711. When the computer program is executed by the central processing unit (CPU) 701, the above-mentioned functions defined in the system of the present invention are executed.
[0125] It should be noted that the computer-readable medium shown in the present invention may be a computer-readable signal medium or a computer-readable storage medium or any combination of the above two. The computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or device, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present invention, a computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in combination with an instruction execution system, device or device. In the present invention, a computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, which carries a computer-readable program code. This propagated data signal may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. Computer-readable signal media may also be any computer-readable medium other than computer-readable storage media, which may send, propagate or transmit a program for use by or in conjunction with an instruction execution system, apparatus or device. The program code contained on the computer-readable medium may be transmitted using any appropriate medium, including but not limited to: wireless, wire, optical cable, RF, etc., or any suitable combination of the above.
[0126] The flow chart and block diagram in the accompanying drawings illustrate the possible architecture, function and operation of the system, method and computer program product according to various embodiments of the present invention. In this regard, each box in the flow chart or block diagram can represent a module, a program segment, or a part of a code, and the above-mentioned module, program segment, or a part of a code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order from the order marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram or flow chart, and the combination of the boxes in the block diagram or flow chart can be implemented with a dedicated hardware-based system that performs a specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.
[0127] The modules involved in the embodiments of the present invention may be implemented by software or hardware. The modules described may also be set in a processor, for example, it may be described as: a processor includes a frame sampling module, a foreground detection module, a background module and a filling module. The names of these modules do not constitute a limitation on the modules themselves in some cases. For example, the frame sampling module may also be described as "a module for sampling key frames of the video to be processed based on a sampling threshold to obtain a key frame set".
[0128] As another aspect, the present invention also provides a computer-readable medium, which may be included in the device described in the above embodiment; or it may exist independently and not be assembled into the device. The above computer-readable medium carries one or more programs. When the above one or more programs are executed by a device, the device includes: based on a sampling threshold, key frame sampling is performed on the video to be processed to obtain a key frame set; wherein the sampling threshold is less than the total number of frames of the video to be processed; foreground detection is performed on each key frame in the key frame set to determine the foreground target information in each key frame; based on foreground detection, a background template and a background library are obtained; wherein the background template is obtained based on the first frame data, and the background library includes the second frame data and its foreground target information; the key frame set includes the first frame data and the second frame data; according to the second frame data in the background library and its foreground target information, the background template is filled with background to obtain the background restoration data of the video to be processed.
[0129] In the embodiment of the present invention, only some key frames are extracted for calculation. The amount of calculation is greatly reduced, and the speed and efficiency are significantly improved. This advantage is more evident in longer videos, because the extracted key frames are still fixed, so the processing time does not change much, while the number of frames of the video increases as the length of time increases, and the processing time increases linearly. In addition, because the quality of the background restoration image processed by most traditional methods is strongly dependent on the moving rate of the foreground: that is, in a video, when the foreground object, such as a person, moves very slowly, or simply does not move, resulting in the foreground occupying most of the video frames in a fixed area, the effect of the traditional method is usually not good, and traces will be left, resulting in low quality. The embodiment of the present invention successfully overcomes the problem of uncertain background restoration quality of the existing method. Moreover, the embodiment of the present invention adopts a uniform sampling method, first of all, to ensure that the sampled frames are reasonably distributed, even if the person's speed is very slow, or the background appears for a short time, as long as there is a corresponding background pixel in the key frame, the background image can be restored. Because it is a direct slice copy, there is no problem of traces. .
[0130] The above specific implementations do not constitute a limitation on the protection scope of the present invention. It should be understood by those skilled in the art that various modifications, combinations, sub-combinations and substitutions may occur depending on design requirements and other factors. Any modification, equivalent substitution and improvement made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. A method for background restoration of video data, characterized in that: include: Based on a sampling threshold, key frame sampling is performed on the video to be processed to obtain a key frame set; wherein the sampling threshold is less than the total number of frames of the video to be processed; Performing foreground detection on each key frame in the key frame set to determine foreground target information in each key frame; Based on the foreground detection, a background template and a background library are obtained; wherein the background template is obtained according to the first frame data, and the background library includes the second frame data and its foreground target information; the key frame set includes the first frame data and the second frame data; According to the second frame data in the background library and its foreground target information, the background template is filled with background to obtain background restoration data of the video to be processed; Among them, based on the foreground detection, the steps of obtaining the background template and the background library include: selecting a frame data in the key frame set as the first frame data; performing foreground subtraction processing on the first frame data according to the foreground target information of the first frame data to obtain the background template; the second frame data other than the first frame data in the key frame set and its foreground target information form a background library, including: selecting a frame data from the second frame data of the background library as iterative frame data; according to the position of the cut-out part in the background template, verifying whether the target pixel point at the corresponding position of the iterative frame data belongs to the foreground target; if not, copying the target pixel point to the background module; otherwise, selecting a new frame data from the second frame data of the background library as iterative frame data again.
2. The method according to claim 1, characterized in that Based on the set sampling threshold, the step of sampling key frames of the video to be processed to obtain a key frame set includes: Obtaining a video to be processed, and determining the total number of frames of the video to be processed; Determine a sampling interval according to the total number of frames and a sampling threshold; wherein the sampling threshold is less than the total number of frames of the video to be processed; According to the sampling interval, key frame sampling is performed on the video to be processed to obtain a key frame set.
3. The method according to claim 1, characterized in that Before selecting a new frame data from the second frame data of the background library as the iterative frame data again, the method further includes: In the case of determining that there are no unfilled pixels in the removed pixel points of the background template, determining whether all the second frame data of the background library are selected as iterative frame data; If yes, the background template is filled according to the surrounding pixels of the unfilled pixels; otherwise, a new frame data is selected from the second frame data of the background library as the iterative frame data.
4. The method according to claim 1, characterized in that: The step of verifying whether the target pixel point at the corresponding position of the iterative frame data belongs to the foreground target includes: Creating a binary matrix mask of the iterative frame data; According to the binary matrix mask, it is verified whether the target pixel point at the corresponding position of the iterative frame data belongs to the foreground target.
5. The method according to claim 1, characterized in that The step of performing foreground detection on each key frame in the key frame set to determine foreground target information in each key frame includes: Based on a deep learning target detection algorithm, performing foreground target detection on each key frame in the key frame set; The position information of the foreground object in each key frame is determined, wherein the foreground object information at least includes the position information of the foreground object.
6. The method according to claim 5, characterized in that The position information is the target frame information of the foreground object; wherein the target frame information at least includes the length, width and diagonal pixel coordinates.
7. A device for restoring the background of video data, characterized in that: include: A frame sampling module, used for sampling key frames of the video to be processed based on a sampling threshold to obtain a key frame set; wherein the sampling threshold is less than the total number of frames of the video to be processed; A foreground detection module, used to perform foreground detection on each key frame in the key frame set to determine foreground target information in each key frame; A background module, used to obtain a background template and a background library based on the foreground detection; wherein the background template is obtained according to the first frame data, the background library includes the second frame data and its foreground target information; the key frame set includes the first frame data and the second frame data; A filling module, used to fill the background template with the second frame data in the background library and its foreground target information, so as to obtain the background restoration data of the video to be processed; The background module is also used to select a frame data in the key frame set as the first frame data; perform foreground subtraction processing on the first frame data according to the foreground target information of the first frame data to obtain a background template; and form a background library with the second frame data and its foreground target information in the key frame set except the first frame data; The filling module is also used to select a frame data from the second frame data of the background library as iterative frame data; based on the position of the cut-out part in the background template, verify whether the target pixel point at the corresponding position of the iterative frame data belongs to the foreground target; if not, copy the target pixel point to the background module; otherwise, select a new frame data from the second frame data of the background library again as iterative frame data.
8. The device according to claim 7, characterized in that The frame sampling module is also used to obtain the video to be processed and determine the total number of frames of the video to be processed; determine the sampling interval according to the total number of frames and a sampling threshold; wherein the sampling threshold is less than the total number of frames of the video to be processed; and perform key frame sampling on the video to be processed according to the sampling interval to obtain a key frame set.
9. The device according to claim 7, characterized in that The filling module is also used to determine whether all the second frame data of the background library are selected as iterative frame data when it is determined that there are no unfilled pixels in the removed pixel points of the background template; If yes, the background template is filled according to the surrounding pixels of the unfilled pixels; otherwise, a new frame data is selected from the second frame data of the background library as the iterative frame data.
10. The device according to claim 7, characterized in that The filling module is further used to create a binary matrix mask of the iterative frame data; and verify whether the target pixel point at the corresponding position of the iterative frame data belongs to the foreground target according to the binary matrix mask.
11. The device according to claim 7, characterized in that The foreground detection module performs foreground target detection on each key frame in the key frame set based on a deep learning target detection algorithm; The position information of the foreground object in each key frame is determined, wherein the foreground object information at least includes the position information of the foreground object.
12. The device according to claim 11, characterized in that The position information is the target frame information of the foreground object; wherein the target frame information at least includes the length, width and diagonal pixel coordinates.
13. An electronic device, characterized in that: include: one or more processors; a storage device for storing one or more programs, When the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1 to 6.
14. A computer readable medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the method according to any one of claims 1 to 6 is implemented.
Citation Information
Patent Citations
Embedded target detection algorithm
CN103049919A
Representation frame acquisition method and representation frame acquisition apparatus
CN105516735A