Method and Apparatus for Processing Image in Compressed Domain
Patent Information
- Application Number
- KR1020190033708
- Authority / Receiving Office
- KR · KR
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2019-03-25
- Publication Date
- 2026-09-22
- Estimated Expiration
- 2039-03-25
Smart Images

Figure R1020190033708_ABST
Abstract
Description
Technology Field
[0001] The present embodiment relates to an image processing method and apparatus in a compression area. More specifically, it relates to an image processing method and apparatus for more efficiently generating image summary information for very long images, such as CCTV images. Background Technology
[0002] The content described in this section merely provides background information regarding the present embodiment and does not constitute prior art.
[0003] Figure 1 is an example diagram illustrating the configuration of a typical video stream. Referring to Figure 1, the video stream has IDR pictures at regular intervals, and the unit of pictures from one IDR picture to the next is called a Group of Picture (GOP). Although IDR pictures are larger than b / p pictures, they have the advantage of being easy to store because a single picture can form a complete screen. On the other hand, they have the disadvantage that, due to their large size, they consume a lot of bandwidth during network transmission and take up a lot of space during storage.
[0004] Figure 2 is a diagram illustrating a procedure for generating thumbnails for video summarization from a conventional video stream. Conventionally, when generating thumbnails for video summarization from a video stream, a method was used to store IDR pictures that are periodically included in the video stream. However, this method is not suitable for storing very long video streams, such as CCTV footage. In particular, when storing a large number of cloud-based sessions, it results in a significant waste of space. That is, because the intervals between IDR pictures in CCTV video streams are wide, thumbnails are generated at intervals that are long for humans to perceive when using the conventional method. Furthermore, since a large number of meaningless pictures are generated among the periodically created pictures, there is a limitation in that storage space is wasted significantly when storing very long periods, such as CCTV footage.
[0005] Accordingly, there is a need for a new technology that can improve the quality of video summaries while reducing storage space usage by frequently saving meaningful thumbnail images when saving video streams transmitted from CCTV cameras and simultaneously saving thumbnails for video summaries. The problem to be solved
[0006] The present embodiment aims to provide a video processing method and apparatus in a compression area that can improve the quality of video summaries while reducing storage space usage by analyzing a video stream transmitted from a CCTV camera in a compression area to detect the movement of an object and generating thumbnail images for video summaries based on this, thereby storing meaningful pictures at intervals suitable for human perception. means of solving the problem
[0007] The present embodiment provides an image processing device characterized by comprising: an input buffer unit that receives an encoded video stream of an image; an extraction unit that analyzes the encoded video stream in a compressed domain to extract motion vector information and DCT coefficients within the encoded video stream; an analysis unit that provides an analysis result of analyzing the motion vector information and the DCT coefficients; and a motion detection unit that detects the movement of an object within the image based on at least one analysis result among the motion vector information and the DCT coefficients.
[0008] In addition, according to another aspect of the present embodiment, an image processing method is provided, characterized by comprising: a process of receiving an encoded video stream of an image; a process of analyzing the encoded video stream in a compression area to extract motion vector information and DCT coefficients within the encoded video stream; a process of providing an analysis result of analyzing the motion vector information and the DCT coefficients; and a process of detecting the movement of an object in the image based on at least one analysis result among the motion vector information and the DCT coefficients.
[0009] In addition, according to another aspect of the present embodiment, a computer program stored on a computer-readable recording medium is provided to execute each step of the image processing method according to claim 12. Effects of the invention
[0010] According to the present embodiment, by analyzing a video stream transmitted from a CCTV camera in a compression area to detect the movement of an object and generating thumbnail images for video summarization based on this, meaningful pictures are stored at intervals suitable for human perception, thereby improving the quality of video summarization while reducing the use of storage space. Brief explanation of the drawing
[0011] Figure 1 is an example diagram illustrating the configuration of a typical video stream. Figure 2 is a diagram illustrating a procedure for generating thumbnails for video summaries from a conventional video stream. FIG. 3 is a block diagram schematically showing an image processing device according to the present embodiment. FIG. 4 is a block diagram schematically showing the execution unit of an image processing device according to the present embodiment. FIG. 5 is a flowchart for explaining an image processing method according to the present embodiment. Specific details for implementing the invention
[0012] Hereinafter, the present embodiment will be described in detail with reference to the attached drawings.
[0013] Recently, there has been an increase in applications that summarize and display videos in thumbnail format. The lengths of the videos to be summarized vary, and for very long videos, such as CCTV footage, usability and storage space are significantly affected by the storage method. CCTV video streams have characteristics different from general video streams. Specifically, to reduce network bandwidth and storage space, a wide GOP (Group of Per Second) is used, and by lowering the frame rate, the intervals between IDR (In-Video Recording) pictures are significantly widened. For example, in the case of 60 GOP and 15 fps, an IDR picture is included once every 4 seconds. On the other hand, when the fps is low and the GOP is high, as in CCTV video streams, IDR is included only once every few seconds, making it highly likely that the thumbnail will miss important scenes.
[0014] In addition, when saving video streams and thumbnail images for video summaries simultaneously, if images are periodically extracted from the video stream, all pictures must be decoded from the video, which consumes significantly more CPU resources, and when saving for a long time, meaningless images are also continuously saved, consuming a large amount of storage space.
[0015] Meanwhile, since the motion vector information of the video stream of the CCTV footage contains movement information for a moving object while the background is stationary, if a motion vector is detected, it is highly likely to be meaningful information. Based on this, the present embodiment proposes a method to analyze the video stream transmitted from the CCTV camera in the compression area to detect the movement of an object through motion vectors, and to generate thumbnail images for video summarization based on this, thereby allowing meaningful pictures to be stored at intervals suitable for human perception, thereby improving the quality of video summarization while reducing the use of storage space.
[0016] FIG. 3 is a block diagram schematically showing an image processing device according to the present embodiment.
[0017] As illustrated in FIG. 3, the image processing device (300) according to the present embodiment includes an input buffer unit (310), an execution unit (320), and an image processing unit (330). Here, the components included in the image processing device (300) according to the present embodiment are not necessarily limited thereto. That is, FIG. 3 illustrates only the essential components for detecting object movement through motion vector extraction in a compressed area and generating thumbnail images for image summarization based thereon, and it should be recognized that such an image processing device (300) may have more or fewer components or different configurations of components than those illustrated to implement other functions.
[0018] The input buffer unit (310) performs the function of buffering a video stream of an image generated from a video recording device. Meanwhile, in the present embodiment, the video recording device may preferably be an intelligent surveillance device such as a CCTV device. Accordingly, the video stream buffered by the input buffer unit (310) according to the present embodiment uses a wide GOP unit and has a low frame rate, and thus the spacing of the IDR picture may be formed widely.
[0019] The input buffer unit (310) according to the present embodiment can receive a video stream encoded according to a preset compression standard from a video capturing device. For example, the input buffer unit (310) can receive a video stream encoded according to a compression standard such as H.264 / H.265 based on block-unit motion compensation.
[0020] Meanwhile, in this embodiment, b / p pictures that are not IDR pictures within the encoded video stream must be decoded from the leading IDR picture to obtain a raw image, and then compressed into a jpeg / png image. Accordingly, the input buffer unit (310) can receive the above encoded video stream, which is buffered in units of Group of Pictures (GOP) defined by the Instantaneous Decoder Refresh (IDR) picture, from the video capturing device.
[0021] The execution unit (320) performs the function of detecting meaningful scenes from the encoded video stream. Meanwhile, in the present embodiment, the video generated from the video capturing device has the characteristic that the background is static and thus the video has little variability as a result of monitoring a designated specific area. Consequently, when movement of an object within the video occurs, it can be classified as a meaningful scene.
[0022] Based on this, the execution unit (320) according to the present embodiment analyzes the encoded video stream to detect the movement of an object and provides the result, thereby enabling video summarization of important events to be performed during the video processing process of the video processing unit (330) thereafter.
[0023] The execution unit (320) can detect motion vector information within the encoded video stream and detect the movement of an object based thereon. That is, in the present embodiment, since the motion vector information within the encoded video stream has a background that is stationary and contains movement information for a moving object, if a motion vector is detected, it is highly likely to be meaningful information.
[0024] The execution unit (320) according to the present embodiment analyzes the encoded video stream in the compressed domain. That is, the execution unit (320) detects whether motion occurs by analyzing the motion vector area in the encoded state without decompressing the encoded video stream. Generally, in the case of media streams, a large amount of CPU consumption occurs during the decompression process, but in the case of the present embodiment, this can be reduced by analyzing in the compressed state.
[0025] In addition, in this embodiment, the execution unit (320) may additionally detect and utilize Discrete Cosine Transform (DCT) coefficients within the encoded video stream to improve the accuracy of detecting object movement through motion vector information. Meanwhile, DCT coefficients are an abbreviation for Discrete Cosine Transform, which refers to a method of reducing the amount of data by expressing n data points as the sum of n cosine functions. This method is fundamentally derived from the Fourier transform, which states that n values in the time domain can be expressed as n trigonometric functions in the frequency domain. In MPEG, an international standard image compression method, video is encoded in units of 8x8, or 64 blocks; expressing these 64 data points as 64 cosine functions is called DCT, and the coefficients of these transformed cosine functions are called DCT coefficients. By using these DCT coefficients, the difference between the preceding and succeeding frames of the video can be extracted.
[0026] In this embodiment, the execution unit (320) utilizes the DCT coefficient as a verification parameter in verifying the validity of the detected motion vector information.
[0027] Hereinafter, with reference to FIG. 4, the specific operation of the execution unit (320) according to the present embodiment will be described. Meanwhile, FIG. 4 (a) is a diagram illustrating the procedure for detecting the movement of an object in a pixel area. FIG. 4 (b) is a block diagram schematically showing the execution unit (320) according to the present embodiment, illustrating the procedure for detecting the movement of an object through motion vector information in a compression area.
[0028] First, referring to Figure 4(a), when detecting the motion of an object in a pixel area, resource-intensive operations such as reordering, IDCT, and motion compensation must be performed. Additionally, since the motion detection of an object in a pixel area requires computation for every pixel, a significant amount of resources are consumed in this process as well.
[0029] On the other hand, referring to Fig. 4(b), in the case of this embodiment, motion detection of an object is performed in a compressed area, so motion detection can be performed while skipping parts that require a large amount of computation (parts indicated by dotted lines in Fig. 4(a)), thus having the effect of reducing resource waste.
[0030] As illustrated in FIG. 4(b), the execution unit (320) according to the present embodiment includes an extraction unit (400), an analysis unit (410), and a motion detection unit (420). Here, the components included in the execution unit (320) are not necessarily limited thereto.
[0031] The extraction unit (400) performs the function of analyzing the encoded video stream in the compression area to extract motion vector information and DCT coefficients within the encoded video stream. This operation of the extraction unit (400) corresponds to a process of decoding a part of the encoded video stream, rather than the whole, to extract only the necessary information for motion detection.
[0032] To this end, according to the present embodiment, the extraction unit (400) may be implemented as an entropy decoder configured to perform entropy decoding on an encoded video stream. That is, the extraction unit (400) distinguishes between cases where the encoded video stream is encoded based on CAVLC (Context-adaptive variable-length coding) and cases where it is encoded based on CABAC (Context-adaptive binary arithmetic coding), and based on this, performs different decoding methods to extract motion vector information and DCT coefficients.
[0033] The analysis unit (410) receives motion vector information and DCT coefficients extracted through the extraction unit (400), analyzes them, and performs the function of providing analysis results for each. That is, in this embodiment, the analysis unit (410) may be implemented by including a first analysis unit (Motion vector analysis, 412) and a second analysis unit (DCT coefficient analysis).
[0034] The first analysis unit (412) performs an analysis of motion vector information and provides the analysis results. In this embodiment, the first analysis unit (412) can perform correction on the motion vector information based on the analysis results of the motion vector information. Meanwhile, the motion vector information contains vector information that finds the most similar macroblock by comparing it with the previous image in macroblock units. This can be 16×16, 8×8, 4×4, etc. When motion vector information is generated in this way, motion may be detected, but it must be corrected for various reasons. That is, motion vector information is found through the motion estimation process, but since blocks favorable for compression are found at this time, the motion vector information may not match the actual motion. Also, in the case of blocks that process by intra-prediction within their own blocks rather than generating motion vector information, motion vector information is not generated. In addition, there may be cases where motion vector information is incorrectly found in areas where the pattern is usually uniform or has almost no features.
[0035] Based on this, the first analysis unit (412) can adaptively assign motion vector information to a specific block that performs processing based on intra prediction among the surrounding blocks of a block that matches the motion vector information through analysis of the motion vector information. This corresponds to a correction process to improve the accuracy of motion detection within an object, and for example, the first analysis unit (412) can assign motion vector information to a specific block based on the motion vector information of the surrounding blocks.
[0036] The first analysis unit (412) transmits the corrected motion vector information and the above analysis results to the motion detection unit (420), thereby enabling the motion detection unit (420) to perform a motion detection procedure.
[0037] The second analysis unit (414) performs an analysis of the DCT coefficients and provides the results of the analysis. The DCT coefficients are coefficients that are generated after performing a discrete cosine transform on the pixel values of the difference from the previous screen for the macroblocks, and can be used for two purposes.
[0038] First, the DCT coefficient can be used as a variable value to detect the movement of an object independently of motion vector information. That is, if movement is detected, there will be a change in the DCT coefficient, and if this amount of change exceeds a specific threshold, it can be said that movement has been detected. To this end, the second analysis unit (414) can provide the amount of change in the DCT coefficient as the above analysis result.
[0039] Additionally, the DTC coefficients can be used to verify the validity of the motion vector information analyzed and corrected through the first analysis unit (212). That is, since the approximate actual pattern of the image can be known through the DCT coefficients, it is possible to determine whether the motion vector information is valid. The second analysis unit (214) transmits the analysis result of the DCT coefficients to the motion detection unit (420), thereby allowing the motion detection unit (420) to determine the validity of the motion vector information through the analysis result of the DCT coefficients.
[0040] The motion detection unit (420) performs the function of detecting the movement of an object in an image based on at least one analysis result among motion vector information and DCT coefficients provided through the analysis unit (210).
[0041] In this embodiment, the motion detection unit (420) can determine whether the motion of an object has been detected by integrating the analysis results of the motion vector information and the DCT coefficients. That is, the motion detection unit (420) performs correction by reflecting the analysis results of the DCT coefficients in the analysis results of the motion vector information.
[0042] In this case, the motion detection unit (420) can determine the validity of motion vector information by utilizing the analysis result of the DCT coefficients and detect the movement of an object according to the determination result. For example, if the weighted value reflecting the predefined weights for the analysis result of the motion vector information and the analysis result of the DCT coefficients corresponding to the motion vector information exceeds a pre-set threshold, the motion detection unit (420) determines that the motion vector information is valid and determines that the movement of an object has been detected. Conversely, if the above weighted value is less than the pre-set threshold, the motion detection unit (420) determines that the motion vector information is invalid and determines that the movement of an object has not been detected.
[0043] In another embodiment, the motion detection unit (420) calculates the amount of change of the DCT coefficient based on the analysis result of the DCT coefficient and can detect the movement of an object based on the calculated amount of change. For example, the motion detection unit (420) determines that the movement of an object has been detected if the calculated amount of change exceeds a preset threshold.
[0044] In another embodiment, the motion detection unit (420) receives a result of recognizing a face or person from a video recording device and can detect the movement of an object based on this.
[0045] The image processing unit (330) generates image summary information based on the object motion detection result of the motion detection unit (420). In this embodiment, the image summary information is preferably a thumbnail image, but is not necessarily limited thereto. The image processing unit (330) may be configured to include a video decoder, an image resize unit, an image decoder, and first and second processing units (Black Hole). Here, the operation of each component included in the image processing unit (330) is the same as in the past, so a detailed description is omitted.
[0046] The image processing unit (330) stores timestamps for scenes in which the movement of an object in the image was detected in a table based on the object movement detection results.
[0047] The image processing unit (330) generates image summary information by storing an image corresponding to a specific point in time based on a stored timestamp. For example, the image processing unit (330) decodes a encoded video stream, performs re-encoding for a time including the time corresponding to the timestamp within the decoded video stream, and then stores it as an image.
[0048] Meanwhile, the image processing unit (330) performs the above re-encoding by additionally combining information about the reference image (e.g., IDR picture) within the decoded video stream when the scene in which the movement of an object in the image is detected is a b / p picture rather than an IDR picture. This has the effect of enabling the decoding of the image summary information without decoding the reference image when subsequently querying the image summary information.
[0049] The image processing unit (330) can change the generation cycle of image summary information as needed. For example, if the image processing unit (330) generates images too frequently, it consumes a lot of CPU and storage space, so it saves them at the most appropriate interval. For example, if the video stream is 15fps, the interval between pictures is 67ms, but when saving as images, saving about one picture per second allows for saving storage space while ensuring that the images do not look awkward to humans when viewed qualitatively.
[0050] FIG. 5 is a flowchart for explaining an image processing method according to the present embodiment.
[0051] The image processing device (300) receives an encoded video stream of an image generated from an image capturing device (S502). In step S502, the image processing device (300) may receive the above encoded video stream, which is buffered in GOP units defined by an IDR picture, from the image capturing device.
[0052] The image processing device (300) analyzes the encoded video stream of step S502 and extracts motion vector information and DCT coefficients within the encoded video stream (S504). In step S504, the image processing device (300) extracts the motion vector region and DCT coefficients in the encoded state without decompressing the encoded video stream. That is, the image processing device (300) performs entropy decoding on the encoded video stream, and thereby can extract motion vector information and DCT coefficients by decoding a part of the encoded video stream rather than the whole of it.
[0053] The image processing device (300) analyzes the motion vector information and DCT coefficients extracted in step S504 to generate an analysis result (S506). In step S506, the image processing device (300) can perform correction on the motion vector information based on the analysis result of the motion vector information. For example, through the analysis of the motion vector information, the image processing device (300) can adaptively assign motion vector information to a specific block that performs processing based on intra-prediction among the surrounding blocks of a block that matches the motion vector information.
[0054] The image processing device (300) detects the movement of an object in the image based on at least one of the analysis results of the motion vector information generated in step S506 and the analysis results of the DCT coefficients (S508). In step S508, the image processing device (300) determines the validity of the motion vector information by utilizing the analysis results for the DCT coefficients, and can detect the movement of the object according to the determination result.
[0055] Additionally, the image processing device (300) can calculate the amount of change of the DCT coefficient based on the analysis result of the DCT coefficient and detect the movement of an object based on the calculated amount of change.
[0056] The image processing device (300) generates image summary information based on the motion detection result of step S508 (S510). Based on the motion detection result of step S508, the image processing device (300) stores a timestamp for the scene in which the movement of an object in the image was detected in a table, and generates image summary information by storing an image corresponding to a specific point in time including the time of the stored timestamp.
[0057] Here, steps S502 to S510 correspond to the operation of each component of the image processing device (300) described above, so further detailed description is omitted.
[0058] Although FIG. 5 describes each process as being executed sequentially, it is not necessarily limited to this. In other words, FIG. 5 is not limited to a chronological order, as it may be applicable to execute the processes described in FIG. 5 by modifying them or to execute one or more processes in parallel.
[0059] As described above, the image processing method described in FIG. 5 can be implemented as a program and recorded on a recording medium (CD-ROM, RAM, ROM, memory card, hard disk, magneto-optical disk, storage device, etc.) that can be read using computer software.
[0060] The above description is merely an illustrative explanation of the technical concept of the present embodiment, and a person skilled in the art to which the present embodiment belongs would be able to make various modifications and variations within the scope of the essential characteristics of the present embodiment. Accordingly, the present embodiments are intended to explain, not limit, the technical concept of the present embodiment, and the scope of the technical concept of the present embodiment is not limited by these embodiments. The scope of protection of the present embodiment shall be interpreted by the claims below, and all technical concepts within an equivalent scope shall be interpreted as being included within the scope of rights of the present embodiment. Explanation of the symbols
[0061] 300: Image processing unit 310: Input buffer section 320: Execution unit 330: Image processing unit 400: Extraction unit 410: Analysis unit 420: Motion detection unit
Claims
Claim 1 An image processing device comprising: an input buffer unit receiving an encoded video stream of an image; an extraction unit that analyzes the encoded video stream in a compressed domain to extract motion vector information and DCT coefficients within the encoded video stream; an analysis unit that provides an analysis result of analyzing the motion vector information and the DCT coefficients; and a motion detection unit that detects the movement of an object in the image based on the analysis result of the motion vector information and the DCT coefficients, wherein the analysis unit adaptively assigns motion vector information to a specific block that performs processing based on intra-prediction among the surrounding blocks of a block that matches the motion vector information through analysis of the motion vector information, and the motion detection unit determines the validity of the motion vector information by utilizing the analysis result of the DCT coefficients and detects the movement of the object according to the determination result. Claim 2 An image processing device according to claim 1, wherein the input buffer receives a video stream encoded based on a compression standard based on block-unit motion compensation. Claim 3 An image processing device according to claim 1, wherein the input buffer unit receives the encoded video stream that is buffered in units of Group of Pictures (GOP) defined by Instantaneous Decoder Refresh (IDR) pictures. Claim 4 An image processing device according to claim 1, wherein the analysis unit comprises a first analysis unit that performs analysis on the motion vector information and a second analysis unit that performs analysis on the DCT coefficients. Claim 5 delete Claim 6 An image processing device according to claim 1, wherein the motion detection unit determines that the motion vector information is valid when the weighted value, which reflects a predefined weight for the analysis result of the motion vector information and the DCT coefficient corresponding to the motion vector information, exceeds a pre-set threshold. Claim 7 delete Claim 8 An image processing device according to claim 1, further comprising an image processing unit that generates image summary information based on the detection result of the motion detection unit. Claim 9 In claim 8, the image processing unit is characterized by storing a timestamp for a scene in which the movement of the object within the image is detected, and storing an image corresponding to a specific point in time based on the timestamp to generate the image summary information. Claim 10 An image processing device according to claim 9, wherein the image processing unit decodes the encoded video stream, performs re-encoding for a time including a time corresponding to the timestamp within the decoded video stream, and then stores it as the image. Claim 11 An image processing device according to claim 10, wherein the image processing unit further combines information regarding a reference image within the decoded video stream to perform the re-coding. Claim 12 A video processing method comprising: a process of receiving an encoded video stream of a video; a process of analyzing the encoded video stream in a compression area to extract motion vector information and DCT coefficients within the encoded video stream; a process of providing an analysis result of analyzing the motion vector information and the DCT coefficients; and a process of detecting the movement of an object in the video based on the analysis result of the motion vector information and the DCT coefficients, wherein the process of providing the analysis result adaptively assigns motion vector information to a specific block that performs processing based on intra-prediction among the surrounding blocks of a block that matches the motion vector information through analysis of the motion vector information, and the process of detecting the movement of an object in the video is characterized by determining the validity of the motion vector information using the analysis result of the DCT coefficients and detecting the movement of the object according to the determination result. Claim 13 A computer program stored on a computer-readable recording medium to execute each step of the image processing method according to Clause 12.
Citation Information
Patent Citations
Imaging apparatus, its control method, and computer program
JP2010010967A
Method and apparatus for shot conversion detection ofvideo encoder
KR1020030082794A
Method and system for encoding by using image analyzer and image analyzer therefor
KR1020090042047A