Method, computer program product, and computer-readable medium for generating a mask for a camera stream

By accumulating differential images and edge images in the camera stream, generating combined images and performing thresholding processing, the problem of difficult to generate masks in real time and efficiently, realizing accurate identification of non-static areas and mask generation, which is robust and efficient.

CN114127784BActive Publication Date: 2025-06-24AIMOTIVE KFT
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN201980098438.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2019-07-17
Filing Date
2019-12-03
Publication Date
2025-06-24
Estimated Expiration
2039-12-03

AI Technical Summary

Technical Problem

The prior art is difficult to efficiently generate masks covering non-static areas in real time when processing camera flow, especially in applications such as real-time signal processing and autonomous vehicles, where there are problems of wasted computing resources and unstable lighting changes.

Method used

By a method of accumulating differential images and cumulative edge images, a combined image is generated, and a mask covering only the region of interest is generated in real time through thresholding and mask generation algorithms.

Benefits of technology

It realizes efficient identification and mask generation of non-static areas in the camera stream, saves time and resources for image processing, and has robustness for lighting changes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114127784B_ABST
    Figure CN114127784B_ABST
Patent Text Reader

Abstract

The present invention is a method for generating a mask (40) for non-static regions based on a camera stream having successive images, comprising the steps of: - generating a cumulative difference image by accumulating difference images, each difference image being obtained by subtracting two images of the camera stream from one another, - generating a cumulative edge image by accumulating edge images, each edge image being obtained by detecting edges in respective images of the camera stream, - generating a combined image (30) by combining the cumulative edge image and the cumulative difference image, defining a first threshold pixel value for the combined image (30), and - generating the mask (40) by including in the mask pixels of the combined image (30) pixels having the same relationship to the first threshold pixel value. The present invention also relates to a computer program product and a computer-readable medium for implementing the above method.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a method for generating a mask for non-static regions of a camera stream. The present invention also relates to a computer program product and a computer-readable medium for implementing the method. Background Art

[0002] Processing multiple images of a camera stream typically involves subtracting consecutive images or frames in order to distinguish the stationary parts and the dynamic or changing parts of the images. Most applications (such as those related to autonomous vehicles) typically need to ignore the stationary parts of the images, resulting in faster and more efficient processing of the dynamically changing parts.

[0003] A method using image subtraction and dynamic thresholding is disclosed in US 6061476. Images are taken before and after applying solder paste on a printed circuit board, and then the before and after images are subtracted to inspect the applied solder paste. The foreground and background of the subtracted image are separated by a dynamic threshold instead of a scalar threshold, and both the negative subtracted image and the positive subtracted image are further processed. The dynamic threshold is considered to be able to detect the foreground more precisely than the scalar threshold. The dynamic threshold processing results in subtracting the positive image and the negative image, and then these images are combined and binarized. Edge detection is performed on the before and after images by a Sobel edge detector, and then true peak detection is carried out to eliminate false edges. Then the resulting image is binarized by a binarization map. The pixels corresponding to the solder paste are determined by subtracting the binarized before and after images.

[0004] The disadvantage of the above method is that due to the complex edge detection algorithm, it is not suitable for real-time signal processing.

[0005] A method for extracting a foreground object image from a video stream is disclosed in US 2011 / 0164823 A1. The method includes separating the background and the foreground of the image by calculating the edges of a frame and a reference background image. The foreground is extracted by subtracting the edges of the frame and the reference background image. The method may also include thresholding to remove noise in the image. The disadvantage of this method is that it requires a reference background image that may be difficult to provide in some applications. For example, the dashboard camera of an autonomous or self-driving vehicle may not be able to capture the reference background image required by this method.

[0006] In US 2012 / 0014608 A1, an apparatus and method for image processing are disclosed. Based on feature analysis of an image, a region of interest (ROI) is determined, and a mask is created for the ROI. According to the present invention, the ROI can be more precisely specified by detecting the edges of the image, because the features to be detected by this method have continuous edges that can be extracted by an edge detection algorithm. The ROI masks generated by different methods are synthesized to generate a mask for the "region of most interest". However, the area covered by the created mask is not limited to the features of interest, but also includes some of the surrounding environment, because the generated mask has a rectangular shape. Therefore, the above method only ensures that the features of interest in the image are part of the mask, but the mask can include additional image portions. This will result in higher data storage requirements, which is not conducive to real-time evaluation of the image.

[0007] In view of the known methods, there is a need for a method by which a mask covering a non-static region of a camera stream can be generated in a more efficient manner than the prior art methods, in order to enable real-time mask generation for the camera stream.

[0008] Description of the Invention

[0009] The main object of the present invention is to provide a method for generating a mask for a non-static camera stream, which is free of the drawbacks of the prior art methods to the greatest possible extent.

[0010] The object of the present invention is to provide a method by which a mask can be created in a more efficient manner than the prior art methods, so as to be able to identify non-static regions of a camera image. Therefore, the object of the present invention is to provide a fast mask generation method that can generate a mask in real time, so as to be able to filter out regions of no interest, mainly static regions, from the camera stream.

[0011] Furthermore, the object of the present invention is to provide a non-transitory computer program product for implementing the steps of the method according to the present invention on one or more computers and a non-transitory computer-readable medium including instructions for executing the steps of the method on one or more computers.

[0012] The object of the present invention can be achieved by the method as described in claim 1. Preferred embodiments of the present invention are defined in the dependent claims. The object of the present invention can be further realized by the non-transitory computer program product according to claim 16 and the non-transitory computer-readable medium according to claim 17.

[0013] Compared with existing technology methods, the main advantage of the method according to the present invention comes from the fact that it generates a mask that only covers the region of interest (i.e., the non-static part of the camera image). The size and shape of the mask correspond to the actual area and shape of the region of interest, so any other pixels can be avoided from being processed. This can save time and resources for further processing of the image.

[0014] Another advantage of the method according to the present invention is that the method is robust to illumination changes because the edges and dynamic parts of the image can always be detected. Therefore, the method can be used in any illumination and weather conditions and is not affected by illumination changes. Even if shadows fall on the substantially static regions of the image, it will not affect mask generation because mask generation starts from the middle of the image, so mask generation will have been completed before reaching these regions because the boundary between the static and non-static regions is always closer to the image center than these regions.

[0015] It has been recognized that having an additional edge detection step provides more feature edges between the static and non-static image parts, enabling the non-static image parts for mask generation to be separated faster and more securely. The provided feature edges can also generate the mask safely and accurately through a simple algorithm such as ray marching. Therefore, contrary to the greater computational requirements due to the additional edge detection step, the overall computational requirements of the method can be reduced by simpler mask generation, and at the same time the generated mask will be more precise.

[0016] The method according to the present invention is capable of generating a mask in real time, so the method can be applied to images recorded by cameras of autonomous or self-driving vehicles, where real-time image processing and discarding static regions from the images to save computational resources are particularly important.

[0017] In the method according to the present invention, no further information or input is required other than the images from the camera stream.

[0018] Brief Description of the Drawings

[0019] The preferred embodiments of the present invention will be described below by way of example with reference to the drawings, where:

[0020] Figure 1 is a flowchart of the steps of an embodiment of the method for generating a cumulative difference image according to the present invention,

[0021] Figure 2 is a flowchart of further steps of an embodiment of the method for generating a cumulative edge image according to the present invention, and

[0022] Figure 3 is a flowchart of even further steps of an embodiment of the method for generating a mask from the cumulative difference image and the cumulative edge image according to the present invention.

[0023] Mode for implementing the invention

[0024] The present invention relates to a method for generating a mask for non-static regions of a camera stream. The camera stream is provided by a camera and has consecutive images 10, preferably a sequence of consecutive images 10. The images 10 are preferably recorded at a given frequency. The camera may be mounted on the dashboard of a car or other vehicle, or on the dashboard of an autonomous or self-driving car or other vehicle. The field of view of a dashboard camera typically covers parts of the vehicle itself, such as parts of the dashboard or window frames, which are static and unchanging parts of the recorded images 10. These parts are outside the region of interest or extent of the images 10, and thus it is beneficial to exclude them from further processing by generating a mask for the region of interest or extent (including the non-static, changing parts of the camera stream).

[0025] The camera may be configured to record color (e.g., RGB) images or grayscale images, and thus the images 10 of the camera stream may be color images or grayscale images.

[0026] The method according to the present invention comprises the following steps, as Figures 1-3 illustrated in

[0027] - Generating a cumulative difference image 14 by accumulating difference images 12, each difference image 12 being obtained by subtracting two images of the camera stream from each other,

[0028] - Generating a cumulative edge image 26 by accumulating edge images 24, each edge image 24 being obtained by detecting edges in respective images 10 of the camera stream,

[0029] -- Generating a combined image 30 by combining the cumulative edge image 26 and the cumulative difference image 14,

[0030] - Defining a first threshold pixel value for the combined image 30, and

[0031] -- Generating a mask 40 by including in the mask 40 of the combined image 30 pixels having the same relationship to the first threshold pixel value.

[0032] An embodiment of the present invention includes the step of acquiring images 10 of the camera stream. After acquiring the images 10 of the camera stream, the steps of a preferred embodiment of the method as Figure 1 illustrated in include generating a cumulative difference image 14 from the camera stream. The camera stream has consecutive images 10, where the consecutive images 10 may be images 10 recorded directly one after another, or images 10 recorded some time after each other (with other images 10 in between).

[0033] In step S100, a difference image 12 is generated by subtracting two consecutive images 10 from each other, where one of the consecutive images 10 is the acquired (preferably current) image 10. Preferably, the image 10 recorded later is subtracted from the previously recorded image 10. The difference image 12 has a lower value for static, unchanging image parts than for non-static, dynamically changing image parts.

[0034] The difference image 12, preferably consecutive difference images 12, are summed in step S110 in order to generate a cumulative difference image 14 that enhances the difference between the static and non-static parts of the image.

[0035] In a preferred embodiment of the method, a previously cumulative difference image 14 is provided, for example, from a previous step of implementing the method, where the previously cumulative difference image(s) 14 is preferably stored by a computer, and the cumulative difference image 14 is generated by adding the difference image 12 to the previously cumulative difference image 14.

[0036] In a preferred embodiment of the method, the cumulative difference image 14 is normalized in a first normalization step S120 by dividing each pixel value by the number of cumulative images. This optional step thus results in a normalized cumulative difference image 16 having pixel values of a normal color (RGB) or grayscale image. Thus, the first normalization in step S120 will prevent the accumulation of pixel values from exceeding a limit and avoid computational problems such as overflow. In the case where the camera records color images, in the case of an RBG image for the red, blue, and green color channels, the first normalization step S120 is performed separately for each color channel.

[0037] If the camera stream includes color images, the method preferably includes a first conversion step S130 for converting the color images to grayscale images. The first conversion step S130 can be performed before the subtraction step S100, between the subtraction step S100 and the summation step S110, between the summation step S110 and the first normalization step S120, or after the first normalization step S120. On the one hand, if the first conversion step S130 is implemented before the summation step S110, the conversion must be performed on all images 10 or all difference images 12, resulting in higher computational requirements. On the other hand, if the first conversion step S130 is implemented after the summation step S110, the conversion only affects one image, i.e., either the cumulative difference image 14 or the normalized cumulative difference image 16. However, each step before the first conversion step S130 must be implemented on each color channel of the image. The first conversion step S130 can be implemented by any conversion method or algorithm known in the prior art.

[0038] In Figure 1In a preferred embodiment of the described method, a first conversion step S130 is performed on the normalized cumulative difference image 16, thereby obtaining a normalized gray-scale cumulative difference image 18.

[0039] According to the present invention, the cumulative difference image 14, or, in the case where the first normalization step S120 is implemented, the normalized cumulative difference image 16, or, in the case where the first conversion step S130 is implemented, the gray-scale cumulative difference image 18 is processed by further steps of the method according to the present invention as Figure 3 explained.

[0040] In addition to Figure 1 the steps explained, edge detection is also performed on the obtained images 10 of the camera stream to detect the edges of the static regions of the camera stream. Possible steps for edge detection of camera images are Figure 2 explained.

[0041] According to a preferred embodiment of the method, the obtained image 10 is blurred in step S200 in order to smooth the original image before edge detection. The exact method chosen for blurring is fitted to the implemented edge detection algorithm. The blurring step S200 can be implemented by any blurring method or algorithm known in the prior art, such as box blurring or calculating the histogram of the image 10. In a preferred embodiment of the method, the blurring step S200 is implemented by Gaussian blurring with a kernel, more preferably by using a 3×3 Gaussian kernel or a 3×3 discrete Gaussian kernel. The standard deviation of the Gaussian blurring is calculated according to the size of the kernel. Blurring is used for noise reduction because those edges that disappear through blurring will only produce artifacts in edge detection, and thus eliminating these artifacts can improve the performance of the edge detection as well as the mask generation method. The edges sought in the method are the edges or boundaries of the static and dynamic regions of the image 10, so on the one hand blurring will not eliminate them, and on the other hand it will result in a more continuous line along the sought edges. The blurring step S200 produces a blurred image 20 having the same size and dimensions as the original image 10.

[0042] In the case where the camera records a color image 10, in a second conversion step S210, the color image is preferably converted to a gray-scale image. The second conversion step S210 can be performed before or after the blurring step S200. According to a preferred embodiment of the method, the second conversion step S210 produces a gray-scale blurred image 22. The second conversion step S210 can be implemented by any conversion method or algorithm known in the prior art.

[0043] In a preferred embodiment of the method, the second conversion step S210 can be omitted if it is directly implemented on the image 10 that generates a gray-scale image Figure 1For the first conversion step S130 described in [description], it can be implemented on the grayscale image Figure 2 The blurring step S200 of [description]. This embodiment also results in a grayscale blurred image 22.

[0044] In step S220, an edge image 24 is generated by detecting the edges on the image 10. If the image 10 is blurred in the blurring step S200, edge detection can be performed on the blurred image 20. If the image 10 or the blurred image 20 is converted to a grayscale image, edge detection can also be implemented on the grayscale blurred image 22. Edge detection can be achieved by any edge detection algorithm known in the prior art, such as first-order edge detectors, such as the Canny edge detector and its variants, the Canny-Deriche detector, edge detectors using the Sobel operator, the Prewitt operator, the Roberts cross operator, or the Frei-Chen operator; second-order edge detectors using the second derivative of the image intensity, such as the differential edge detector (detecting the zero-crossing of the second directional derivative in the gradient direction), algorithms, or edge detectors.

[0045] According to the present invention, edge detection is preferably implemented by the Laplacian algorithm with a kernel. The advantage of the Laplacian algorithm is that it finds edges faster than other more complex algorithms and requires less computing power. The Laplacian algorithm contains both negative and positive values around the detected edges, and these values cancel each other out for the edges of moving objects, so the static edges become significant faster. Laplacian kernels of different sizes can be used. Larger kernels result in slower edge detection, however, the detected edges do not change sharply, so smaller kernels provide similar good results as larger kernels. It is found that the edge detection algorithm using a 3×3 Laplacian kernel is optimal. The edge detection step S220 results in a continuous edge image 24 having the same size as the image 10.

[0046] In step S230, a cumulative edge image 26 is generated by summing the edge image 24 (preferably the continuous edge image 24). The cumulative edge image 26 has high values for static edges. Other edges and other parts of the image (e.g., edges of moving objects or edges of dynamically changing regions) have values lower than those of static edges. The size of the cumulative edge image 26 is the same as the size of the image 10. In a preferred embodiment of this method, for example, a previously cumulative edge image 26 is provided from a previous step of implementing this method, where the previously cumulative edge image 26 is preferably stored by a computer, and the cumulative edge image 26 is generated by adding the edge image 24 to the previously cumulative edge image 26.

[0047] To avoid an excessive increase in pixel values, the edge image 24 can be normalized in a second normalization step S240 by dividing each pixel value of the image by the number of accumulated images. The result of this second normalization step S240 is a normalized image, preferably a normalized accumulated edge image 28 having values in the same range as any image 10 from the camera stream.

[0048] The accumulated edge image 26 or the normalized accumulated edge image 28 will be further processed in the steps described in Figure 3 to generate a mask 40 for the non-static regions of the camera stream.

[0049] Figure 3 A preferred embodiment of the subsequent steps for generating the mask 40 for the non-static regions of the camera stream is described in Figure 1 The method uses the result of the calculation steps of the differential images described in Figure 2 such as the accumulated differential image 14, the normalized accumulated differential image 16 or the gray-scale accumulated differential image 18, and the edge detection results described in

[0050] The accumulated differential image 14 and the accumulated edge image 26 can be interpreted as histograms having the same size as the input images 10 of the camera stream. This interpretation also applies to the normalized accumulated differential image 16, the gray-scale accumulated differential image 18 and the normalized accumulated edge image 28. The histogram of the accumulated differences (accumulated differential image 14, normalized accumulated differential image 16 or gray-scale accumulated difference image 18) tends to show low values for static regions because these regions remain unchanged in the image and thus the difference will be close to zero. The histogram of the accumulated edges (accumulated edge image 26 or normalized accumulated edge image 28) has high values for static edges and low values for changing dynamic regions. For the mask to be generated, the changing regions are to be determined. Thus, in a combining step S300, the histogram of the accumulated edges is combined with the histogram of the accumulated differences to emphasize the contours of the static regions. The combining step S300 produces a combined image 30 generated by the combination of the histograms, where the combination of the histograms can be achieved by simple subtraction. In the case where the histograms are normalized, i.e., if the normalized accumulated edge image 28 is combined with the normalized accumulated differential image 16, or preferably if the normalized accumulated edge image 28 is subtracted from the normalized accumulated differential image 16, the emphasizing effect can be more significant. This latter preferred embodiment is described in Figure 3 where in the combining step S300, the normalized accumulated edge image 28 is subtracted from the normalized accumulated differential image 16, resulting in a combined image 30 with irrelevant low values for the contours of the static regions.

[0051] In an embodiment of the method of calculating the edge image 24 by the Laplace edge detection algorithm, the absolute value of the cumulative edge image 26 is to be calculated. Since the Laplace edge detection includes both negative and positive values, the negative values are eliminated by replacing them with their additive inverse values.

[0052] The result of the combining step S300 is a combined image 30 with irrelevant values. For example, in the case of subtracting the contour of a static region, there are irrelevant low values. A first threshold pixel value is defined for the combined image 30 to distinguish the irrelevant values of the combined image 30.

[0053] In the mask generation step S340, a mask 40 is generated by including the pixels of the combined image 30 whose relationship with the first threshold pixel value is the same as the relationship between the central pixel of the combined image 30 and the first threshold pixel value. Regarding the relationship between the pixel value of the central pixel and the first threshold pixel value, the pixel value of the central pixel can be greater than, greater than or equal to, less than, or less than or equal to the first threshold pixel value. Preferably, the mask 40 will only include those pixels of the combined image 30 whose relationship with the first threshold pixel value is the same as the relationship between the central pixel and the first threshold pixel value and which form a continuous region within the combined image 30. Therefore, other pixels (e.g., near the outer periphery of the combined image 30) whose relationship with the first threshold pixel value is the same as the relationship between the central pixel and the first threshold pixel value will not be included in the mask 40. The central pixel can be any pixel close to or located at the geometric center of the combined image 30.

[0054] In a preferred embodiment of the method according to the invention, the generation of the mask 40 starts from any central pixel of the combined image 30 and proceeds towards the outer periphery of the combined image 30. The inclusion of the pixels of the mask 40 stops at the pixels whose relationship with the first threshold pixel value is not the same as the relationship between the central pixel of the combined image 30 and the first threshold pixel value. This automatically excludes additional (random) pixels whose relationship with the first threshold pixel value is the same as the relationship between the central pixel and the first threshold pixel value from the mask 40.

[0055] The method according to the invention may include further steps as Figure 3 shown. Figure 3Preferred embodiments of the method shown include a first thresholding step S310, where if the combining step S300 is implemented by subtraction, the first threshold pixel value is preferably chosen to be close to zero because the features and regions to be marked have extremely low values compared to other parts of the image. If the combined image 30 is calculated from the normalized image, for example, the normalized cumulative edge image 28 is combined with (or subtracted from) the normalized cumulative difference image 16 or the normalized gray-scale cumulative difference image 18, then due to normalization, the absolute value of the combined image 30 is preferably in the range of [0, 1], while the values corresponding to static edges can be on the order of less than 1. As the number of cumulative images increases, the normalized pixel values of static edges decrease and get closer and closer to zero. In a preferred embodiment of the method according to the invention, the first threshold pixel value for marking the contours of non-static edges with extremely low values is chosen to be at least one order of magnitude less than 1. For example, the first threshold pixel value can be less than 0.1. For optimal results, the first threshold pixel value is set to be at least two orders of magnitude less than 1 or more preferably three orders of magnitude less than 1. For example, the first threshold pixel value is set to be less than 0.01 or more preferably less than 0.001 or even less than 0.0001. In the first thresholding step S310, a thresholded binary image 32 is generated by thresholding and binarizing the combined image 30. Pixels having a value lower than the first threshold pixel value are set to a first pixel value, while pixels having a value higher than the first threshold pixel value are set to a second pixel value, where the first pixel value and the second pixel value are different from each other. Preferably, the first pixel value is the maximum pixel value or 1, and the second pixel value is the minimum pixel value or 0. Thus, the first thresholding step S310 simultaneously accomplishes the thresholding and binarization of the combined image 30. The static regions of the original image 10 have the same pixel value (the second pixel value), while the non-static parts of the original image 10 have different pixel values (the first pixel value).

[0056] The thresholded binary image 32 is preferably blurred by calculating a histogram, preferably by calculating two or more histograms with different orientations of the thresholded binary image 32, and more preferably by calculating a horizontal histogram 34 and a vertical histogram 36 for horizontal and vertical blurring respectively in the histogram calculation step S320. The histograms (e.g., the horizontal histogram 34 and the vertical histogram 36) have a size smaller than that of the image 10 of the camera stream. The bins of the horizontal histogram 34 and the vertical histogram 36 are preferably set in proportion to the width and height of the image respectively, and the size of the bins determines the feature size of the image features to be blurred or filtered out. Smaller bins can only blur smaller features, while larger bins can filter out larger image structures. Preferably, the bins of the horizontal histogram 34 have a vertical size of 1 pixel in the vertical direction, and the horizontal size is a fraction of the horizontal size (width) of the thresholded binary image 32. More preferably, the horizontal size of the bins is in the range of 1 / 10 - 1 / 100 of the width of the thresholded binary image 32. The bins of the vertical histogram 36 preferably have a horizontal size of 1 pixel and a vertical size that is a fraction of the vertical size (height) of the thresholded binary image 32. More preferably, the vertical size of the bins is in the range of 1 / 10 - 1 / 100 of the height of the thresholded binary image 32. Preferably, the horizontal size of the bins of the horizontal histogram 34 and the vertical size of the bins of the vertical histogram 36 are determined by dividing the respective sizes of the thresholded binary image 32 by the same number. For example, if the horizontal size of the bins of the horizontal histogram 34 is 1 / 10 of the horizontal size of the thresholded binary image 32, then the vertical size of the bins of the vertical histogram 36 is also 1 / 10 of the vertical size of the thresholded binary image 32. For each bin of the horizontal histogram 34 and the vertical histogram 36, a Gaussian kernel having the same size as the bin is applied, and the standard deviation of the Gaussian kernel is determined by the size of the kernel.

[0057] The calculation of the horizontal histogram 34 and the vertical histogram 36 achieves blurring of the thresholded binary image 32 in order to remove weak features or noise of the thresholded binary image 32. The horizontal and vertical blurring is preferably achieved by calculating the horizontal histogram 34 and the vertical histogram 36 respectively.

[0058] In one embodiment of the method, the blurring (calculation of the histogram) in step S320 is performed in the rotational direction, preferably in directions having angles of 45° and -45° with respect to the sides of the thresholded binary image 32.

[0059] In another embodiment of the method, the blurring (calculation of the histogram) in step S320 is performed in the horizontal direction, the vertical direction, and the rotational direction, preferably in directions having angles of 45° and -45° with respect to the sides of the thresholded binary image 32.

[0060] The method preferably includes a supplementary step in which the histogram is stretched to the size of the original image 10. In step S330, the horizontal histogram 34 and the vertical histogram 36 are combined together, preferably by simple addition, to obtain a combined histogram image 38 having the same size as the original image 10.

[0061] The method according to the invention preferably includes a second thresholding step for further noise reduction ( Figure 3 not illustrated in). Depending on the selected second threshold pixel value, the thickness of the edges may be affected. Preferably, the second threshold pixel value is determined based on the maximum pixel value of the combined histogram image 38. For example, the second threshold pixel value is set to half or approximately half of the maximum pixel value of the combined histogram image 38. Selecting a higher second threshold pixel value results in the elimination of more features of the image and also eliminates part of the edges, usually resulting in thin edges. Selecting a lower second threshold pixel value results in the image having more remaining features and thicker edges. Preferably, a compromise is made between eliminating interfering features from the image and having reasonably thick edges. It has been found that a second threshold pixel value of half of the value of the maximum pixel value eliminates the interfering and noise-like features of the combined histogram image 38 and the edges also have an optimal thickness. In the second thresholding step, similar to the first thresholding step S310, the combined histogram image 38 is binarized in the following manner. Pixels having values higher than the second threshold pixel value are set to a specific value, preferably the maximum pixel value, more preferably the value 1. Pixels having values less than the second threshold pixel value are set to the minimum pixel value or zero. The resulting image has a maximum value for the static edges and a minimum value for the other parts of the image, and the size of this image is the same as the size of the original image 10.

[0062] According to Figure 3 an embodiment of, a mask 40 is generated in the following manner. Starting from the center of the image and proceeding towards the outer periphery of the image, the mask 40 includes all pixels having the same pixel value as the center. The inclusion of pixels stops at pixels having different pixel values (i.e., another pixel value of the image).

[0063] The starting point of the mask generation step S340 can be any other pixel around the center of the image because the area around the center of the image has the highest probability of not having static edge pixels.

[0064] In one embodiment of the method, the mask 40 can be generated starting from different starting points around the center of the image, so the method will be more robust and possible remaining artifacts or features of the image will not hinder the generation of the mask 40.

[0065] In another embodiment of the method, the mask generation step S340 is implemented by a ray marching algorithm that follows a straight line from the center of the image towards the outer periphery of the image. The ray marching algorithm takes longer and different steps along each direction line and thus it finds different pixel values faster than other algorithms. In this embodiment, all pixels in the mask 40 are included until those pixels whose relationship with the first pixel value is not the same as the relationship of the central pixel of the composite image 30 with the first pixel value.

[0066] For the further acquired images, preferably, for each newly acquired image 10 of the camera stream, the mask 40 can be regenerated by repeating one or more steps of the above method. Due to the uncertainties in the recorded images 10, camera jitter or other reasons, the consecutively generated masks 40 can be different from each other. However, as the number of the accumulated images 10 increases, the mask 40 will cover more and more of the real region of interest, i.e., the dynamically changing part of the image, and thus the generation of the mask 40 can be stopped after a certain period of time because the newly generated masks 40 will not be different from each other and are expected to cover substantially the same region and the same pixels.

[0067] In a preferred embodiment of the method, when a stable mask 40 is reached, the generation of the mask 40 can be stopped to save time and computing resources. By applying a stopping condition, the stopping of the mask generation can be implemented manually or automatically. In the case of automatic stopping, if the stopping condition is met, the generation of the mask 40 automatically stops. Once the generation of the mask 40 stops, the last generated mask 40 can be used for another image 10 of the camera stream. The applied stopping condition can ensure that the generated mask 40 is stable enough.

[0068] In an embodiment of the method, the generation of the mask 40 can be automatically stopped after a predetermined period of time or after a predetermined number of images 10 are obtained.

[0069] In another embodiment of the method, the stopping condition is implemented by evaluating a function with respect to a predetermined limit value. This embodiment of the method includes steps of calculating a metric value for each generated mask 40, a function for calculating the metric value, and stopping the generation of the mask 40 based on a comparison of the result of the function with the predetermined limit value. In a preferred embodiment of the method, the metric value of the mask 40 is the pixel count of the mask 40 (the total number of pixels included in the mask 40), and the function includes the average of the difference between the pixel counts of consecutive masks 40 and a predetermined number of subtracted pixel counts. If the result of the function is lower than the predetermined limit value, the stopping condition is met and the generation of the mask 40 stops. Thus, the stopping condition is reached if the mask 40 has become stable enough such that the number of pixels belonging to the mask 40 has not changed significantly.

[0070] In addition to averaging, other functions can also be used for the stopping condition. For example, these functions can include suitable filtering methods such as Kalman filters, GH filters (also known as alpha-beta filters), etc. Compared with more complex filtering methods, the method including averaging under the stopping condition has the advantage that it is reliable enough and requires less computing resources, so it is suitable for real-time mask generation.

[0071] In the case of camera misalignment, the method according to the present invention must be restarted to ensure that the mask 40 covers the true region of interest.

[0072] For safety reasons, each time the camera stream is stopped and then restarted, the mask generation method can be restarted. In this way, possible camera misalignment or other effects may be corrected by generating a new mask 40.

[0073] Some dashboard cameras are called fisheye cameras and have ultra-wide-angle lenses called fisheye lenses. Such fisheye lenses can obtain a very wide viewing angle, making them desirable for autonomous or self-driving cars or other vehicles. Due to the large viewing angle, the camera image is bound to record blank areas, such as static parts of a car or other vehicle, and these blank areas are located near the outer side of the image. The following compensation method can be applied to exclude the blank areas from the mask of the fisheye camera.

[0074] The fisheye camera can record color or grayscale fisheye images 10.

[0075] Successive fisheye images 10 are directly accumulated from the first recorded fisheye camera image. More preferably, the color values of the images 10 are accumulated, thereby generating a successive accumulated fisheye image. The accumulated fisheye image is normalized by the number of accumulated images.

[0076] If the blank areas of the fisheye image are not continuous on the outer periphery of the fisheye image, the third threshold pixel value can be calculated by calculating the average value of the pixel values on the outer periphery of the image. It has been found that the calculation of the third threshold pixel value is beneficial for thresholding the accumulated fisheye image because this third threshold pixel value maintains a balance between removing and retaining pixels. The threshold fisheye image includes the features of the image, while the blank areas of the fisheye image are excluded. The implementation of this third thresholding is similar to the first thresholding step S310 and also includes binarization of the image.

[0077] The thresholded fisheye image can undergo a fisheye mask generation step, but contrary to the mask generation step S340 described above, this mask is generated starting from the outer peripheral pixels of the image and traversing inward. In the fisheye mask generation step, pixels having the same value as the outer peripheral pixels are stored until the first pixel having a different value is reached. A convex hull is established around the stored pixels surrounding the blank areas of the fisheye image. The generated mask includes the pixels outside the convex hull shell.

[0078] If the generated mask becomes stable, for example by implementing any of the previously described stopping conditions, the generation of the continuous series of masks for the fish-eye image can be stopped.

[0079] Furthermore, the present invention relates to a computer program product comprising instructions which, when the program is executed by a computer, cause the computer to execute an embodiment of the method according to the present invention.

[0080] The computer program product can be executed by one or more computers.

[0081] The present invention also relates to a computer-readable medium comprising instructions which, when executed by a computer, cause the computer to execute an embodiment of the method according to the present invention.

[0082] The computer-readable medium can be a single medium or can comprise a plurality of individual pieces.

[0083] The present invention is of course not limited to the preferred embodiments described in detail above, but additional variations, modifications and developments are possible within the scope of protection determined by the claims. Furthermore, all embodiments defined by any arbitrary combination of dependent claims belong to the present invention.

[0084] List of reference numerals

[0085] 10 Image

[0086] 12 Difference image

[0087] 14 Cumulative difference image

[0088] 16 (Normalized) cumulative difference image

[0089] 18 (Gray-scale) cumulative difference image

[0090] 20 (Blurred) image

[0091] 22 (Gray-scale blurred) image

[0092] 24 Edge image

[0093] 26 Cumulative edge image

[0094] 28 (Normalized) cumulative edge image

[0095] 30 Composite image

[0096] 32 Thresholded binary image

[0097] 34 Horizontal histogram

[0098] 36 Vertical histogram

[0099] 38 Composite Histogram Image

[0100] 40 Mask

[0101] S100 (Subtraction) Step

[0102] S110 (Summation) Step

[0103] S120 (First Normalization) Step

[0104] S130 (First Transformation) Step

[0105] S200 (Blur) Step

[0106] S210 (Second Transformation) Step

[0107] S220 (Edge Detection) Step

[0108] S230 (Summation) Step

[0109] S240 (Second Normalization) Step

[0110] S300 (Combination) Step

[0111] S310 (First Thresholding) Step

[0112] S320 (Histogram Calculation) Step

[0113] S330 (Addition) Step

[0114] S340 (Mask Generation) Step

Claims

1. A method for generating a mask (40) for a non-static region based on a camera stream having consecutive images (10), comprising the steps of: - generating a cumulative difference image (14) by accumulating difference images (12), each difference image (12) being obtained by subtracting two images (10) of the camera stream from each other, - generating a cumulative edge image (26) by accumulating edge images, each edge image (24) being obtained by detecting edges in respective images (10) of the camera stream, - generating a combined image (30) by combining the cumulative edge image (26) and the cumulative difference image (14), - defining a first threshold pixel value for the combined image (30), and - generating the mask (40) by including in the mask (40) of the combined image (30) pixels having the same relationship to the first threshold pixel value.

2. The method according to claim 1, characterized in that, Comprising the steps of: - obtaining an image (10) of the camera stream, - generating a difference image (12) which is generated by subtracting two consecutive images (10) from each other, one image being the obtained image (10), - providing a previous cumulative difference image (14) constituted by the sum of previously generated consecutive difference images (12), and generating the cumulative difference image (14) by adding the difference image (12) to the previous cumulative difference image (14), - generating an edge image (24) by detecting edges in the obtained image (10), - providing a previous cumulative edge image (26) constituted by the sum of consecutively previously generated edge images (24), and generating the cumulative edge image (26) by adding the edge image (24) to the previous cumulative edge image (26), and - generating the mask (40) by including in the mask pixels of the combined image (30) having the same relationship to the first threshold pixel value as the central pixel of the combined image (30) has to the first threshold pixel value.

3. The method according to claim 2, wherein The generation of the mask (40) starts from the central pixel of the combined image (30) and proceeds towards the outer periphery of the combined image (30), and stops including pixels in the mask (40) at pixels whose relationship to the first threshold pixel value is not the same as the relationship of the central pixel of the combined image (30) to the first threshold pixel value.

4. The method according to claim 1, wherein Repeat the steps of claim 1 and stop the mask generation when a stop condition is met.

5. The method according to claim 4, characterized in that, Check the stop condition by means of an evaluation function relative to a predetermined limit value, comprising the steps of: - calculating a metric value for each generated mask (40), - calculating a function of the metric value, and - stopping the generation of the mask (40) based on a comparison of the result of the function with the predetermined limit value.

6. The method according to claim 5, characterized in that: - the metric value is the number of pixels counted for each generated mask (40), - the function comprises: - the difference in the pixel count of consecutive masks (40), and - The average of the subtracted pixel counts of a predetermined quantity, and - If the result of the function is lower than the predetermined limit value, stop the generation of the mask (40).

7. The method according to claim 1, characterized in that, Including another step after defining the first threshold pixel value, the step includes thresholding and binarizing the combined image (30) by setting pixels having values lower than the first threshold pixel value to a first pixel value and setting pixels having values higher than the first threshold pixel value to a second pixel value, to generate a thresholded binary image (32), wherein the first pixel value and the second pixel value are different from each other.

8. The method according to claim 7, wherein Including another step after generating the thresholded binary image (32), the other step includes: - Calculating two or more histograms of different orientations of the thresholded binary image (32), the histograms being calculated using a smoothing kernel, - Generating a combined histogram image (38) by combining the histograms, the size of the combined histogram image (38) being smaller than the original image (10), and - Stretching the combined histogram image (38) to the size of the image (10) of the camera stream.

9. The method according to claim 8, wherein Implementing a second thresholding for noise reduction on the stretched combined histogram image (38) by using a second threshold pixel value, and setting pixels having values higher than the second threshold pixel value to the first pixel value or the second pixel value, and setting pixels having values lower than the second threshold pixel value to another pixel value.

10. The method according to claim 9, wherein The second threshold pixel value is set to half of the maximum pixel value of the combined histogram image (38).

11. The method according to claim 1, characterized in that, Using the normalized cumulative difference image (16) as the cumulative difference image (14), and using the normalized cumulative edge image (28) as the cumulative edge image (26), wherein - The number of the cumulative difference images and the number of the cumulative edge images are respectively counted, - Generating the normalized cumulative difference image (16) by dividing the cumulative difference image (14) by the number of the cumulative difference images, and - Generating the normalized cumulative edge image (28) by dividing the cumulative edge image (26) by the number of the cumulative edge images.

12. The method according to claim 1, characterized in that, Generating the combined image (30) by subtracting the cumulative edge image (26) from the cumulative difference image (14).

13. The method according to claim 1, characterized in that, The step of detecting the edges is implemented by a Laplace edge detection algorithm.

14. The method according to claim 13, wherein Including an additional step of preferably blurring the image (10) by a blurred Gaussian kernel before the step of detecting the edges.

15. The method according to claim 1, characterized in that, The step of generating the mask is implemented by a ray tracing algorithm starting from the center of the image (10).

16. A non-transitory computer program product including instructions which, when the program is executed by a computer, cause the computer to perform the method according to any one of claims 1 to 15.

17. A non-transitory computer-readable medium including instructions which, when executed by a computer, cause the computer to perform the method according to any one of claims 1 to 15.

Citation Information

Patent Citations

  • Video object extraction apparatus and method

    US20110164823A1

  • Apparatus and method for image processing

    US20120014608A1

  • Method and apparatus using image subtraction and dynamic thresholding

    US6061476A

  • A metallographic image edge detection method based on mathematical morphology

    CN109544571A

  • Method and device for detecting edge

    JP1999102442A