Method and apparatus for stitching image data
By dividing the image data according to the level of interest and selecting an appropriate image band for each part for multi-band mixing, the problem of high cost of multi-band mixing in the prior art is solved, and a more efficient image quality and stitching appearance are achieved.
Patent Information
- Application Number
- CN202411591693.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-11-14
- Filing Date
- 2024-11-08
- Publication Date
- 2025-05-16
AI Technical Summary
The existing multi-band hybrid stitching method has high processing costs, making it difficult to achieve improved stitching appearance and improved image quality at low cost.
The overall processing cost is reduced by dividing the image data to be mixed into multiple parts according to its level of interest, and selecting one or more image bands for each part for multi-band mixing.
This improves image quality and stitching appearance, while reducing the use of processing resources, achieving a more efficient multi-band hybrid stitching method.
Smart Images

Figure CN120017980A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of image processing, and in particular to a method for processing video images using multi-band hybrid splicing. Background Art
[0002] So-called panoramic images are often generated by stitching two or more images with overlapping fields of view. For example, a multi-sensor camera can be used to capture several images simultaneously and stitch these images together to output a panoramic image. The panoramic images can form a video sequence.
[0003] Image stitching is generally the process of combining multiple images with overlapping fields of view. The stitching process can be divided into several process stages. First, the images are aligned so that they have matching viewpoints. This can be achieved using transformations. For example, if two images are to be stitched, one of the images can be transformed to match the viewpoint of the other image. The alignment stage can then be followed by a blending stage in which the image data of the multiple images are combined in the overlapping areas of the images, for example by forming a linear combination of the image data. The purpose of blending is to smooth the transition between the images so that a user viewing the stitched image experiences it as a single image.
[0004] There are different variants of blending techniques, of which multiband blending is one of them. In multiband blending, different versions of the image dataset to be blended are generated, representing different image bands. The generation can be achieved in different ways, for example by a technique known as Laplacian pyramid decomposition, in which the image dataset is decomposed into images of different resolutions corresponding to different image bands. The generated images are then blended separately, wherein known methods such as linear combinations or more complex methods of weighting the contributions of the image pixel values can be used. Thereafter, the blended images of the different bands are combined to generate a blended region of the final panoramic image.
[0005] Multi-band blending has the advantages of increased image quality and improved stitching appearance in the stitched areas of the final panoramic image, but has the disadvantage of higher processing costs.Therefore, there is a need to introduce a method in which the advantages of multi-band blending are achieved at low processing costs.
[0006] CA2671894 discloses a method for splicing input images. A low-frequency band sub-image and a high-frequency band sub-image are obtained from each input image. The low-frequency band sub-images are blended, and the high-frequency band sub-images are sliced or blended considering edge features to reduce artifacts.
[0007] XP031772756, "Building a Videoorama with Shallow Depth of Field" by Soonmin Bae et al., discloses a method for generating a video with a stitched background on which a dynamic foreground is rendered. The foreground and the stitched background are blended using a multi-band blending method.
[0008] CN115524343 discloses a method for obtaining a panoramic image of an ice slice for characterizing the physical structure of ice crystals. Summary of the invention
[0009] An object of the invention is to provide a method for stitching video images providing an improved stitching appearance and / or increased image quality at reduced processing costs compared to known solutions.Another object of the invention is to provide an improved multi-band hybrid stitching method with respect to processing efficiency.
[0010] The invention is defined by the appended claims.
[0011] According to a first aspect, the above and other objects are fully or at least partially achieved by claim 1 defining a method for stitching image data from one or more image sensors.
[0012] In order to exploit the advantages of multi-band mixing in a processing efficient manner, the method comprises performing multi-band mixing per portion of the image data and for one or more image bands selected individually for each portion. The invention is based on the recognition that image data having different levels of interest (e.g., different information values for a user viewing the image) can be mixed differently with respect to which image bands and how many image bands. Thus, the image data to be mixed is divided into a plurality of portions based on the level of interest, and a determination is made for each portion which bands to use based on its level of interest.
[0013] Thus, processing resources can be used more to provide good quality in image regions of interest, and less on regions of less interest. The overall processing level is also reduced compared to conventional multi-band mixing, which mixes all bands for all image data.
[0014] Here, the level of interest of the image data means the level of interest of the end application or end user of the generated stitched image in the depicted object. For example, for an object analysis application or a surveillance camera operator, image data depicting an object may have a high level of interest compared to image data depicting the sky or natural elements. Similarly, image data depicting moving objects may have a higher level of interest than image data depicting stationary objects. Therefore, the level of interest may depend on how the stitched image will be used, and in other words, is implementation-specific. The detailed description will provide examples of different implementations in which the level of interest has different bases.
[0015] The selection of the image bands to be mixed for the different parts can be free or restricted. In free selection, the image bands are freely determined with respect to the position and size of the bands within the image frequency interval of the image frequency values of the image data. In other words, the endpoints of the image bands are freely determined. In restricted selection, the image bands are selected among a predetermined set of image bands (i.e., image bands with fixed endpoints).
[0016] There are two types of dividing image data into multiple parts: spatial division and color component division. In other words, image data can be divided into parts of different spatial locations, or divided into parts of different components in a color model (e.g., RGB, CMYK, and YCbCr).
[0017] Understanding the details of the spatial partitioning, different partitioning bases can be implemented. Examples of these implementations include:
[0018] Image segmentation produces spatial regions with objects or things of varying levels of interest;
[0019] The motion analysis produces spatial regions with different motion levels corresponding to different levels of interest;
[0020] The predictive coding process produces spatial regions that are predicted to be encoded using different coding parameter values, such as different QP values, corresponding to different interest levels;
[0021] The image frequency analysis produces spatial regions of image data having different distributions of image frequency values corresponding to different levels of interest;
[0022] User input generates user-defined spatial regions having different levels of interest; and
[0023] The historical input produces spatial regions corresponding to the historical partitioning of the spatial regions into different levels of interest.
[0024] In one embodiment, the images are acquired simultaneously (i.e., at substantially the same point in time) by different image sensors. The image sensors may be arranged in a single multi-sensor camera. A benefit of acquiring the images simultaneously is that the alignment process is simplified in the presence of one or more moving objects in a mixed area. In addition, since computationally efficient stitching algorithms are particularly desirable in cameras implementing real-time stitching, the present invention is particularly advantageous in multi-sensor cameras in which the generation of the stitched image is performed in real-time, i.e., shortly after the image is captured and before it is sent from the camera.
[0025] The method according to the first aspect may be implemented as computer code instructions stored on a computer program product, wherein the computer code instructions are adapted to carry out the method when executed by a device having processing capabilities.
[0026] According to a second aspect, these and other objects are achieved in whole or in part by claim 14 defining an apparatus for stitching image data from one or more image sensors. The apparatus may be a surveillance camera. The apparatus of the second aspect may generally be implemented in the same manner as the method of the first aspect, with attendant advantages.
[0027] Further scope of applicability of the present invention will become apparent from the detailed description given hereinafter.
[0028] It must be noted that, as used in the specification and the appended claims, the articles "a", "an", "the" and "said" are intended to indicate the presence of one or more elements, unless the context clearly dictates otherwise. Thus, for example, references to "an object" or "the object" may include several objects, etc. Furthermore, the word "comprising" does not exclude other elements or steps. BRIEF DESCRIPTION OF THE DRAWINGS
[0029] The invention will now be described in more detail by way of example and with reference to the accompanying schematic drawings, in which:
[0030] Figure 1 illustrates a multi-sensor camera whose image sensors have overlapping fields of view of a scene;
[0031] Figure 2 illustrating an image having image data depicting overlapping portions of a scene;
[0032] Figure 3 Picture only Figure 2 The overlapping part of
[0033] Figure 4 is a flowchart of a method for stitching image data according to an embodiment; and
[0034] Figure 5is an overview of components in a camera apparatus suitable for performing stitching of image data according to an embodiment. DETAILED DESCRIPTION
[0035] Figure 1 The configuration of a multi-sensor camera 12 is illustrated with a plurality of image sensors 14, 16 that acquire a video depicting a scene 10. In this example, the image sensors 14, 16 are located in a single camera device, however it is equally possible to have the image sensors located in separate camera devices. The image sensors 14, 16 are configured with overlapping fields of view 15, 17 for generating a panoramic video (i.e., a video composed of panoramic images generated by stitching images from the image sensors 14, 16). The scene 10 in this example includes a person 18 walking in an outdoor environment. In Figure 1 In FIG. 1 , the person 18 is located in the overlapping region 19 of the fields of view 15 , 17 , which means that the person 18 will be depicted by both image sensors 14 , 16 .
[0036] Figure 2 The images 24, 26 acquired by the image sensors 14, 16 are shown. In order to generate a panoramic image from the images 24, 26, a stitching process is applied. First, Figure 2 As shown in the lower part of , the images 24, 26 are aligned. The purpose of the alignment is to determine image data of the images 24, 26 that correspond to the same depicted content. In other words, the image data of each image 24, 26 describing the scene content located in the overlap region 19 is determined and matched. The alignment is performed using conventional methods (e.g., by determining matching viewpoints or feature points in each of the images 24, 26 and aligning these points). The images 24, 26 may also be subjected to transformations during the alignment.
[0037] Thereafter, a blended area 28 represented by the image data sets in each of the images 24, 26 is determined. The image data sets will be blended to form a stitched area in the panoramic image. In the present embodiment, the blended area 28 is equal to the overlap area, however, it should be understood that the blended area 28 can be constructed differently, such as being larger or smaller than the overlap area, and thus including more or less image data from the images 24, 26 in the image data set. In addition, in the present embodiment, the blended area 28 forms a vertical extension, however, it should be understood that the extension of the blended area 28 can have different angles, such as being slightly tilted to the left or right. Such an embodiment can improve the stitching appearance when depicting tilted objects in the overlap between the images 24, 26, wherein the blended area 28 is tilted in the same manner as the object.
[0038] A first image data set is determined in the first image 24, and a second image data set is determined in the second image 26, both corresponding to a blending region 28. After the first image data set and the second image data set have been determined, these data sets are blended using a variant of multi-band blending in order to generate a smooth transition between the first image 24 and the second image 26 in the panoramic image to be formed later. In general, multi-band blending is a known technique in which a first image data set and a second image data set are blended in a plurality of different image bands. For example, each image data set is decomposed into a plurality of decomposed image data sets of different resolutions, thereby representing different image bands. The decomposition may be performed according to a Laplacian pyramid decomposition. Then, for each band, the decomposed image data sets from the first image data set and the second image data set are blended, i.e., the decomposed set of a particular band of the first image data set is blended with the decomposed set of the same particular band of the second image data set. The result is a plurality of blended image data sets that are combined to form a blending region of the stitched image data. It should be understood that for the purpose of multi-band blending, other techniques besides Laplacian pyramid synthesis may be applied.
[0039] Now refer to Figure 4 A general variant of multi-band mixing using the inventive concept is described. The object of the invention is to exploit the advantages of general multi-band mixing techniques, but with lower processing costs than known methods. To this end, a conscious selection of which image bands are included in the multi-band mixing based on the level of interest of the image data to be mixed is introduced.
[0040] In a first step, image data of two images to be stitched are obtained 402. As described above, the image data obtained represent a determined mixed region of the image. Next, the obtained image data is divided 404 into parts with different levels of interest. These parts can be divided spatially or represent different color components of a color space. Different examples of how to achieve the division will be provided later. Next, each part is subjected to a premixing process 405. For each part, one or more image bands are determined 406, and image data of the determined image bands are obtained S408. The image data can be represented in the spatial domain or the frequency domain. Next, according to conventional multi-band mixing, for each part within its image band, the image data of the determined image bands are mixed 410. Therefore, for each part, only the image data of the determined image bands are mixed. This can, for example, mean that one part is mixed by mixing the image data of a single image band, while another part is mixed by mixing the image data values of multiple image bands. Then, the mixed parts are combined to form a mixed region of the stitched image data.
[0041] Preferably, the blending 410 is performed in the spatial domain, for example by combining pixel values of the different image data sets using a weighting function that determines the contribution of each of the image data sets to the blended image data.
[0042] For a more in-depth understanding of the details of the method, reference will be made to the diagram of the mixing region 28. Figure 2 An embodiment is disclosed. This embodiment implements the spatial division of a mixed region 28. The image data of the mixed region 28 is divided into spatial portions 32, 34, 36 of different interest levels, and a frequency band is determined for each spatial portion 32, 34, 36. The interest level of the image data depends on the application. For example, in a surveillance application, a portion depicting a person has a higher interest level than a portion depicting the sky. In the illustrated example, a first portion 32 corresponds to a sky region, a second portion 34 corresponds to a person, and a third portion 36 corresponds to a ground region. For each of these portions 32, 34, 36, one or more image frequency bands within an image frequency interval of the image data set are determined based on the interest level of the portion. As described above, these portions are mixed separately by applying multi-band mixing to each portion. Then, the mixed image data sets of the portions 32, 34, 36 are spatially combined to form mixed image data of the mixed region 28. The mixed image data of the mixed region 28 is spatially combined with the remaining image data of the images 24, 26 to form a panoramic image. The formation of the panoramic image may include, for example, cropping or otherwise adjusting the images 24 , 26 for a better appearance.
[0043] Going deeper into the details of dividing the mixed region 28 into multiple parts, different embodiments of how the division can be implemented will be described. The purpose of the division is to divide the image data into multiple parts based on its interest level. Depending on the application, the interest level can be defined differently, which will be illustrated, and therefore, there are multiple variants of how to divide the image data into multiple parts. In general, there are two possible ways. In a first variant, the mixed region 28 is divided into spatial parts of different interest levels. In a second variant, the mixed region 28 is divided into color components of different interest levels, which means that the part forms one or more different color components of the same color model. These variants can be combined so that the mixed region 28 is divided into multiple spatial parts, which are in turn divided into multiple color components, and vice versa. In addition, it should be noted that the same division is applied to both image data sets. This can be achieved by determining the division based on the first image data set and applying the same division to the second image data set.
[0044] A variant of the division into spatial portions will now be described in more detail.
[0045] In a first embodiment, image data depicting different object types have different levels of interest. To this end, the division is based on image segmentation. A conventionally constructed image segmentation algorithm is applied to the image data of the mixed region 28 to output a segmentation result. Either the first image data set or the second image data set or both can be analyzed. These possible selections of image data to be analyzed are also valid for other embodiments disclosed herein including image analysis.
[0046] The segmentation algorithm can be implemented using, for example, a neural network. The segmentation result can, for example, include a pixel-level segmentation result indicating the most likely object class per pixel. The image data of the mixed area 28 is divided into pixel portions having the most likely object class, and the most likely object classes have the same level of interest. In surveillance applications, the object class of a person typically has a high level of interest, and the object class of the sky typically has a low level of interest. Therefore, the pixel that is most likely to be the sky will be part of another portion in addition to the pixel that is most likely to be a person.
[0047] Alternatively, a segmentation algorithm can be implemented using a background model that segments foreground pixels from background pixels. In this case, the portions will become one or more portions with background pixel image data and one or more portions with foreground pixel image data, wherein, for many applications, the foreground portion will have a higher level of interest than the background portion.
[0048] It will be appreciated that other variations of segmentation algorithms exist, including variations utilizing artificial intelligence that can be used to divide image data into spatial portions.
[0049] In a second embodiment, the image data of things with different motion levels are depicted with different levels of interest. For this reason, the division is based on motion analysis. The image data of mixed area 28 is analyzed to separate the image data containing different motion levels. This can be achieved by comparing the image data with the image data previously acquired (preferably, the image data in the previous image of the same scene area). By analyzing the change in pixel value, the motion level can be determined at the pixel level or pixel block level. The pixels or pixel blocks with the motion level in the same motion level interval can be associated, and together form a spatial portion, thereby realizing the division. In some embodiments, the part with a higher motion level can be considered to have a higher interest than the part with a lower motion level.
[0050] In a third embodiment, image data to be encoded using different encoding parameter values has different levels of interest. To this end, the image data is divided based on how it may be encoded in a subsequent encoding process that implements predictive video coding. The parts into which the image data is divided are spatial parts corresponding to different encoding parameter values (e.g., QP values). QP is a coding parameter used in conventional codecs (e.g., H.264 and H.265) and indicates the level of compression. A high QP value corresponding to high compression results in low-quality image data, and vice versa.
[0051] Predicting which encoding parameter values the image data will use to encode can be achieved by analyzing the image data or receiving information from the encoding process. The encoding process can, for example, provide information representing the most recently applied QP value or the most recent motion level of the encoded image data. In a standard environment, the encoded image data will be image data of one or several image frames acquired before the image data to be divided. This information can be used to determine spatial regions with different QP values. These parts can be divided to correspond to different intervals of QP values, such as a part with QP values 0-29 and a part with QP values 30 and higher. In this embodiment, the level of interest of the spatial part corresponding to the lower QP value (or lower QP value interval) is higher than that of the spatial part corresponding to the higher QP value (or higher QP value interval). This is because high-quality image data may have greater value than low-quality image data, and therefore, it may be more useful to have a better splicing appearance in the high-quality image data part.
[0052] In a fourth embodiment, image data of different image frequencies have different levels of interest. To this end, the image data is divided based on its frequency content. The image data can be analyzed by a conventional algorithm based on, for example, a discrete Fourier transform (DFT) to determine its image frequency content. The analysis can produce the number of different image bands per image area. The image area can include one or more pixels. According to the number of different image bands, the division into spatial parts can be completed by collecting image areas with the same main image band. In other words, the spatial part will include image data with the same main image band. In this embodiment, the image bands used in multi-band mixing can be determined before or after the analysis. In other words, either the image data is analyzed to determine the amount of image frequency content of a predetermined image band, or the image band is determined based on the result of the image frequency analysis. The latter option provides a more flexible solution because the image band can adapt to the frequency content of the image data. In addition, the number of image bands can adapt to the image frequency content of the image data and can vary between image data sets. On the other hand, the first option of the predetermined image band may require less processing resources. For image data with few image frequencies, it may be sufficient to determine only two image bands. For image data comprising more varying image frequency content, the image frequency interval of the frequency content may be divided into more than two image frequency bands.
[0053] In the present fourth embodiment, image data having a higher main image frequency can have a higher level of interest than image data having a lower main image frequency.
[0054] A detailed variant of the partitioning into portions corresponding to different color components of a color space will now be disclosed. In one embodiment, a YCbCr color space is used as the basis for the partitioning, wherein the Y (luminance) component and the CbCr (chrominance) component have different interest levels. A representation of the image data in the YCbCr color space is generated in a known manner, and portions of the image data representing different color components are formed. A portion may include image data of the Y color component, and a portion may include image data of both the Cb component and the Cr component (i.e., the chrominance component). In this embodiment, the portion of the Y color component may have a high interest level, and the portion of the chrominance color component may have a low interest level. In surveillance applications, since the Y color component represents interesting structural information from a forensic perspective, the portion of the Y color component is generally more interesting than the portion of the chrominance color component. In other embodiments, a partition based on the RGB color space can be applied in a similar manner. For example, in the case where the image sensor has a better resolution in one of the color components (e.g., the G color component), the portion of the color component is more interesting than the portion of the other color components (e.g., the R and B color components).
[0055] The interest level of a portion of a color component may depend on which object the depicted thing is. For example, the image data may first be divided into spatial portions using segmentation. Thereafter, each spatial portion is divided into sub-portions of the image data of the spatial portion including different color components (e.g., Y and CbCr of the YCbCr color space). These sub-portions have different interest levels depending on the segmentation label of the spatial portion (i.e., depending on what type of thing the portion describes). In this example, the segmentation label of the portion may be "sky", for which the sub-portions of the Y color component and the CbCr color component have different importance in the representation. The Y sub-portion has a low interest level for the sky because it has a low impact on the representation of the sky, and the CbCr sub-portion has a high interest because it has a high impact on the representation of the sky. Therefore, a single image band for mixing the Y sub-portion can be determined. A greater number of image bands for mixing the CbCr sub-portions may be determined. For the segmentation labels of "person" or "license plate", the sub-portions of the Y color component and the CbCr color component may have other interest levels. For the spatial portions of any of these segmentation labels, the sub-portion of the Y color component may have a high level of interest and the sub-portion of the CbCr color component may have a low level of interest because the Y color component provides more valuable information for depicting human objects or license plate characters.
[0056] In one embodiment, the division into portions of different color components may be performed, followed by the division of some or each portion into sub-portions of different spatial components. For example, the division into color components Y and CbCr may be performed first, followed by the division of the portions of the Y color component into sub-portions of different segmentation labels mixed using separately determined frequency bands. Sub-portions determined to be "sky" may be considered of low interest, and sub-portions determined to be "ground" may be considered of high interest, and the frequency bands determined accordingly. In this example, portions of the CbCr color component may not be divided into sub-portions at all, while portions of the Y component are divided into spatial portions. Again, this is an example of how the method can be customized to the needs of a particular implementation without introducing unnecessary processing.
[0057] The mixed sub-portions are then combined to form stitched image data of the portion, which in turn may be blended with other portions according to any of the embodiments described herein.
[0058] For two common variants of partitioning, i.e., partitioning into spatial parts or parts of different color components, a user-defined partitioning can be applied. In the present embodiment, an input of a user via, for example, a user interface can be received. The input can indicate that the user is interested in the spatial parts and / or color components, and therefore, should be assigned a high level of interest. The input can, for example, indicate an image area of a doorway, road, or parking lot of interest to be monitored.
[0059] Alternatively, and also for any of the general variants of partitioning, a historical partitioning of the image data set in previous image data sets representing the same image area can be used to determine portions of different interest levels. This embodiment may include determining these portions based on historical partitioning in cases where the depicted scene has not changed sufficiently. When sufficient changes are detected, another embodiment (e.g., a partitioning based on segmentation) may be temporarily applied until the scene changes become low again and therefore there are not sufficient changes. For example, the amount of scene change may be determined based on a motion level determined for the image or based on data from an external sensor such as a motion sensor, a lidar sensor, or a radar sensor.
[0060] Going into more detail regarding the determination of one or more frequency bands for a portion based on the level of interest in the portion, there are different embodiments of the determination. First, it can be noted that the determination of one or more image frequency bands can include selecting from predetermined image frequency bands, or can include the determination of image frequency bands. The predetermined image frequency bands can be fixed, meaning that they are the same for all processed image data, which can reduce overall processing requirements because the decomposition can be performed by a hardware module.
[0061] In one embodiment, for portions having a higher level of interest, a larger number of frequency bands are determined.
[0062] In another embodiment, one or more high frequency bands are determined for the portion with a higher interest level compared to the portion with a lower interest level. For example, the portion with a high interest level may be mixed in the high frequency, medium frequency and low frequency bands, while the portion with a low interest level may be mixed only in the low frequency band.
[0063] In one embodiment, due to limited processing power, the number of bands that can be mixed is limited. For example, in a camera device that streams live panoramic video, the time resources and processing resources for splicing more than a certain number of image bands are limited. In one example, each part can mix up to two bands. According to the inventive concept, it is determined which two image bands are based on the interest level of the part. In an embodiment with a predetermined image frequency level of a high image band, a medium image band, and a low image band, for the purpose of multi-band mixing, a high-interest part can be assigned a high image band and a medium image band, and a low-interest part can be assigned a medium image band and a low image band. For the low-interest level part, it is preferred to select the image band with the lowest frequency because this generally uses the least memory bandwidth. In some embodiments, memory bandwidth may be a bottleneck, and it is therefore advantageous to reduce the amount of memory bandwidth required.
[0064] In the case where the frequency bands are not fixed, the frequency bands may be determined for each portion of each image data set, or may be determined from time to time and thus applicable to several image data sets. The determination of the frequency bands may be based on the content of the spatial portion. For example, the spatial portion corresponding to the segmentation label "person" may be mixed with one or more frequency bands determined to be suitable for depicting a person in the image. The determination may be performed by analyzing the spatial portion using a neural network trained to determine suitable image frequency bands for the image data.
[0065] Figure 5Components of a camera device 5 implementing an embodiment of the proposed hybrid method are illustrated. The camera device 5 includes a pair of image sensors 52, 54 arranged to acquire image data depicting a scene. The image sensors 52, 54 are arranged with overlapping fields of view of the scene. The acquired image data is fed to an image processing unit 56, which is sometimes referred to as an image processing pipeline, which is arranged to process the image data in terms of, for example, noise reduction, white balance or tone mapping. The image processing unit 55 may also include image analysis such as object detection from which data representing rough segmentation, motion analysis or object tracking can be obtained. The processed image data is transmitted to an image stitcher 56 suitable for performing image data stitching. The image stitcher 56 may include an image data obtainer (ID obtainer) suitable for determining and obtaining an image data set to be blended from the received image data. The image stitcher 56 further includes a divider suitable for dividing the acquired image data set into a plurality of parts according to any one of the embodiments described herein. The image stitcher 56 further includes an image frequency obtainer (IF obtainer) adapted to perform processing to determine one or more image bands of each portion based on the interest level of the portion, and to obtain image data of the determined one or more image bands from the image data of the portion. The divider and the image frequency obtainer may retrieve information (e.g., information from the image processing unit 55, a memory unit (not shown), or other processing units (not shown)) from outside the image stitcher 56 to perform their tasks. The image stitcher 56 further includes a mixer adapted to mix the image data of the mixed region using multi-band mixing. For each portion, as described in the embodiments herein, the mixer only mixes the obtained image data of the determined one or more image bands. The mixer also combines the mixed portions of the image data into stitched image data of the mixed region, and the stitched image data is combined with the rest of the received image data to form a panoramic image.
[0066] The image data corresponding to the panoramic image is sent to the encoder 58 to be video-encoded according to a conventional video compression standard such as H.264 or H.265. Thereafter, the encoded panoramic image can be transmitted from the camera device 5 to a receiver such as a streaming server. A panoramic video can be generated by continuously acquiring and processing the image data encoded as a video stream and transmitted from the camera device 5.
[0067] The image stitcher 56 of the camera device 5 may be implemented as software, hardware or a combination of these.
[0068] In a hardware implementation, the image stitcher 56 may correspond to a circuit that is dedicated and specifically designed to provide the functions of the components in the image stitcher (i.e., the image data obtainer, the divider, the image frequency obtainer, and the mixer). The circuit may be in the form of one or more integrated circuits such as one or more application specific integrated circuits or one or more field programmable gate arrays.
[0069] In a software implementation, the circuitry may alternatively be in the form of a processor such as a microprocessor, which in association with computer code instructions stored on a (non-transitory) computer-readable medium (e.g., non-volatile memory) causes the processor to perform (a portion of) any of the methods disclosed herein. Examples of non-volatile memory include read-only memory, flash memory, ferroelectric RAM, magnetic computer storage devices, and optical disks, among others. In the case of software, therefore, the components of the image stitcher 56 may each correspond to a portion of computer code instructions stored on a computer-readable medium, which, when executed by the processor, causes the image stitcher 56 to implement the functionality of that component.
[0070] It should be understood that it is also possible to have a combined implementation of hardware and software, ie the functionality of some of the components in the image stitcher 56 is implemented in hardware and others in software.
[0071] Those skilled in the art will appreciate that the present invention is not limited to the above-described embodiments. On the contrary, many modifications and variations are possible within the scope of the appended claims. By studying the drawings, the disclosure, and the appended claims, those skilled in the art can understand and implement these modifications and variations when practicing the claimed invention. For example, it is possible to partition into spatial portions based on other information in addition to the information provided by the embodiments herein, as long as the information can be obtained by the image stitcher used to perform the partitioning.
Claims
1. A method for generating stitched image data for an application, wherein: The image data is acquired by one or more image sensors arranged to acquire image data depicting at least partially overlapping perspectives of a scene, the method comprising: obtaining a first image data set and a second image data set, wherein each image data set represents a mixed region; defining different levels of interest in the image data for different spatial regions or different color components of a color space, depending on the application for which the stitched image data is generated; Each image data set is divided into portions of defined different levels of interest corresponding to the same plurality of spatial regions of the mixed region or corresponding to the same plurality of color components, for each portion of the image data: determining one or more image bands of the portion taking into account the level of interest of the portion; and obtaining image data for the determined one or more image bands from the portion of the image data; The first image data set and the second image data set are mixed by multi-band blending, wherein, for each portion of the image data, only the obtained image data of the determined one or more image bands are blended to generate stitched image data of the blended area for the application.
2. The method according to claim 1, wherein: The step of determining one or more image frequency bands includes selecting one or more image frequency bands from a fixed set of image frequency bands.
3. The method according to claim 1, wherein: The image data set is divided into portions corresponding to the same plurality of spatial regions of the mixed region and comprises: performing image segmentation on the image dataset; and The image data set is divided into spatial portions corresponding to segmented regions defined at different levels of interest.
4. The method according to claim 3, wherein: The image segmentation is performed using a background model.
5. The method according to claim 3, wherein: The image segmentation is performed using an image classification algorithm or an image recognition algorithm.
6. The method according to claim 1, wherein: The image data set is divided into portions corresponding to the same plurality of spatial regions of the mixed region, the method further comprising: A motion analysis is performed on the image data set and the image data set is divided into spatial regions corresponding to different motion levels having defined different levels of interest.
7. The method according to claim 1, wherein: The image data set is divided into a plurality of spatial regions, and the method further comprises: receiving information indicative of encoding parameter values for said image data set from a predictive encoding process; and The image data set is divided into spatial regions corresponding to different encoding parameters having different defined levels of interest.
8. The method according to claim 1, wherein: The image data set is divided into a plurality of spatial regions, and the method further comprises: analyzing the image data set to determine the image frequency content; and The image data set is divided into spatial regions having different dominant frequencies with defined different levels of interest.
9. The method according to claim 1, wherein: The image data set is divided into the same plurality of color components, wherein the color components are Y, Cb, and Cr of a YCbCr color space.
10. The method according to claim 1, wherein: The image data set is divided into a plurality of portions based on a user-defined partition.
11. The method according to claim 1, wherein: The image data set is divided into a plurality of regions, and the method further comprises: determining a historical partitioning of the image data set in previous image data sets representing the same image region; and The image data set is divided into a plurality of parts according to the historical division.
12. A method for generating a sequence of stitched images, comprising: obtaining a first sequence of images and a second sequence of images from a plurality of image sensors arranged to acquire image data depicting at least partially overlapping perspectives of a scene; For a temporally corresponding image data set, the method according to claim 1 is performed, and A sequence of stitched images is generated based on the first sequence of images, the second sequence of images and the generated stitched image data of the mixed area.
13. A computer program product comprising a computer readable medium having computer code instructions stored thereon, the computer code instructions being adapted to implement the method of claim 1 when executed by a device having processing capabilities.
14. A device for generating stitched image data for an application, wherein: The image data is acquired by one or more image sensors arranged to acquire image data depicting at least partially overlapping perspectives of a scene, the apparatus comprising: an image data obtainer arranged to obtain a first image data set and a second image data set, wherein each image data set represents a mixed region; a divider arranged to divide each image data set into portions of defined different levels of interest corresponding to the same plurality of spatial regions of the mixed region or to the same plurality of color components, wherein the different levels of interest of the image data of different spatial regions or different color components of the color space are defined depending on the application for which the stitched image data is generated; Image frequency obtainer, adapted for each portion of image data: determining one or more image bands of the portion taking into account the level of interest of the portion; and obtaining image data of said determined one or more image bands from said portion of image data, A mixer adapted to mix the first image data set and the second image data set by multi-band mixing, wherein, for each portion of the image data, only the obtained image data of the determined one or more image bands are mixed to generate stitched image data of the mixed area for the application.
15. A surveillance camera adapted to provide a live stream of panoramic video, wherein: The surveillance camera device comprises a device for stitching image data according to claim 14.