Camera system and method for encoding video image frames
By identifying and adjusting the compression level of overlapping areas in video image frames, the problem of excessive bit count in existing technologies is solved, achieving the effect of reducing bit count while maintaining or improving quality.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- AXIS
- Filing Date
- 2023-06-12
- Publication Date
- 2026-04-10
AI Technical Summary
In the overlapping region between video image frames captured by two image sensors, existing technologies struggle to effectively reduce the number of bits while maintaining the desired quality for each segment.
By identifying overlapping regions in two video image frames, a higher compression level is selectively applied to the overlapping regions, while a lower compression level is applied to the unselected regions, in order to limit the number of bits after encoding, while maintaining or improving the quality of the overlapping regions.
It effectively reduces the number of bits after encoding, while maintaining or improving quality in overlapping areas, adapting to the importance and application requirements of different areas.
Smart Images

Figure CN117255202B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present invention relates to encoding video image frames captured by respective ones of two image sensors, and in particular to adjusting a compression level in an overlapping area between video image frames captured by the two image sensors. BACKGROUND
[0002] When transmitting and / or storing video image frames, it is of interest to limit the number of bits to be stored and / or transmitted, and to provide a desired quality for each part of the video image frames. This is achieved by means of compression, for example, when encoding video image frames in an image processing apparatus comprising an encoder. If the video image frames are captured by respective ones of two or more image sensors, the amount of data to be transmitted and / or stored becomes larger, and the desire to limit the amount of data becomes even more desirable. SUMMARY
[0003] It is an object of the present invention to facilitate a reduction of the number of bits of encoded video image frames captured by respective ones of two or more image sensors. It is a further object of the present invention to facilitate providing a desired quality for each part of video image frames of encoded video image frames captured by respective ones of two or more image sensors.
[0004] According to a first aspect, there is provided a method for encoding two video image frames captured by respective ones of two image sensors, wherein each of the two video image frames depicts a respective part of a scene. The method of the first aspect comprises identifying a respective overlapping area in each of the two video image frames, the two overlapping areas depicting a same subpart of the scene, and selecting a video image frame of the two video image frames. The method of the first aspect further comprises setting a compression level for the two video image frames, wherein a respective compression level is set for a block of pixels in the selected video image frame based on a given principle, and wherein the respective compression level of a block of pixels in the overlapping area in the selected video image frame is selectively set higher than the respective compression level that would have been set based on the given principle. The method of the first aspect further comprises encoding the two video image frames.
[0005] Describing a same subpart of a scene means that each overlapping area is a two-dimensional representation of a same three-dimensional subpart of the scene. This does not necessarily mean that the overlapping areas of the two video image frames are identical, as the two sensors can be positioned at an angle to each other and / or the two video image frames can have undergone different projection transformations.
[0006] The compression level refers to a value indicative of a level of compression, such that the higher the compression level, the higher the level of compression.
[0007] The inventors have realized that if each of the two video image frames comprises a respective overlap region, wherein the two overlap regions depict the same sub-portion of the scene, the bit number of the encoded versions of the two video image frames can be limited by increasing the compression level set for the pixel blocks in the overlap region of a selected one of the two video image frames. At the same time, since the compression level is selectively increased for the pixel blocks in the overlap region in the selected video image frame, this does not affect the compression level for the pixel blocks in the overlap region of the non-selected video image frame. This can even accommodate a decrease in the compression level in the overlap region of the non-selected video image frame. Thus, the compression level set for the pixel blocks of the overlap region of the non-selected video image frame can remain determined based on a given principle. Thus, since the overlap region of the non-selected video image frame depicts the same sub-portion of the scene as the overlap region of the selected video image frame, the overlap region of the non-selected video image frame can be used to contribute to providing the desired quality for the overlap region.
[0008] The given principle can be any principle for setting the compression value for a pixel block. "Given" means that the principle is predetermined, e.g. in the sense that it is the principle implemented in the apparatus with which the method is related, or in the sense that the user has selected the principle from alternative principles implemented in the apparatus with which the method is related. In other words, the given principle is the principle according to which the compression value of a pixel block will be set to a default value, i.e. unless another instruction related to the compression value of a pixel block is given.
[0009] The given principle may, for example, be that a respective compression level is set for a pixel block in the selected video image frame based on a respective property or property value associated with the pixel block in the selected video image frame. Such a property may, for example, be a respective interest level associated with the pixel block in the selected video image frame.
[0010] The interest level represents the relative interest or importance of different regions of an image frame. What is considered to have a relatively high interest or importance, and what is considered to have a relatively low interest or importance, will depend on the application.
[0011] In the method of the first aspect, the act of selecting the video image frame of the two video image frames can include selecting the video image frame of the two video image frames that has one or more of the following image properties in the respective overlapping region: lower quality focus, lowest resolution, lowest angular resolution, lowest dynamic range, lowest light sensitivity, most motion blur, and lower quality color representation. Thereby, the compression level of the pixel blocks is selectively set higher in the overlapping region of the video image frame that has one or more of the lower quality focus, lowest resolution, lowest angular resolution, lowest dynamic range, lowest light sensitivity, most motion blur, and lower quality color representation in its overlapping region. This is beneficial because the overlapping region of the selected video image frame can be the least usable region, e.g., for image analysis.
[0012] The method of the first aspect can further include identifying one or more objects in the respective identified overlapping region of each of the two video image frames. The act of selecting the video image frame of the two video image frames can include selecting the video image frame of the two video image frames in which the one or more objects are most occluded or the objects are least identifiable. Thereby, the compression level of the pixel blocks is selectively set higher in the overlapping region of the video image frame in which the one or more objects are most occluded or the objects are least identifiable in its overlapping region. This is beneficial because the overlapping region of the selected video image frame can be the least usable region for at least one of the identification and analysis of the objects.
[0013] In the method of the first aspect, the act of selecting the video image frame of the two video image frames can include selecting the video image frame of the two video image frames that has one or more of the following image properties in the respective overlapping region: lower quality focus, lowest resolution, lowest angular resolution, lowest dynamic range, lowest light sensitivity, most motion blur, and lower quality color representation. Thereby, the compression level of the pixel blocks is selectively set higher in the overlapping region of the video image frame that has one or more of the lower quality focus, lowest resolution, lowest angular resolution, lowest dynamic range, lowest light sensitivity, most motion blur, and lower quality color representation in its overlapping region. This is beneficial because the overlapping region of the selected video image frame can be the least usable region, e.g., for image analysis.
[0014] In the method of the first aspect, the act of selecting the video image frame of the two video image frames can include selecting the video image frame of the two video image frames that is captured by the image sensor that has the farthest distance to the sub-portion of the scene. This is beneficial because the overlapping region of the selected video image frame can be the least usable region, e.g., for image analysis, because if other parameters are the same, the longer the distance to the sub-portion of the scene will result in a lower resolution of the sub-portion of the scene in the overlapping region.
[0015] The method of the first aspect can further comprise identifying one or more objects in the respective identified overlap region in each of the two video image frames. The act of selecting a video image frame of the two video image frames can comprise selecting a video image frame of the two video image frames that has an image sensor that is farthest away from the identified one or more objects in the scene. This is beneficial because the overlap region of the selected video image frame can be the least usable region, e.g. for image analysis, because the longer the distance to the one or more objects, if other parameters are the same, will result in a lower resolution of the one or more objects in the overlap region.
[0016] The method of the first aspect can further comprise identifying one or more objects in the respective identified overlap region in each of the two video image frames. The act of selecting a video image frame of the two video image frames can comprise selecting a video image frame of the two video image frames that has a poor object classification, a poor object identification, a poor object pose, or a poor re-identification vector. For example, a poor pose can mean that the object is least close to the front. This is beneficial because the overlap region of the selected video image frame can be the least usable region in terms of at least one of object classification, object identification, object pose, and object re-identification.
[0017] In the act of setting the compression levels, the respective compression levels of the blocks of pixels in the overlap region in the selected video image frame can further be set higher than the respective compression levels of the blocks of pixels in the overlap region in the non-selected video image frame of the two video image frames. This is beneficial because the overlap region of the non-selected video image frame can be the most usable region, e.g. for image analysis.
[0018] In the act of setting the compression levels, the respective compression levels of the blocks of pixels in the overlap region in the non-selected video image frame can be set based on a given principle, and wherein the respective compression levels of the blocks of pixels in the overlap region in the non-selected video image frame are selectively set lower than the respective compression levels that would have been set based on the given principle, wherein the combined bit number of the encoded two video image frames is equal to or lower than the bit number that the two video image frames would have had if the compression levels would have been set based on the given principle only. By keeping the bit number equal or lower for the two video image frames, a higher quality can be achieved in the overlap region of the non-selected video image frame without increasing and even optionally reducing the bit number of the two video image frames. This is beneficial, e.g. when the same subpart of the scene depicted by the overlap region is of particular interest.
[0019] The method of the first aspect can further comprise transmitting the encoded two video image frames to a common receiver.
[0020] According to a second aspect, there is provided a method for encoding two video image frames captured by respective ones of two image sensors, wherein each of the two video image frames depicts a respective portion of a scene. The method of the second aspect comprises identifying a respective overlap region in each of the two video image frames, the two overlap regions depicting a same sub-portion of the scene, and selecting a video image frame of the two video image frames. The method of the second aspect further comprises setting a compression level for the two video image frames, wherein a respective compression level is set for a block of pixels in the selected video image frame based on a given principle, and wherein the respective compression level of a block of pixels in the overlap region in the selected video image frame is selectively set lower than the respective compression level that would have been set based on the given principle. This is for example beneficial when the same sub-portion of the scene that the overlap region describes is of particular interest.
[0021] The given principle may, for example, be that a respective compression level is set for a block of pixels in the selected video image frame based on a respective attribute value associated with the block of pixels in the selected video image frame. Such an attribute may, for example, be a respective interest level associated with the block of pixels in the selected video image frame.
[0022] The interest level represents a relative interest level or importance of different regions of an image frame. What is considered to have a relatively high interest or importance, and what is considered to have a relatively low interest or importance, will depend on the application.
[0023] The method of the second aspect can further comprise encoding the two video image frames.
[0024] In the act of setting a compression level of the method of the second aspect, a respective compression level of a block of pixels in a non-overlap region in the selected video image frame can be selectively set higher than the respective compression level that would have been set based on the given principle, wherein the number of bits of the selected video image frame is equal to or lower than the number of bits the selected video image frame would have had if the compression levels would have been set based on the given principle only. By keeping the number of bits of the selected video image frame equal or lower, a higher quality can be achieved in the overlap region of the selected video image frame without increasing, even optionally decreasing, the number of bits of the selected video image frame. This is for example beneficial when the overlap region of the selected video image frame is of particular interest.
[0025] According to a third aspect, there is provided a non-transitory computer- readable storage medium having stored thereon instructions for implementing the method of the first aspect or the second aspect when executed in a system having at least two image sensors, at least one processor, and at least one encoder.
[0026] The above optional features of the method according to the first aspect also apply to this third aspect, when applicable.
[0027] According to a fourth aspect, there is provided an image processing apparatus for encoding two video image frames captured by respective ones of two image sensors, wherein each of the two video image frames depicts a respective portion of a scene. The image processing apparatus comprises circuitry configured to perform: an identifying function configured to identify respective overlapping areas in each of the two video image frames, the two overlapping areas depicting a same sub-portion of the scene; a selecting function configured to select a video image frame of the two video image frames; and a setting function configured to set compression levels for the two video image frames, wherein a respective compression level is set for a block of pixels in the selected video image frame based on a given principle, and wherein the respective compression level of a block of pixels in the overlapping area in the selected video image frame is selectively set higher than the respective compression level that would have been set based on the given principle. The image processing apparatus further comprises at least one encoder for encoding the two video image frames.
[0028] The given principle may, for example, be that the respective compression level is set for a block of pixels in the selected video image frame based on a respective property value associated with the block of pixels in the selected video image frame. Such a property may, for example, be a respective interest level associated with the block of pixels in the selected video image frame.
[0029] The interest level represents a relative interest level or importance of different areas of an image frame. What is considered to have a relatively high interest or importance, and what is considered to have a relatively low interest or importance, will depend on the application.
[0030] The image processing apparatus of the fourth aspect, the selecting function may further be configured to select the video image frame of the two video image frames based on one of an image property of the respective overlapping area in each of the two video image frames, an image sensor property of each of the two image sensors, and an image content of the respective overlapping area in each of the two video image frames.
[0031] In the setting function, the respective compression level of a block of pixels in the overlapping area in the selected video image frame may further be set higher than the respective compression level of a block of pixels in the overlapping area in the non-selected video image frame of the two video image frames.
[0032] In the setting function, the respective compression levels for the pixel blocks in the unselected video image frames can be set based on the given principle, and the respective compression levels for the pixel blocks in the overlapping area in the unselected video image frames can be selectively set to be lower than the respective compression levels that would have been set based on the given principle, wherein the combined bit number of the two encoded video image frames is equal to or lower than the bit number that the two video image frames would have had if the compression levels would have been set based on the given principle only.
[0033] The image processing apparatus of the fourth aspect can further comprise a transmitter for transmitting the two encoded video image frames to a common receiver.
[0034] The above further optional features of the method of the first aspect also apply to this third aspect, when applicable.
[0035] According to a fifth aspect, there is provided a camera system. The camera system comprises the image processing apparatus of the fourth aspect and two image sensors configured to capture a respective one of the two video image frames.
[0036] It should be understood, therefore, that the invention is not limited to the particular components described or the actions described as such apparatus and methods can vary. It is also to be understood that the terminology used herein is for the purpose of describing particular embodiments only and is not intended to be limiting. It must be noted that, as used in the specification and the appended claims, the articles "a," "an," and "the" are intended to mean one or more unless otherwise indicated. Furthermore, the words "comprise," "include," and "contain" do not exclude other elements or steps. BRIEF DESCRIPTION OF DRAWINGS
[0037] The above and other aspects of the present invention will now be described in more detail, with reference to the appended drawings. The drawings should not be considered limiting; the purpose is to explain and understand.
[0038] Figure 1 A flowchart is shown in relation to an embodiment of a method of the present disclosure.
[0039] Figure 2 A flowchart is shown in relation to an embodiment of another method of the present disclosure.
[0040] Figures 3a to 3d A schematic is shown in relation to an overlapping area in two video image frames.
[0041] Figure 4 A schematic is shown in relation to an embodiment of an image processing apparatus of the present disclosure, which image processing apparatus is optionally comprised in a camera system of the present disclosure. DETAILED DESCRIPTION
[0042] The present application will now be described in connection with the following drawings in which:
[0043] The present application is applicable to a scenario in which two video image frames are captured by respective ones of two image sensors, each of the two video image frames depicts a respective portion of a scene, and each of the two video image frames comprises a respective overlap region, the two overlap regions depicting a same sub-portion of the scene. Each of the two video image frames can for example be one of a sequence of independent video image frames. The two video image frames are to be encoded, stored and / or transmitted, and thus it is of interest to limit the number of bits to be stored and / or transmitted. At the same time, it can be important (e.g. in relation to subsequent analysis) that each portion of the video image frames still has a desired quality after storage and / or transmission. In the method of the first aspect, the redundancy in the form of the overlap regions in the two video image frames is used, and the total number of bits after encoding of the two video image frames is reduced by selectively setting a higher compression in the overlap region in a selected one of the two video image frames. Because the overlap region in the non-selected one of the two video image frames depicts the same sub-portion of the scene, the overlap portion of the non-selected one of the two video image frames can provide the desired quality in relation to the same sub-portion of the scene. In the method of the second aspect, the identification of the overlap regions in the two video image frames is used, and the quality is improved by selectively setting a lower compression in the overlap region in a selected one of the two video image frames. The total number of bits after encoding of the two video image frames can then optionally be reduced by increasing the compression in the non-overlap region in the selected one of the two video image frames or in the overlap region in the non-selected one of the two video image frames.
[0044] In connection with Figure 1 and Figures 3a to 3d An embodiment of the method 100 of the first aspect for encoding two video image frames captured by respective ones of two image sensors will be discussed. Each of the two video image frames depicts a respective portion of a scene. The steps of the method 100 can be performed by the image processing apparatus 400 described in connection with Figure 4 The image processing apparatus 400 described in connection with
[0045] Two image sensors can be arranged in two separate cameras or in a single camera. In the former case, the two different cameras can be, for example, two fixed cameras, arranged such that they capture two video image frames comprising an overlapping area depicting the same sub-section of the scene. For example, the two cameras can be two surveillance cameras, arranged to depict corresponding parts of the scene. In the latter case, a single camera can be, for example, arranged to capture two video image frames in the same sub-section of the scene, which are to be used to generate a panoramic video image frame by stitching the two overlapping video image frames together. Furthermore, when the sensors are arranged in two separate cameras, the two video image frames can be intended to be used to generate a panoramic video image frame by stitching the two overlapping video image frames together.
[0046] Method 100 includes identifying corresponding overlapping regions in each of two video image frames in S110, the two overlapping regions depicting the same sub-part of the scene.
[0047] Depicting identical sub-parts of a scene, each overlapping region is a two-dimensional representation of the same three-dimensional sub-part of the scene. Figures 3a to 3d A schematic diagram illustrating the relationship between overlapping regions in two video image frames is shown. Figure 3a In the image, an example is shown, with the first video image frame 310 represented by dashed lines and the second video image frame 315 represented by solid lines. Each of the first video image frame 310, captured by the first sensor, and the second video image frame 315, captured by the second sensor, includes an overlapping region 320 depicting the same sub-part of the scene. The first image sensor can be arranged in camera 330 and the second image sensor can be arranged in camera 335, such that the sensors are arranged as... Figure 3b The figures shown are parallel in the horizontal direction. Figure 3b The overlapping portion 320 produced by the arrangement is the result of the overlap 340 between the viewpoints of the first camera 330 and the second camera 335. Figure 3b As illustrated in the diagram, an alternative arrangement could be made, where the first and second image sensors are housed within a single camera. The first image sensor could be positioned in camera 350 and the second image sensor in camera 355, such that the sensors... Figure 3c The diagram shown is perpendicular in the horizontal direction. Figure 3c The overlapping portion 320 produced by the arrangement is the result of the overlap 360 between the horizontal viewpoint of the first camera 330 and the horizontal viewpoint of the second camera 335. Figure 3c The first camera 330 and the second camera 335 each have a fisheye lens with a 180° horizontal field of view. Relative to Figure 3bThe arrangement of the elements with a smaller horizontal viewing angle will result in a relatively large overlap area 320. Figure 3c The arrangement can be part of an arrangement including two additional cameras, each with a fisheye lens having a 180° horizontal field of view and arranged perpendicular to each other. Each of the two additional cameras includes a corresponding image sensor and is arranged perpendicular to the [image sensor]. Figure 3c One of the two cameras 350 and 355. Therefore, the four image sensors will capture video image frames that collectively depict a 360° view. The first image sensor can be arranged in camera 370 and the second image sensor can be arranged in camera 375, such that the sensors... Figure 3c The two cameras, as shown in the diagram, are facing each other horizontally. This could be, for example, when two cameras are positioned at opposite corners or sides of a square. Figure 3d The overlapping portion 380 produced by the arrangement is the result of the overlap 348 between the viewpoints of the first camera 370 and the second camera 375. The first and second image sensors can be arranged at any angle to each other, and each of the first video image frames 310 and 315 will include an overlapping portion, provided that the horizontal viewpoints of the first and second cameras are sufficiently large. The similarity of the overlapping regions of the two video image frames will depend on the angle between the first image sensor 310 and the second image sensor 315 and / or on the two video image frames after different projection transformations. It should be noted that even in... Figure 3d and Figure 3b The camera shown is only related to the angle between the sensors in the horizontal direction. The camera could, of course, be arranged such that the first and second image sensors are also arranged at an angle to each other in the vertical direction. The overlap 320 will also depend on the angle between the first and second image sensors, as well as the vertical viewing angles of the first and second image sensors.
[0048] The overlapping region in each of the two video image frames can be identified by means of real-time computation, based on the current position, orientation and focal length of each of the cameras comprising the two image sensors that captured the two video image frames. In an alternative, the overlapping region in each of the two video image frames can be identified in advance, if the two video image frames are captured by two image sensors arranged in cameras that each have a known mounting position, a known pan position, a known tilt position, a known vertical angle of view and a known horizontal angle of view, which are all fixed. If computed in advance, the identification of the respective overlapping region in each of the two video image frames, which both depict the same sub-portion of the scene, is performed once and then can be reused at each subsequent execution of the method 100, as long as the cameras are fixedly mounted and are not pan-tilt-zoom cameras. As a further alternative, the overlapping region can be determined based on image analysis of the two video image frames. For example, from object detection / tracking in the two video image frames, it can be inferred that the same object is located in both video image frames at the same time. From this it can in turn be inferred that there is an overlap between the two video image frames and that the overlapping region can be determined.
[0049] The method further comprises selecting S120 a video image frame of the two video image frames. In its broadest sense, the selection does not need to be based on any particular criteria, as long as one of the video image frames is selected.
[0050] However, the selection can be based on one or more criteria that result in the selection of the video image frame whose overlapping region is least suitable for a particular application, such as a surveillance application. For example, in a surveillance application, one or more criteria related to properties that affect the likelihood of image analysis can be used.
[0051] The selection S120 can be based on image properties of the respective overlapping region in each of the two video image frames. Such image properties can for example be focus, resolution, angular resolution, dynamic range, sensitivity, motion blur and color representation. In accordance with the method 100, in the overlapping region of the selected video image frame, the compression level is selectively set S130 to be higher, the video image frame of the two video image frames can be selected S120 that has one or more of the lowest focus, the lowest resolution, the lowest angular resolution, the lowest dynamic range, the lowest sensitivity, the most motion blur and the lowest quality of color representation in its overlapping region.
[0052] Alternatively or additionally, the selection S120 can be based on image sensor properties of each of the two image sensors. Such image properties can for example be resolution, dynamic range, color representation and light sensitivity. According to the method 100, in the overlapping region of the selected video image frames, the compression level is selectively set S130 to be higher, the video image frame of the two video image frames captured by the image sensor can be selected S120 which has one or more of the lowest resolution, the lowest dynamic range, the lower quality color representation and the lowest light sensitivity.
[0053] Alternatively or additionally, the selection S120 can be based on image content of the respective overlapping region in each of the two video image frames. For example, the selection S120 can be based on one or more objects in the overlapping region in each of the two video image frames. The method 100 then further comprises identifying S115 the one or more objects in the respective identified overlapping region in each of the two video image frames. According to the method 100, in the overlapping region of the selected video image frames, the compression level is selectively set S130 to be higher, the video image frame of the two video image frames can be selected S120 in which the one or more objects are most occluded, the one or more objects are least identifiable, the object classification is poorer, the object identification is poorer, the re-identification vector is poorer, or the selected video image frame is captured by the image sensor having the farthest distance to the identified one or more objects in the scene. The re-identification vector being poorer can mean that the re-identification vector has a shorter re-identification distance compared to a reference vector, or that the re-identification vector forms a better separated cluster of re-identification vectors.
[0054] Alternatively or additionally, the selection S120 can be based on a distance from each image sensor to a sub-portion of the scene. The distance from the image sensor to the sub-portion of the scene will affect the resolution in which each object is reproduced in the video image frame captured by the image sensor. According to the method 100, in the overlapping region of the selected video image frames, the compression level is selectively set S130 to be higher, the video image frame of the two video image frames having the farthest distance to the sub-portion of the scene can be selected S120 if higher resolution is preferred.
[0055] Once a video image frame has been selected, a compression level is set for the two image frames. The compression level is set for a pixel block in the selected video image frame based on a given principle. The given principle can be any compression principle for setting a compression value for a pixel block. The given principle can for example be that a respective compression level is set for a pixel block in the selected video image frame based on a respective property or property value associated with the pixel block in the selected video image frame. Such a property can for example be a respective interest level associated with the pixel block in the selected video image frame. Depending on the encoding standard, a pixel block is also referred to as a macroblock, a coding tree unit or others. The interest level or importance level can be set individually for each pixel block or for each region covering multiple pixel blocks. Depending on the application, the interest level or importance level can be determined based on different criteria. For video surveillance, a region with a certain type of motion, such as a moving person or an object entering the scene, can be considered to have a higher importance compared to other objects generating persistent motion such as a tree swaying in the wind, bushes, grass, flags, etc. Therefore, a pixel block related to such a higher importance object can have a higher interest level. Other examples of a pixel block that can be given a higher interest level are if it involves an identifiable part of an object (e.g. a face, a car license plate, etc.). An example of a pixel block that can be given a lower interest level is if the pixel block involves a part of the image where there is very much motion resulting in motion blur, too low or too high light level (overexposure), out of focus, high noise level, etc. Another example of a pixel block that can be given a higher interest level is if it is located in a region comprising one or more detected objects or in a region indicated as an area of interest, e.g. based on image analysis.
[0056] Other given principles can be, for example at pixel block level, to set a compression level for a pixel block in the selected video image frame based on:
[0057] - frame type of the frame the pixel block is located in (I-frame, P-frame, B-frame, intra refresh, fast forward frame, non-reference P-frame, empty frame)
[0058] - probability of motion in the pixel block
[0059] - probability that the pixel block is part of the background (change in frequency distribution of the pixel block, i.e. not based on motion detection but on spatial detection)
[0060] - compression value of the pixel block in a previous frame
[0061] - compression value of neighboring pixel blocks of the pixel block
[0062] - Number of previous frames in which image data of the pixel block has occurred (for pan-tilt-zoom cameras)
[0063] - Number of subsequent frames in which image data of the pixel block is estimated to occur (for pan-tilt-zoom cameras)
[0064] - Whether quality improvement is required
[0065] - Short-term and / or long-term bitrate target
[0066] - Luminance level of the pixel block
[0067] - Variance of the pixel block
[0068] - Detection based on video analysis or other analysis such as radar and audio analysis
[0069] - Focus level of the pixel block
[0070] - Motion blur level of the pixel block
[0071] - Noise level of the pixel block
[0072] - Number of frames in which the pixel block is expected to be encoded as an I-block
[0073] In the setting of compression levels, the respective compression levels of the pixel blocks in the overlapping area in the selected video image frame are selectively set S130 to be higher than the respective compression levels that would have been set based on the given principle, such as based on the respective interest level associated with the pixel blocks in the overlapping area in the selected video image frame. Thus, a reduction of the number of bits of the selected video image frame after compression will be achieved compared to the case where the compression levels are both set based on the given principle, such as based on the respective interest level associated with the pixel blocks in the overlapping area in the selected video image frame.
[0074] When encoding an image frame divided into a plurality of pixel blocks into a video, a compression level (e.g. in the form of a compression value) is set for each pixel block. An example of such a compression value is the quantization parameter (QP) value used in the H.264 and H.265 video compression standards. The compression level can be absolute or relative. If the compression level is relative, it is expressed as a compression level offset relative to a reference compression level. The reference compression level is the compression level that has been chosen as a reference from which the level offset is set. The reference compression level can for example be the expected average or median compression value over time, the maximum compression value, the minimum compression value, etc. The reference compression level offset can be a negative value, a positive value or "0". For example, an offset QP value can be set relative to a reference QP value. Setting the respective compression value higher is selectively setting the offset QP value higher. The set offset QP value for each pixel block can then be provided in a quantization parameter map (QMAP) that is used to instruct an encoder to encode the image frame using the set offset QP value according to the QMAP.
[0075] The respective compression level of the pixel blocks in the overlap region in the selected video image frame can be set higher than the respective compression level of the pixel blocks in the overlap region in the unselected video image frame of the two video image frames.
[0076] The respective compression level of the pixel blocks in the unselected video image frame can be set based on a given principle, such as based on the respective interest level associated with the pixel blocks in the unselected video image frame. Furthermore, the respective compression level of the pixel blocks in the overlap region in the unselected video image frame can be selectively set lower than the respective compression level that would have been set based on the given principle, such as based on the respective interest level associated with the pixel blocks in the overlap region in the unselected video image frame. Selecting a higher compression level results in a combined bit number of the encoded two video image frames being equal to or lower than the bit number the two video image frames would have had if the compression levels were set based on the given principle only, such as based on the respective interest level associated with the pixel blocks in the overlap region in the selected video image frame and in the unselected video image frame only. This results in a higher quality of the overlap portion of the unselected video image frame than would have been obtained if the respective compression level would have been set based on the given principle, such as based on the respective interest level associated with the pixel blocks in the overlap region in the unselected video image frame. At the same time, the combined bit number of the encoded two video image frames is not increased and can be reduced.
[0077] The method 100 further comprises encoding S140 the two video image frames. The two video image frames can be encoded by two independent encoders or they can be encoded by a single encoder. When encoding, the compression level set for the pixel blocks of the video image frames is used for the compression encoding of the video image frames.
[0078] In case the two video image frames are encoded with two independent encoders, the selection S120 of one of the two video image frames can be based on properties of the two independent encoders. For example, in the act of selection S120, the video image frame of the two video image frames to be encoded in the encoder with the lower compression level of the two encoders can be selected. Further, in the act of selection S120, the video image frame of the two video image frames to be encoded in the encoder not supporting the desired encoding standard can be selected. The latter is relevant if the not selected video image frame is to be encoded in the encoder supporting the desired encoding standard.
[0079] The method 100 can further comprise storing the two encoded video image frames in a memory and / or transmitting to a common receiver. After storing and / or transmitting, the two encoded video image frames can for example be decoded and watched and / or analyzed (e.g. at the common receiver) or they can be stitched together to form a single video image frame.
[0080] In connection with Figure 3c and Figure 2 embodiments of a method 200 for encoding two video image frames captured by respective ones of two image sensors will be discussed. Each of the two video image frames depicts a respective portion of a scene. The steps of the method 200 can be performed by the image processing apparatus 400 described in connection with Figures 3a to 3d The steps of the method 200 can be performed by the image processing apparatus 400 described in connection with
[0081] The two image sensors can be arranged in two independent cameras or in a single camera. In the former case, the two different cameras can for example be two stationary cameras arranged such that they capture two video image frames comprising an overlapping area depicting the same sub-portion of a scene. For example, the two cameras can be two surveillance cameras arranged to depict respective portions of a scene. In the latter case, the single camera can for example be arranged to capture two video image frames at the same sub-portion of a scene, the two video image frames to be used to produce a panoramic video image frame by stitching the two overlapping video image frames together. Further, when the sensors are arranged in two independent cameras, the two video image frames can be intended to be used to produce a panoramic video image frame by stitching the two overlapping video image frames together.
[0082] Method 200 includes identifying corresponding overlapping regions in each of two video image frames in S210, the two overlapping regions depicting the same sub-part of the scene.
[0083] Depicting identical sub-parts of a scene, each overlapping region is a two-dimensional representation of the same three-dimensional sub-part of the scene. Figure 4 A schematic diagram illustrating the relationship between overlapping regions in two video image frames is shown. Figures 3a to 3d In the image, an example is shown, with the first video image frame 310 represented by dashed lines and the second video image frame 315 represented by solid lines. Each of the first video image frame 310, captured by the first sensor, and the second video image frame 315, captured by the second sensor, includes an overlapping region 320 depicting the same sub-part of the scene. The first image sensor can be arranged in camera 330 and the second image sensor can be arranged in camera 335, such that the sensors are arranged as... Figure 3a The figures shown are parallel in the horizontal direction. Figure 3b The overlapping portion 320 produced by the arrangement is the result of the overlap 340 between the viewpoints of the first camera 330 and the second camera 335. Figure 3b As illustrated in the diagram, an alternative arrangement could be made, where the first and second image sensors are housed within a single camera. The first image sensor could be positioned in camera 350 and the second image sensor in camera 355, such that the sensors... Figure 3b The diagram shown is perpendicular in the horizontal direction. Figure 3c The overlapping portion 320 produced by the arrangement is the result of the overlap 360 between the horizontal viewpoint of the first camera 330 and the horizontal viewpoint of the second camera 335. Figure 3c The first camera 330 and the second camera 335 each have a fisheye lens with a 180° horizontal field of view. Relative to Figure 3c The arrangement of the elements with a smaller horizontal viewing angle will result in a relatively large overlap area 320. Figure 3b The arrangement can be part of an arrangement including two additional cameras, each with a fisheye lens having a 180° horizontal field of view and arranged perpendicular to each other. Each of the two additional cameras includes a corresponding image sensor and is arranged perpendicular to the [image sensor]. Figure 3c One of the two cameras 350 and 355. Therefore, the four image sensors will capture video image frames that collectively depict a 360° view. The first image sensor can be arranged in camera 370 and the second image sensor can be arranged in camera 375, such that the sensors... Figure 3c The two cameras, as shown in the diagram, are facing each other horizontally. This could be, for example, when two cameras are positioned at opposite corners or sides of a square. Figure 3cThe overlap 380 created by the arrangement in Fig. 3 is a result of the overlap 348 between the field of view of the first camera 370 and the field of view of the second camera 375. The first and second image sensors can be arranged at any angle to each other, and as long as the horizontal field of view of the first camera and the horizontal field of view of the second camera are large enough, each of the first and second video image frames 310, 315 will comprise an overlap. The similarity of the overlap area of the two video image frames will depend on the angle between the first and second image sensors 310, 315 and / or the two video image frames being subject to different projection transformations. It should be noted that even in the case where the first and second image sensors are arranged at an angle to each other in the horizontal direction only, the cameras can of course also be arranged such that the first and second image sensors are arranged at an angle to each other also in the vertical direction. The overlap 320 will also depend on the angle between the first and second image sensors and the vertical field of view of the first image sensor and the vertical field of view of the second image sensor. Figure 3d and Figure 3d The cameras shown in Fig. 3 are only related to the angle between the sensors in the horizontal direction, the cameras can of course also be arranged such that the first and second image sensors are arranged at an angle to each other also in the vertical direction. The overlap 320 will also depend on the angle between the first and second image sensors and the vertical field of view of the first image sensor and the vertical field of view of the second image sensor.
[0084] The overlap area in each of the two video image frames can be identified by means of real-time computation based on the current position, orientation and focal length of each of the cameras comprising the two image sensors capturing the two video image frames. In an alternative, the overlap area in each of the two video image frames can be identified in advance if the two video image frames are captured by two image sensors arranged in cameras each having a known mounting position, a known pan position, a known tilt position, a known vertical field of view and a known horizontal field of view, all of which are fixed. If computed in advance, identifying the respective overlap area in each of the two video image frames (the two overlap areas depict the same sub-portion of the scene) is performed once and can then be reused every time the method 100 is performed subsequently, as long as the cameras are fixed mounted and not pan-tilt-zoom cameras. As a further alternative, the overlap area can be determined based on image analysis of the two video image frames. For example, from object detection / tracking in the two video image frames, it can be inferred that the same object is located in both video image frames at the same time. From this it can in turn be inferred that there is an overlap between the two video image frames and the overlap area can be determined.
[0085] The method 200 further comprises selecting S220 a video image frame of the two video image frames. In its broadest sense, the selection does not need to be based on any particular criteria, as long as one of the video image frames is selected.
[0086] However, the selection can be based on one or more criteria that result in the selected overlap region being best suited for a particular application, such as a surveillance application. For example, in a surveillance application, one or more criteria related to properties that affect the likelihood of image analysis can be used.
[0087] The selection S220 can be based on image properties of the respective overlap region in each of the two video image frames. Such image properties can for example be focus, resolution, angular resolution, dynamic range, sensitivity, motion blur, and color representation. According to the method 200, in the overlap region of the selected video image frame, the compression level is selectively set S230 to be lower, the selection S220 can select the video image frame of the two video image frames that has one or more of highest focus, highest resolution, highest angular resolution, highest dynamic range, highest sensitivity, least motion blur, and higher quality color representation in its overlap region.
[0088] Alternatively or additionally, the selection S220 can be based on image sensor properties of each of the two image sensors. Such image properties can for example be resolution, dynamic range, color representation, and sensitivity. According to the method 200, in the overlap region of the selected video image frame, the compression level is selectively set S230 to be lower, the selection S220 can select the video image frame of the two video image frames captured by the image sensors that has one or more of highest resolution, highest dynamic range, higher quality color representation, and highest sensitivity.
[0089] Alternatively or additionally, the selection S220 can be based on image content of the respective overlap region in each of the two video image frames. For example, the selection S220 can be based on one or more objects in the overlap region in each of the two video image frames. The method 200 then further comprises identifying S215 the one or more objects in the respective identified overlap region in each of the two video image frames. According to the method 200, in the overlap region of the selected video image frame, the compression level is selectively set S230 to be lower, the selection S220 can select the video image frame of the two video image frames in which the one or more objects are most occluded, the one or more objects are least identifiable, the object classification is poorer, the object identification is better, the re-identification vector is better, or the selected video image frame is captured by the image sensor that has the shortest distance to the identified one or more objects in the scene. The re-identification vector is better can mean that the re-identification vector has a longer re-identification distance compared to a reference vector, or that the re-identification vector forms a more poorly separated cluster of re-identification vectors.
[0090] Alternatively or additionally, the selection S220 can be based on a distance from each image sensor to the sub-portion of the scene. The distance from an image sensor to the sub-portion of the scene will affect the resolution in which each object is reproduced in the video image frames captured by the image sensor. According to the method 200, in the overlapping area of the selected video image frames, the compression level is selectively set S230 to be lower, and if higher resolution is preferred, the video image frame of the two video image frames with the shortest distance to the sub-portion of the scene can be selected S220.
[0091] Once a video image frame has been selected, the compression level is set for the two image frames. The respective compression level is set for the blocks of pixels in the selected video image frame based on a given principle. The given principle can for example be that the respective compression level is set for the blocks of pixels in the selected video image frame based on a respective property or property value associated with the blocks of pixels in the selected video image frame. Such a property can for example be a respective interest level associated with the blocks of pixels in the selected video image frame. Depending on for example the coding standard, the blocks of pixels can also be referred to as macroblocks, coding tree units or others. The interest level or importance level can be set individually for each block of pixels, or for each region covering a plurality of blocks of pixels. Depending on the application, the interest level or importance level can be determined based on different criteria. For video surveillance, a region with a certain type of motion, such as a moving person or an object entering the scene, can be considered to have a higher importance compared to other objects generating persistent motion such as a tree swaying in the wind, bushes, grass, flags, etc. Therefore, a block of pixels related to such a higher importance object can have a higher interest level. Other examples of blocks of pixels that can be given a higher interest level are if it involves an identifiable part of an object (e.g. a face, a car license plate, etc.). An example of a block of pixels that can be given a lower interest level is if it involves a part of the image in which there is very much motion resulting in motion blur, a too low or a too high (overexposed) light level, a defocus, a high noise level, etc.
[0092] In the setting of the compression level, the respective compression level of the blocks of pixels in the overlapping area in the selected video image frame is selectively set 230 to be lower (as opposed to higher in the method 100) than the respective compression level that would have been set based on the given principle, such as based on the respective interest level associated with the blocks of pixels in the overlapping area in the selected video image frame. Thus, in the method 200, an improvement of the quality of the overlapping area of the selected video image frame after compression will be achieved compared to the case where the compression levels are both set based on the given principle, such as based on the respective interest level associated with the blocks of pixels in the overlapping area in the selected video image frame as well.
[0093] When encoding an image frame divided into a plurality of pixel blocks into a video, a compression level (e.g. in the form of a compression value) is set for each pixel block. An example of such a compression value is the quantization parameter (QP) value used in the H.264 and H.265 video compression standards. The compression level can be absolute or it can be relative. If the compression level is relative, it is expressed as a compression level offset relative to a reference compression level. The reference compression level is the compression level that has been chosen as a reference, from which the level offset is set. The reference compression level can for example be the expected average or median compression value over time, the maximum compression value, the minimum compression value, etc. The reference compression level offset can be a negative value, a positive value or “0”. For example, an offset QP value can be set relative to a reference QP value. Setting the respective compression value higher is selectively setting the offset QP value higher. The set offset QP value for each pixel block can then be provided in a quantization parameter map (QMAP) for instructing an encoder to encode the image frame using the set offset QP value according to the QMAP.
[0094] The respective compression level of the pixel blocks in the non-overlapping region in the selected video image frame can be selectively set higher than it would have been set based on a given principle, such as based on the respective interest level associated with the pixel blocks in the non-overlapping region in the selected video image frame. Selecting a higher compression level results in the encoded selected video image frame having a number of bits equal to or lower than the number of bits the selected video image frame would have had if the compression level had been set based on the given principle only, such as based on the respective interest level associated with the pixel blocks in the non-overlapping region and the overlapping region in the selected video image frame only. Thus, the combined number of bits of the encoded selected video image frames is not increased and can be reduced.
[0095] The method 200 further comprises encoding S240 the two video image frames. The two video image frames can be encoded by two separate encoders or they can be encoded by a single encoder.
[0096] In connection with Figure 3b Embodiments of the image processing apparatus 400 of the fourth aspect for encoding two video image frames captured by respective ones of two image sensors and embodiments of the camera system 450 of the fifth aspect will be discussed. Each of the two video image frames depicts a respective portion of a scene. The steps of the method 100 can be performed by the image processing apparatus 400 described in connection with Figure 3c the fourth aspect.
[0097] The image processing apparatus 400 comprises an encoder 410 and a circuit 420. The circuit 420 is configured to perform the functions of the image processing apparatus 400. The circuit 420 can comprise a processor 422, such as a central processing unit (CPU), a microcontroller, or a microprocessor. The processor 422 is configured to execute program code. The program code can for example be configured to perform the functions of the image processing apparatus 400.
[0098] The image processing apparatus 400 can further comprise a memory 430. The memory 430 can be one or more of a buffer, a flash memory, a hard drive, a removable media, a volatile memory, a non-volatile memory, a random access memory (RAM), or other suitable device. In a typical arrangement, the memory 430 can include a non-volatile memory for long term data storage and a volatile memory used as a device memory for the circuit 420. The memory 430 can exchange data with the circuit 420 over a data bus. There can also be accompanying control lines and an address bus between the memory 430 and the circuit 420.
[0099] The functions of the image processing apparatus 400 can be implemented in the form of executable logic routines, e.g., lines of code, software programs, etc., stored in a non-transitory computer readable medium (e.g., the memory 430) of the image processing apparatus 400 and executed by the circuit 420 (e.g., using the processor 422). Further, the functions of the image processing apparatus 400 can be standalone software applications or form part of software applications that perform additional tasks related to the image processing apparatus 400. The described functions can be considered methods that the processing unit (e.g., the processor 422 of the circuit 420) is configured to perform. Moreover, while the described functions can be implemented in software, such functions can also be performed via special-purpose hardware or firmware, or some combination of hardware, firmware, and / or software.
[0100] The encoder 410 can for example be adapted for encoding according to the H.264 or H.265 video compression standard.
[0101] The circuit 420 is configured to perform an identifying function 442, a selecting function 444, and a setting function 446.
[0102] The identifying function 442 is configured to identify a respective overlap region in each of the two video image frames, the overlap region depicting a same sub-portion of the scene.
[0103] The selecting function 444 is configured to select a video image frame of the two video image frames.
[0104] The setting function 446 is configured to set compression levels for the two image frames, wherein respective compression levels are set for pixel blocks in the selected video image frames based on a given principle. The given principle can for example be that respective compression levels are set for pixel blocks in the selected video image frames based on respective properties or property values associated with the pixel blocks in the selected video image frames. Such a property can for example be a respective interest level associated with the pixel blocks in the selected video image frames. The respective compression levels of the pixel blocks in the overlap region in the selected video image frames are selectively set to be higher than the respective compression levels that would have been set based on the given principle, such as based on the respective interest levels associated with the pixel blocks in the overlap region in the selected video image frames.
[0105] The at least one encoder 410 is configured to encode the two video image frames.
[0106] The functions performed by the circuit 420 can further be adapted to corresponding steps of embodiments of the method described with respect to Figure 4 、 Figure 4 and Figure 1 .
[0107] In particular, the setting function 446 can be configured to set compression levels for the two image frames, wherein respective compression levels are set for pixel blocks in the selected video image frames based on a given principle, such as based on respective interest levels associated with the pixel blocks in the selected video image frames, and wherein the respective compression levels of the pixel blocks in the overlap region in the selected video image frames are selectively set to be lower than the respective compression levels that would have been set based on the given principle, such as based on the respective interest levels associated with the pixel blocks in the overlap region in the selected video image frames, for performing the method 200 described with respect to Figure 2 .
[0108] The image processing apparatus can optionally be arranged in a camera system 450 comprising the image processing apparatus 400 and two image sensors 460, 470 configured to capture respective ones of the two video image frames.
[0109] It is noted that even though Figures 3a to 3d Figure 2 Figure 4 the camera system 450 is depicted as comprising only one image processing apparatus 400, in an alternative the camera system 450 can comprise two image processing apparatuses 400, each processing a separate one of the two video image frames captured by a respective one of the first image sensor 460 and the second image sensor 470.
[0110] Those skilled in the art will realize that the application is not limited to the examples described above. Rather, the scope of the application is limited only by the claims that follow, within the full scope of equivalents to which such claims are entitled. Modifications and variations of the described embodiments are possible, as those skilled in the art will readily appreciate. Such modifications and variations are considered to be within the scope of the application as disclosed in the specification and claimed in the claims.
Claims
1. A method for encoding two video image frames, wherein, each of the two video image frames is captured by a respective one of two image sensors, wherein each of the two video image frames depicts a respective portion of a scene, the method comprising: identifying a respective overlap region in each of the two video image frames, the two overlap regions depicting a same sub-portion of the scene; selecting a video image frame of the two video image frames; setting compression values for the two image frames, wherein respective compression values for pixel blocks in a selected video image frame are set based on respective interest levels associated with the pixel blocks in the selected video image frame, and wherein respective compression values for pixel blocks in the overlap region in the selected video image frame are selectively set higher than respective compression values that would otherwise be set for the pixel blocks in the overlap region based on the respective interest levels associated with the pixel blocks in the overlap region; and encoding the two video image frames.
2. The method of claim 1, wherein, The act of selecting a video image frame of the two video image frames includes: selecting the video image frame of the two video image frames based on one of image properties of the respective overlap region in each of the two video image frames, image sensor properties of each of the two image sensors, and image content of the respective overlap region in each of the two video image frames.
3. The method of claim 1, further comprising: identifying one or more objects in the respective identified overlap region in each of the two video image frames, and wherein the act of selecting a video image frame of the two video image frames includes: selecting the video image frame of the two video image frames in which the one or more objects are most occluded or the objects are least identifiable.
4. The method of claim 1, wherein, The act of selecting a video image frame of the two video image frames includes: selecting the video image frame of the two video image frames that has one or more of the following image properties in the respective overlap region: lower quality focus; lowest resolution; lowest angular resolution; lowest dynamic range; lowest sensitivity; most motion blur; and lower quality color representation.
5. The method of claim 1, wherein, The act of selecting a video image frame of the two video image frames includes: selecting the video image frame of the two video image frames that is captured by the image sensor that is farthest from the sub-portion of the scene.
6. The method of claim 1, further comprising: identifying one or more objects in the respective identified overlap region in each of the two video image frames, and wherein the act of selecting a video image frame of the two video image frames includes: selecting the video image frame of the two video image frames that is captured by the image sensor that is farthest from the identified one or more objects in the scene.
7. The method of claim 1, further comprising: identifying one or more objects in the respective identified overlapping region in each of the two video image frames, and wherein the act of selecting a video image frame of the two video image frames comprises: selecting the video image frame of the two video image frames that is poorer in object classification, poorer in object identification, or poorer in re-identification vectors.
8. The method of claim 1, wherein, In the act of setting compression values, the respective compression values of the blocks of pixels in the overlapping region in the selected video image frame are further set to be higher than respective compression values of blocks of pixels in the overlapping region in the non-selected video image frame of the two video image frames.
9. The method of claim 1, wherein, In the act of setting compression values, respective compression values of blocks of pixels in the overlapping region in the non-selected video image frame are set based on respective interest levels associated with the blocks of pixels in the overlapping region, and wherein respective compression values of blocks of pixels in the overlapping region in the non-selected video image frame are set to be lower than respective compression values that would have been set for the blocks of pixels in the overlapping region based on the respective interest levels associated with the blocks of pixels in the overlapping region, wherein a combined bit number of the two video image frames that are encoded is equal to or lower than a bit number that the two video image frames would have if the compression values for the blocks of pixels in the overlapping region were set based only on the respective interest levels associated with the blocks of pixels in the overlapping region.
10. The method of claim 1, further comprising: transmitting the two video image frames that are encoded to a common receiver.
11. A non-transitory computer readable storage medium having instructions stored thereon, the instructions, when executed in an apparatus having at least one processor and at least one encoder, for implementing the method of claim 1.
12. An image processing apparatus for encoding two video image frames, wherein, each of the two video image frames is captured by a respective one of two image sensors, wherein each of the two video image frames depicts a respective portion of a scene, the image processing apparatus comprising: circuitry configured to perform: an identifying function configured to identify a respective overlapping region in each of two video image frames, the two overlapping regions depicting a same sub-portion of the scene; a selecting function configured to select a video image frame of the two video image frames; and a setting function configured to set compression values for the two image frames, wherein respective compression values of blocks of pixels in the selected video image frame are set based on respective interest levels associated with the blocks of pixels in the selected video image frame, and wherein respective compression values of blocks of pixels in the overlapping region in the selected video image frame are selectively set to be higher than respective compression values that would have been set for the blocks of pixels in the overlapping region based on the respective interest levels associated with the blocks of pixels in the overlapping region; and at least one encoder to encode the two video image frames.
13. A camera system comprising: the image processing apparatus of claim 12; and a camera configured to capture the two video image frames. The two image sensors are configured to capture respective ones of the two video image frames.
Citation Information
Patent Citations
Image data encoding / decoding method and apparatus
US20190230337A1
Image processing device, image processing method, program, and image transmission system
US20210152848A1