Fisheye image compression, fisheye video stream compression, and panoramic video generation method
By distinguishing between rendered and non-rendered areas of the fisheye image based on the positioning information of the rendering area at the decoding end, and by adopting a differentiated compression method, the problems of long time consumption and low clarity in panoramic video stitching are solved, achieving efficient image compression and saving hardware performance.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- ARASHI VISION INC
- Filing Date
- 2021-07-09
- Publication Date
- 2026-05-19
AI Technical Summary
Panoramic video stitching is time-consuming, has low resolution after compression, and requires high hardware performance. Existing methods cannot effectively reduce the time consumption while ensuring high resolution.
Based on the location information of the rendering area at the decoding end, the rendering area and non-rendering area on the fisheye image are determined, and they are compressed using different compression ratios. The rendering area may not be compressed or is under-compressed to generate a compressed image.
This reduces image transmission time, ensures high definition in the rendered area, and lowers the performance requirements of the compression hardware.
Smart Images

Figure CN115604528B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of video compression technology, and in particular to a method for fisheye image compression, fisheye video stream compression, and panoramic video generation. Background Technology
[0002] A fisheye lens is a lens with a focal length of 16mm or less and an angle of view close to, equal to, or greater than 180°. It is an extreme wide-angle lens, and "fisheye lens" is its common name. To achieve the maximum shooting angle, the front lens element of this type of lens has a very short diameter and protrudes forward in a parabolic shape, much like a fish's eye, hence the name "fisheye lens".
[0003] Currently, panoramic video stitching cameras typically use fisheye lenses as the acquisition device for panoramic video images, which are highly favored in the market due to their wide angle of view and high resolution. However, the high resolution of panoramic video images is not conducive to network transmission. Therefore, video image compression is necessary. Currently, panoramic video images are generally compressed directly, i.e., the obtained fisheye images are first stitched together to obtain the panoramic image, and then the panoramic image is compressed. However, the current method still has the following problems: 1. Panoramic stitching is time-consuming; 2. During the panoramic stitching process, interpolation sampling of the original fisheye images results in the loss of some information, leading to lower clarity after compression; 3. The generated panoramic stitched images are generally very large, placing high demands on the hardware performance of the compression end.
[0004] Therefore, it is necessary to provide an image compression method that reduces time consumption, ensures high resolution, and has low dependence on the performance of the compression end hardware. Summary of the Invention
[0005] Based on this, it is necessary to provide a method and apparatus for fisheye image compression, fisheye video stream compression, and panoramic video generation, as well as computer equipment and storage media, that can reduce time consumption, ensure high definition, and have low dependence on the performance of compression end hardware, in order to address the above-mentioned technical problems.
[0006] A method for compressing fisheye video streams, the method comprising:
[0007] Obtain the location information of the rendering area on the decoding end;
[0008] The fisheye rendering area on the corresponding fisheye image is determined based on the positioning information, and the area on the fisheye image other than the fisheye rendering area is the non-rendered fisheye area.
[0009] The fisheye image is compressed to obtain a compressed image; wherein the compression ratio of the fisheye rendered region in the compressed image is less than the compression ratio of the fisheye non-rendered region, and / or the fisheye rendered region is not compressed.
[0010] A fisheye video stream compression method includes:
[0011] Acquire fisheye video stream;
[0012] The fisheye image compression method described in the above embodiments is used to compress each frame of the fisheye video stream to obtain a compressed image of each frame of the fisheye image.
[0013] The video stream is compressed based on the compressed images of each frame of the fisheye video stream to obtain a compressed fisheye video stream.
[0014] A panoramic video generation method, comprising:
[0015] Obtain a compressed fisheye video stream; wherein the compressed fisheye video stream is obtained by processing the fisheye video stream compression method described in the above embodiments;
[0016] Based on the compressed fisheye video stream, a compressed image corresponding to multiple frames of the original fisheye image is obtained;
[0017] The compressed image is restored to obtain the original fisheye image;
[0018] By stitching together the original fisheye images, a panoramic video is obtained.
[0019] A fisheye video stream compression device, the device comprising:
[0020] The information transmission module is used to obtain the positioning information of the rendering area at the decoding end;
[0021] The compression region determination module is used to determine the fisheye rendering region on the corresponding fisheye image based on the positioning information, wherein the area on the fisheye image other than the fisheye rendering region is the fisheye non-rendering region.
[0022] A compression module is used to compress the fisheye image to obtain a compressed image; wherein the compression ratio of the fisheye rendering area in the compressed image is less than the compression ratio of the fisheye non-rendering area, and / or the fisheye rendering area is not compressed.
[0023] A fisheye video stream compression device, comprising:
[0024] The video stream acquisition module is used to acquire fisheye video streams;
[0025] The image compression module is used to compress each frame of the fisheye image in the fisheye video stream using the fisheye image compression method described in the above embodiments, so as to obtain a compressed image of each frame of the fisheye image.
[0026] The video stream compression module is used to compress the video stream based on the compressed images of each frame of the fisheye video stream to obtain a compressed fisheye video stream.
[0027] A panoramic video generation device, comprising:
[0028] A video stream acquisition module is used to acquire a compressed fisheye video stream; wherein the compressed fisheye video stream is obtained by processing the fisheye video stream compression method as described in the above embodiments;
[0029] The video stream decompression module is used to obtain compressed images corresponding to multiple frames of the original fisheye images based on the compressed fisheye video stream.
[0030] The restoration module is used to restore the compressed image to obtain the original fisheye image;
[0031] The stitching module is used to stitch together the original fisheye images to obtain a panoramic video.
[0032] A computer device includes a memory and a processor, the memory storing a computer program, the processor executing the computer program to implement the steps of any of the above methods.
[0033] A computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the steps of any of the methods described above.
[0034] The above-mentioned fisheye image compression method determines the fisheye rendering area on the fisheye image based on the positioning information of the rendering area at the decoding end. During compression, the compression ratio of the fisheye rendering area in the compressed image is less than the compression ratio of the non-rendered fisheye area, and / or the fisheye rendering area is not compressed, thereby achieving the purpose of compression. At the same time, it saves image transmission bandwidth, and the compression ratio of the compressed fisheye rendering area is small, or lossless compression, so that the final rendered image is not excessively compressed, resulting in a decrease in image quality, thus ensuring the clarity of the rendered area image. Attached Figure Description
[0035] Figure 1 This is a diagram illustrating the application environment of a fisheye video stream compression method in one embodiment.
[0036] Figure 2 This is a flowchart illustrating a fisheye video stream compression method in one embodiment;
[0037] Figure 3 This is a flowchart illustrating the steps of determining the fisheye rendering region on the corresponding fisheye image based on positioning information in one embodiment.
[0038] Figure 4 This is a schematic diagram of the distribution of the second two-dimensional point set on two fisheyes in one embodiment;
[0039] Figure 5 This is a schematic diagram of the distribution of a second two-dimensional point set on a fisheye in one embodiment;
[0040] Figure 6 In one embodiment Figure 4 The corresponding rendering diagram;
[0041] Figure 7 In one embodiment Figure 5 The corresponding rendering diagram;
[0042] Figure 8 This is a structural block diagram of a fisheye video stream compression device in one embodiment;
[0043] Figure 9 This is an internal structural diagram of a computer device in one embodiment. Detailed Implementation
[0044] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.
[0045] The fisheye video stream compression method provided in this application can be applied to, for example... Figure 1 In the application environment shown, the decoding end 102 communicates with the encoding end 104 via a network. The encoding end 104 is specifically an image acquisition device with a fisheye lens set up at the acquisition site. The decoding end is a processing device that receives the fisheye video stream and stitches it into a panoramic image; it can be VR glasses, a camera, a mobile phone, a computer, an iPad, etc., and is not limited to any particular device in this invention. In one embodiment, the decoding end is a VR glasses device. In this embodiment, at least two fisheye lenses are used, set up at the acquisition site, capable of stitching together a 360-degree panoramic view of the acquisition site. The VR glasses acquire the user's head movements, determine the current user's viewpoint, and determine the rendering area based on the current user's viewpoint. Specifically, the encoding end 104 acquires the positioning information of the rendering area from the decoding end 102; determines the corresponding fisheye rendering area on the fisheye image based on the positioning information; the area on the fisheye image other than the fisheye rendering area is the non-fisheye rendering area; compresses the fisheye image to obtain a compressed image; wherein the compression ratio of the fisheye rendering area in the compressed image is less than the compression ratio of the non-fisheye rendering area, and / or the fisheye rendering area is not compressed.
[0046] In one embodiment, such as Figure 2 As shown, a fisheye video stream compression method is provided, which is applied to... Figure 1 Taking the encoding end of the code as an example, the following steps are included:
[0047] Step S202: Obtain the positioning information of the rendering area at the decoding end.
[0048] Fisheye lenses typically have a field of view of 220° or 230°. Multiple fisheye lenses capture fisheye images, which are then stitched together to obtain a panoramic image. The decoder is the device used for decoding; it can be a VR headset that follows the VR viewpoint, rendering the image corresponding to that viewpoint. In other words, the rendering area is the region in the decoder's rendering screen where the fisheye image or the stitched panoramic image appears. Positioning information is used to locate the rendering area, such as Euler angles yaw and pitch, which represent the viewpoint direction and are used to identify the horizontal field of view (hFOV) and vertical field of view (vFOV) of the rendering area, respectively. Positioning information can be understood as defining the boundary of a region; in this embodiment, the positioning information defines the rendering area of the decoder.
[0049] Specifically, before compressing the acquired fisheye image, the encoding end obtains the location information of the current frame's rendering region from the decoding end. This location information can be obtained through communication between the decoding and encoding ends, with the decoding end sending it to the encoding end. Alternatively, the encoding end can obtain the location information of the decoding end's rendering region through communication with a third-party device. It should be understood that a set of location information can correspond to a single frame of fisheye image or multiple frames of fisheye image. The location information can be sent in every frame or every other frame.
[0050] Step S204: Determine the fisheye rendering area on the corresponding fisheye image based on the positioning information. The area on the fisheye image other than the fisheye rendering area is the non-rendering area of the fisheye.
[0051] Fisheye images are images captured by a fisheye lens. The biggest characteristic of a fisheye lens is its wide field of view, typically reaching 220° or 230°. Therefore, the resulting fisheye images are images with an extremely wide field of view at the capture site. Fisheye lenses can be used to capture fisheye video streams, and each frame in the fisheye video stream is a fisheye image.
[0052] Specifically, whenever the encoding end receives the location information of the rendering region of the current frame in the fisheye video stream, since the location information of the rendering region is sent frame by frame, the rendering region is projected onto the fisheye image of the frame corresponding to that location information. The area projected onto the fisheye image is the rendering region determined for this compression, i.e., the fisheye rendering region. The area on the fisheye image other than the fisheye rendering region is the non-rendered fisheye region. It can be understood that as the viewpoint of the VR glasses changes, the location information changes, and thus the requested rendering region changes, resulting in a different fisheye rendering region for each frame of the fisheye image.
[0053] Step S206: Compress the fisheye image to obtain a compressed image; wherein the compression ratio of the fisheye rendering area in the compressed image is less than the compression ratio of the non-rendered fisheye area, and / or the fisheye rendering area is not compressed.
[0054] Specifically, after the encoding end determines the fisheye rendering area on the fisheye image through positioning information projection, the fisheye image is compressed. The fisheye rendering area is the content requested by the current viewpoint. By not compressing the fisheye rendering area, or compressing the fisheye rendering area at a compression ratio lower than that of the non-rendered fisheye area, the compression purpose can be achieved, saving image transmission bandwidth. At the same time, the compression ratio of the compressed fisheye rendering area is small, or not compressed at all, so that the final rendered image is not excessively compressed, resulting in a decrease in image quality, thus ensuring the clarity of the rendered area image.
[0055] The above-mentioned fisheye video stream compression method determines the fisheye rendering area on the fisheye image based on the positioning information of the rendering area at the decoding end. During compression, the compression ratio of the fisheye rendering area in the compressed image is less than the compression ratio of the non-rendered fisheye area, and / or the fisheye rendering area is not compressed, thereby achieving the purpose of compression. At the same time, it saves the bandwidth of image transmission, and the compression ratio of the compressed fisheye rendering area is small, or lossless compression, so that the final rendered image is not excessively compressed, which would lead to a decrease in image quality, thus ensuring the clarity of the rendered area image.
[0056] In one embodiment, such as Figure 3 As shown, the fisheye rendering region on the corresponding fisheye image, determined based on the positioning information, includes:
[0057] Step S302: Collect points at equal intervals on the boundary of the rendering area to obtain the first two-dimensional point set.
[0058] Specifically, multiple points are collected at equal intervals along the boundary of the rendering area at the decoding end, for example, 100 points can be collected. Then, the collected points are stored in a specific order, clockwise or counterclockwise, forming a first two-dimensional point set P. Since the rendering area at the decoding end is usually square (rectangular), to ensure that each edge on the boundary is collected, at least four points are collected, that is, each of the four vertices of the square is a sampling point. The sampling interval is determined based on the size of the rendering area and the determined number of points. In this embodiment, it is preferable to collect at least 12 points, because the more sampling points, the more accurate the boundary of the rendering area is located on the fisheye image.
[0059] Step S304: Based on the positioning information of the rendering area, project the first two-dimensional point set onto the spherical coordinate system to obtain the three-dimensional point set.
[0060] Specifically, after obtaining the first two-dimensional point set P acquired at equal intervals, the first two-dimensional point set P is projected onto a spherical coordinate system based on the positioning information of the rendering area, resulting in the three-dimensional points projected onto the spherical coordinate system, forming a three-dimensional point set Ps. The projection method used can be any existing projection method, such as spherical perspective projection or spherical isometric projection. This embodiment uses spherical perspective projection to project the first two-dimensional point set P onto the spherical coordinate system to obtain the three-dimensional point set Ps.
[0061] Step S306: Project the three-dimensional point set onto the fisheye image corresponding to the positioning information to obtain the second two-dimensional point set.
[0062] Specifically, after the encoding end obtains the three-dimensional point set Ps, it projects the three-dimensional point set Ps onto the fisheye image corresponding to the positioning information. The point set formed by the two-dimensional points obtained from this projection is used as the second two-dimensional point set Pf. In this embodiment, a spherical equidistant projection method is preferably used to project the three-dimensional point set Ps onto each fisheye image to obtain the second two-dimensional point set Pf.
[0063] Furthermore, depending on the viewpoint of the rendering (decoding) end, the second two-dimensional point set Pf projected onto the fisheye image may be distributed across one fisheye image, or it may be distributed across two or more fisheye images. (See reference...) Figure 4-7 , Figure 4 The distribution diagram shown illustrates the positioning of the rendered area (white area) projected onto the fisheye image from a certain viewpoint direction when the horizontal field of view of the rendered area is 100 degrees; that is, it is distributed across two fisheye images. Figure 5 The distribution map shown represents the positioning of the rendered region on the fisheye image after projection at 60 degrees, i.e., it is distributed only on one fisheye image. Therefore, when the second two-dimensional point set Pf is distributed on two fisheye images, the second two-dimensional point set of the left fisheye image can be denoted as Pf0, and the second two-dimensional point set of the right fisheye image can be denoted as Pf1, where Pf = {Pf0, Pf1}. Figure 6 and Figure 7 They are respectively Figure 4 and Figure 5 The corresponding rendered panoramic view (left) and rendered area view (right).
[0064] Step S308: Determine the fisheye rendering area of the fisheye image based on the second two-dimensional point set.
[0065] Specifically, the encoding end determines the fisheye rendering region R based on the second two-dimensional point set Pf obtained from the projection.
[0066] In one embodiment, the step of determining the fisheye rendering region of a fisheye image based on a second two-dimensional point set includes: determining whether the second two-dimensional point set is a closed point set based on the Euclidean distance between the first and last points in the second two-dimensional point set; when the second two-dimensional point set is a closed point set, using the internal region defined by the closed boundary obtained by connecting the points in the second two-dimensional point set in sequence as the fisheye rendering region of the fisheye image; when the second two-dimensional point set is not a closed point set, constructing a closed second two-dimensional point set, and using the internal region defined by the closed boundary obtained by connecting the points in the constructed second two-dimensional point set in sequence as the fisheye rendering region of the fisheye image.
[0067] Specifically, due to the different viewpoint orientation at the decoding end, the second two-dimensional point set Pf may or may not be closed. In the case of an open point set, a specific method is needed to find additional sampling points on the fisheye image to make it closed. Therefore, before determining the fisheye rendering region R using the second two-dimensional point set Pf, it is first necessary to determine whether the second two-dimensional point set Pf is a closed point set. If it is determined to be a closed point set, the points in the second two-dimensional point set Pf are directly connected in sequence; the area occupied by the polygon formed by connecting all the points is the fisheye rendering region. If it is determined to be a non-closed point set, after constructing a closed point set, the points in the constructed closed point set are connected in sequence to obtain the fisheye rendering region.
[0068] In one embodiment, constructing a closed second two-dimensional point set includes: acquiring points at equal intervals on the field-of-view boundary of the fisheye lens in the fisheye image to obtain an additional point set; and merging the additional point set with the second two-dimensional point set to obtain a closed second two-dimensional point set.
[0069] Specifically, when constructing a closed second two-dimensional point set Pf, points are collected at equal intervals along the field of view (FOV) boundary of the fisheye lens in the fisheye image. For example, when the subset Pf0 of the second two-dimensional point set Pf distributed on the left fisheye image is not a closed point set, multiple points (e.g., 500 points) are collected at equal intervals along the FOV boundary of a certain field of view angle of the fisheye lens in the left fisheye image to form an additional point set Pfe. The selected field of view angle can be greater than 180 degrees and less than the maximum FOV (field of view boundary) of the fisheye lens. The "field of view boundary" refers to the area covered by a certain FOV in the fisheye image, which can be idealized as a circular area defined by a point C in the central region of the fisheye lens as the center and R as the radius. R is calculated from the FOV, and C is obtained through calibration. The boundary of this circular area is the "field of view boundary". Then, the extra point set Pfe is back-projected. Back-projection involves first projecting the extra point set Pfe onto a predetermined spherical coordinate system using spherical isometric projection to form a 3D point set. Then, spherical perspective is used to project this 3D point set onto the plane containing the rendering area. Points projected into the rendering area are added to the set to form point set Pr. That is, point set Pr is obtained by projecting a subset pfe0 of the extra point set Pfe. Therefore, the union of point sets Pfe0 and Pf0 can form a closed point set. Thus, by collecting points at equal intervals along the field of view boundary of the fisheye lens in a fisheye image, the extra point set Pfe is obtained. The subsets Pfe0 and Pf0 of the extra point set can be merged to form a closed point set.
[0070] In this embodiment, the fisheye rendering area corresponding to the decoding rendering area is determined by the mapping method of projection and back projection, which can obtain the compressed area corresponding to the decoding rendering area and ensure the clarity of the compressed rendering area.
[0071] In one embodiment, compressing a fisheye image to obtain a compressed image includes: determining the area of the compressed image based on a preset compression ratio and the resolution of the fisheye image; downsampling the fisheye image according to a first compression ratio to obtain a fisheye thumbnail; and storing the pixels in the fisheye rendering area and the fisheye thumbnail in the compressed image when the total number of pixels in the fisheye thumbnail and the fisheye rendering area is less than or equal to the total number of pixels in the compressed image.
[0072] The preset compression ratio is a pre-defined value determined by the transmission performance of the fisheye video stream. A suitable compression ratio is typically chosen to enable low-latency transmission of the video stream. The compressed image is then used to store the fisheye thumbnail and the fisheye rendering area.
[0073] Specifically, the size of the compressed image is calculated using a preset compression ratio and the resolution of the fisheye image. For example, if the preset compression ratio is K:1 and the resolution is Wf*Hf, the compressed image area S = Wf*Hf / K can be obtained. The corresponding compressed image is generated according to the determined compressed image area S. The fisheye image is downsampled according to a first compression ratio, such as 500:1, to obtain a fisheye thumbnail. In other words, the preset compression ratio is related to the first compression ratio, and the first compression ratio is usually greater than the preset compression ratio.
[0074] Before storage, the relationship between the total number of pixels in the fisheye thumbnail and the fisheye rendering area and the total number of pixels in the compressed image is first determined. Based on this relationship, it is decided whether to compress the fisheye rendering area. Specifically, when the total number of pixels in the fisheye thumbnail and the fisheye rendering area is less than or equal to the total number of pixels in the compressed image (i.e., the total area Sr of all fisheye rendering areas ≤ (area of compressed image S - area of fisheye thumbnail w*h), it means the compressed image can store the pixels of both the fisheye thumbnail and the fisheye rendering area simultaneously. In this case, the pixels of the fisheye thumbnail and the fisheye rendering area are directly stored row by row in the compressed image. At this point, the pixels in the fisheye rendering area are not compressed, achieving lossless compression of the fisheye rendering area. In practical applications, the positioning information of the rendering area can also be stored in the compressed image. This positioning information can also be stored in other ways, as long as it enables the transmission of positioning information between the encoding and decoding ends.
[0075] In another embodiment, when the total number of pixels in the fisheye thumbnail and the fisheye rendering area is greater than the total number of pixels in the compressed image, the fisheye rendering area is compressed using a second compression ratio, and the pixels in the compressed fisheye rendering area and the pixels in the fisheye thumbnail are stored in the compressed image; wherein, the second compression ratio is less than the first compression ratio.
[0076] Specifically, when the total number of pixels in the fisheye thumbnail and the fisheye rendering area is greater than the total number of pixels in the compressed image (i.e., Sr > Sw*h), it indicates that the compressed image cannot simultaneously store the pixels of the fisheye thumbnail and the fisheye rendering area. In this case, the fisheye rendering area is downsampled using a second compression ratio, and then the downsampled pixels of each fisheye rendering area and the fisheye thumbnail are stored row-by-row in the compressed image. The second compression ratio is less than the first compression ratio; for example, the second compression ratio K' is K' = Sr / (Sw*h). In this embodiment, before storing in the compressed image, the size relationship is judged to determine whether to downsample again before storage, preventing the amount of data that the generated compressed image can store from exceeding its capacity. The first and second compression ratios are related to a preset compression ratio; the first compression ratio is greater than the preset compression ratio, and the second compression ratio is less than the preset compression ratio.
[0077] In another embodiment, the method of storing pixels in the fisheye rendering region into a compressed image includes: extracting pixels from the fisheye rendering region sequentially in a preset direction, and storing the extracted pixels into the compressed image sequentially in the extraction order; the preset direction includes rows or columns.
[0078] In this embodiment, pixels in the fisheye rendering region corresponding to the fisheye image are stored in a densely packed manner into the compressed image, thereby completing the compression. The arrangement of the pixels stored in the compressed image can be arbitrary, as long as it follows the principle of easy storage and easy decoding. Specifically, pixels in the fisheye rendering region are extracted sequentially in a preset direction, such as by row or column, and the extracted pixels are stored in the compressed image in the extraction order. Using this method, pixel information can be densely stored in the compressed image, wherein the last pixel in the nth row / column of the fisheye rendering region in the compressed image is immediately followed by the first pixel in the (n+1)th row / column.
[0079] It is understandable that the method of storing the pixels of the fisheye thumbnail into the compressed image is the same as the method of storing the pixels of the fisheye rendering area into the compressed image, and will not be repeated here.
[0080] The difference between dense storage and normal storage is that dense storage destroys the original image but preserves the positional relationships of the images. Normal storage stores images in a block-like manner, while dense storage destroys the concept of blocks and eliminates the concept of rows and columns. For example, in the memory of a compressed image, the last pixel of the nth row is immediately followed by the first pixel of the (n+1)th row. This method reduces the image size, making it easier to store and restore.
[0081] In another embodiment, the fisheye image is compressed to obtain a compressed image, including:
[0082] Generate at least one downsampling mapping table, each recording the mapping relationship between the non-rendered fisheye region and the compressed image, as well as the mapping relationship between the rendered fisheye region and the compressed image; perform image remapping on the non-rendered fisheye region and the rendered fisheye region according to each downsampling mapping table to obtain the compressed image.
[0083] Specifically, a mapping table with downsampling functionality is generated for the corresponding region, namely the downsampling mapping table in this embodiment. Each downsampling mapping table records the mapping relationship between the non-rendered fisheye region and the compressed image, as well as the mapping relationship between the rendered fisheye region and the compressed image. Since it is a mapping table with downsampling functionality, it can be understood that downsampling of the region has been completed during the generation of the mapping table. This downsampling mapping table includes the original positioning information of each pixel in its corresponding region in the fisheye image. Then, based on the original positioning information in the generated mapping table, image remapping is performed on the non-rendered fisheye region and the rendered fisheye region to complete compression and obtain a compressed image. The image stored in the compressed image is the mapping result of the image remapping. Alternatively, multiple mapping tables can be generated simultaneously. The difference between each mapping table lies in the region it corresponds to. That is, the mapping results obtained by performing image remapping on multiple mapping tables are each a part of the compressed image, and all the mapping results are stored in the same compressed image to obtain a complete compressed image.
[0084] The mapping table can be generated using any method, primarily aiming to ensure multi-resolution downsampling and maintain the local continuity of the fisheye image. Multi-resolution downsampling refers to applying a lower compression ratio to the defined fisheye rendering area or not applying downsampling at all (i.e., when the compressed image can fit the complete rendering area, downsampling can be omitted), while applying a higher compression ratio to the non-rendered areas of the fisheye. Maintaining the local continuity of the fisheye image means that the relative positional relationship between any two pixels in a certain area of the fisheye remains unchanged in the compressed image. This objective is beneficial for video stream encoding compression when converting the compressed image to a video stream.
[0085] In this embodiment, compression is performed using a mapping table. Whether it is the fisheye rendering area or the non-rendering area, the downsampling step can be included in the process of generating the mapping table. The entire compressed image can be obtained in one step or in several steps directly through one or more mapping tables.
[0086] In one embodiment, a fisheye video stream compression method is also provided, which is applied to, for example... Figure 1 The encoding end shown includes the following methods: acquiring a fisheye video stream; using the fisheye image compression method of each embodiment to compress each frame of the fisheye video stream to obtain a compressed image of each frame of the fisheye image; and performing video stream compression based on the compressed images of each frame of the fisheye video stream to obtain a compressed fisheye video stream.
[0087] Here, fisheye video stream refers to video images captured using a fisheye lens. It can be understood that each frame of the fisheye video stream is the fisheye image mentioned in the previous embodiments.
[0088] The fisheye image compression methods for each embodiment have been described in the preceding embodiments and will not be elaborated here. It is understood that, for each frame of the compressed fisheye image, the fisheye rendering area of the fisheye image is determined based on the positioning information of the rendering area, and the fisheye rendering area is compressed at a small compression ratio or with lossless compression, thereby achieving the purpose of compression. At the same time, the fisheye rendering image is not excessively compressed, which would lead to a decrease in image quality, thus ensuring the clarity of the image in the fisheye rendering area.
[0089] In other words, in a dynamic fisheye video stream, the actual fisheye rendering area required can be determined in real time according to the change of viewpoint, so that the fisheye video stream can be dynamically compressed in real time according to the change of viewpoint at the decoding end.
[0090] For each frame of the fisheye video stream, a compressed image is then further compressed using a video stream compression method, such as H.264. This video stream compression process yields the compressed fisheye video stream.
[0091] The aforementioned fisheye video stream compression can dynamically adjust the rendering area of the fisheye based on the position of the real-time preview area of the rendering end in the panorama. It can compress the fisheye video stream in real time according to the changes in the viewpoint of the decoding end, which not only ensures the high definition of the rendering area, but also achieves a high compression ratio, greatly saving the transmission bandwidth of the video stream.
[0092] In one embodiment, a panoramic video generation method is also provided, which is applied to, for example... Figure 1 The decoding end shown includes: acquiring a compressed fisheye video stream; wherein the compressed fisheye video stream is obtained by processing the fisheye video stream compression method described above; obtaining compressed images corresponding to multiple frames of original fisheye images based on the compressed fisheye video stream; restoring the compressed images to obtain the original fisheye images; and stitching together the original fisheye images to obtain a panoramic video.
[0093] Specifically, the decoding end obtains the fisheye video stream processed by the encoding end. The method for obtaining the fisheye video stream by the encoding end has been described in the previous specification and will not be elaborated here.
[0094] Understandably, the decoder may obtain multiple compressed fisheye video streams, depending on the number of fisheye lenses set up at the acquisition site. If two fisheye lenses are set up at the acquisition site, the decoder will obtain two fisheye video streams.
[0095] For the received compressed fisheye video stream, a decompression method corresponding to the video stream compression method is used to decompress the fisheye video stream to obtain a compressed image corresponding to multiple frames of the original fisheye image.
[0096] Specifically, the compressed image is restored using a restoration method corresponding to the compression method to obtain the original fisheye image.
[0097] Specifically, for the original fisheye image, matching points are found, and the images are stitched together based on the matching points to obtain a panoramic video.
[0098] The aforementioned panoramic video generation method performs video stream decompression, compressed image restoration, and stitching on the compressed fisheye video stream to obtain the panoramic video. On one hand, since the panoramic video is obtained by processing the compressed fisheye video stream, the rendering area of the fisheye can be dynamically adjusted according to the real-time preview area's position within the panorama. This allows for real-time compression of the fisheye video stream based on changes in the decoder's viewpoint, ensuring both high definition of the rendered area and a high compression ratio, significantly saving video stream transmission bandwidth. On the other hand, the encoding end only compresses the fisheye image and video stream, without stitching; instead, the decoding end decompresses and then stitches the data, reducing the high hardware performance requirements of the encoding end.
[0099] After compression is completed at the encoding end, the compressed image can be encapsulated into a data stream and transmitted to the decoding end for decoding and display. Display recovery requires the location information of the rendering area. As mentioned earlier, the location information of the rendering area can be stored in the compressed image and transmitted to the decoding end. Other storage methods can also be used, as long as the transmission of location information between the encoding and decoding ends is possible. After obtaining the location information of the rendering area, the decoding end decodes the compressed image according to the original location information to recover the original fisheye image.
[0100] In the case of a scheme where the positioning information of the rendering area is stored in a compressed image, restoring the compressed image to obtain the original fisheye image includes: parsing the compressed image to obtain the original positioning information of each pixel in the fisheye rendering area in the original fisheye image; and decoding the compressed image based on the original positioning information to recover the original fisheye image.
[0101] Specifically, when the compressed image at the encoding end is obtained through dense storage, the original positioning information of each pixel in the compressed image on the original fisheye image needs to be sent to the decoding end along with the compressed image. Subsequently, the decoding end copies the number of pixels indicated by the original positioning information in the compressed image to the corresponding positions in the fisheye image to complete the decoding and restoration of the rendered area in the original fisheye image. The non-rendered area can be obtained by upsampling the fisheye thumbnail stored in the compressed image. The rendered and non-rendered areas together constitute the restored original fisheye image. If the pixels are stored row-wise, the original positioning information of each pixel can also be recorded row-wise. That is, the original positioning information of each row of pixels can be represented by three data points: the row index and column index of the first pixel in the row in the original fisheye image, and the total number of pixels in that row.
[0102] In another embodiment, restoring the compressed image to obtain the original fisheye image includes: obtaining an inverse mapping table generated based on a downsampling mapping table; and remapping the compressed image based on the inverse mapping table to restore the original fisheye image.
[0103] Specifically, when the compressed image is obtained through image remapping using a mapping table, the inverse mapping table or the parameter information for constructing the inverse mapping table must also be transmitted to the decoding end along with the compressed image. The decoding end then uses the inverse mapping table to remap the pixels in the compressed image to decode and recover the original fisheye image. Finally, the decoding end performs panoramic stitching and rendering on the decompressed original fisheye image.
[0104] In this embodiment, the corresponding decoding information is transmitted to the decoding end according to different compression methods to ensure that the decoding end can efficiently complete the decoding and recovery of the fisheye image.
[0105] It should be understood that, although Figure 2-3 The steps in the flowchart are shown sequentially as indicated by the arrows, but these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise specified in this document, there is no strict order in which these steps are executed, and they can be performed in other orders. Furthermore, Figure 2-3 At least some of the steps in the process may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but may be executed at different times. The execution order of these steps or stages is not necessarily sequential, but may be executed in turn or alternately with other steps or at least some of the steps or stages in other steps.
[0106] In one embodiment, such as Figure 8As shown, a fisheye video stream compression device is provided, including: an information transmission module 802, a compression region determination module 804, and a compression module 806, wherein:
[0107] Information transmission module 802 is used to obtain the positioning information of the rendering area of the decoding end;
[0108] The compression region determination module 804 is used to determine the fisheye rendering region on the corresponding fisheye image based on the positioning information, wherein the area on the fisheye image other than the fisheye rendering region is the fisheye non-rendering region.
[0109] Compression module 806 is used to compress the fisheye image to obtain a compressed image; wherein the compression ratio of the fisheye rendering area in the compressed image is less than the compression ratio of the fisheye non-rendering area, and / or the fisheye rendering area is not compressed.
[0110] In one embodiment, the compression region determination module 804 is further configured to collect points at equal intervals on the boundary of the rendering region to obtain a first two-dimensional point set; project the first two-dimensional point set onto a spherical coordinate system according to the positioning information of the rendering region to obtain a three-dimensional point set; project the three-dimensional point set onto the fisheye image corresponding to the positioning information to obtain a second two-dimensional point set; and determine the fisheye rendering region of the fisheye image according to the second two-dimensional point set.
[0111] In one embodiment, the compression region determination module 804 is further configured to determine whether the second two-dimensional point set is a closed point set based on the Euclidean distance between the first and last points in the second two-dimensional point set; when the second two-dimensional point set is a closed point set, the internal region defined by the closed boundary obtained by connecting the points in the second two-dimensional point set in sequence is used as the fisheye rendering region of the fisheye image; when the second two-dimensional point set is not a closed point set, a closed second two-dimensional point set is constructed, and the internal region defined by the closed boundary obtained by connecting the points in the constructed second two-dimensional point set in sequence is used as the fisheye rendering region of the fisheye image.
[0112] In one embodiment, the compression region determination module 804 is further configured to collect points at equal intervals on the field of view boundary of the fisheye lens in the fisheye image to obtain an additional point set; and to merge the additional point set with the second two-dimensional point set to obtain a closed second two-dimensional point set.
[0113] In one embodiment, the compression module 806 is further configured to determine the area of the compressed image based on a preset compression ratio and the resolution of the fisheye image; downsample the fisheye image according to the first compression ratio to obtain a fisheye thumbnail; and store the pixels in the fisheye rendering area and the fisheye thumbnail in the compressed image when the total number of pixels in the fisheye thumbnail and the fisheye rendering area is less than or equal to the total number of pixels in the compressed image.
[0114] In one embodiment, the compression module 806 is further configured to compress the fisheye rendering region using a second compression ratio when the total number of pixels in the fisheye thumbnail and the fisheye rendering region is greater than the total number of pixels in the compressed image, and store the pixels in the compressed fisheye rendering region and the pixels in the fisheye thumbnail in the compressed image; wherein the second compression ratio is less than the first compression ratio; and the preset compression ratio is related to the first compression ratio and the second compression ratio.
[0115] In another embodiment, the compression module is further configured to extract pixels from the fisheye rendering area sequentially in a preset direction, and store the extracted pixels in the compressed image in the extraction order, wherein the preset direction includes rows or columns.
[0116] In the compressed image, the last pixel of the nth row / column of the fisheye rendering area is immediately followed by the first pixel of the (n+1)th row / column.
[0117] In one embodiment, the compression module 806 is further configured to generate at least one downsampling mapping table, each downsampling mapping table recording the mapping relationship between the non-rendered fisheye region and the compressed image, and the mapping relationship between the rendered fisheye region and the compressed image; and to perform image remapping on the non-rendered fisheye region and the rendered fisheye region according to each downsampling mapping table to obtain a compressed image.
[0118] Specific limitations regarding the fisheye video stream compression device can be found in the limitations of the fisheye video stream compression method described above, and will not be repeated here. Each module in the aforementioned fisheye video stream compression device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in hardware or independently of the processor in the computer device, or stored in software in the memory of the computer device, so that the processor can call and execute the corresponding operations of each module.
[0119] In another embodiment, a fisheye video stream compression apparatus is also provided, comprising:
[0120] The video stream acquisition module is used to acquire fisheye video streams;
[0121] The image compression module is used to compress each frame of the fisheye image in the fisheye video stream using the fisheye image compression method described in the above embodiments, so as to obtain a compressed image of each frame of the fisheye image.
[0122] The video stream compression module is used to compress the video stream based on the compressed images of each frame of the fisheye video stream to obtain a compressed fisheye video stream.
[0123] In another embodiment, a panoramic video generation apparatus includes:
[0124] A video stream acquisition module is used to acquire a compressed fisheye video stream; wherein the compressed fisheye video stream is obtained by processing the fisheye video stream compression method as described in the above embodiments;
[0125] The video stream decompression module is used to obtain compressed images corresponding to multiple frames of the original fisheye images based on the compressed fisheye video stream.
[0126] The restoration module is used to restore the compressed image to obtain the original fisheye image;
[0127] The stitching module is used to stitch together the original fisheye images to obtain a panoramic video.
[0128] The restoration module is used to parse the compressed image to obtain the original positioning information of each pixel in the fisheye rendering area in the original fisheye image; and to decode the compressed image according to the original positioning information to restore the original fisheye image.
[0129] The restoration module is used to obtain an inverse mapping table generated based on the downsampling mapping table; and to remap the compressed image according to the inverse mapping table to restore the original fisheye image.
[0130] In one embodiment, a computer device is provided, which may be a terminal, and its internal structure diagram may be as follows: Figure 9As shown, the computer device includes a processor, memory, communication interface, display screen, and input devices connected via a system bus. The processor provides computing and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The communication interface is used for wired or wireless communication with external terminals; wireless communication can be achieved through Wi-Fi, carrier networks, NFC (Near Field Communication), or other technologies. When executed by the processor, the computer program implements a fisheye video stream compression method, a fisheye video stream compression method, or a panoramic video generation method. The display screen can be an LCD screen or an e-ink screen. The input devices can be a touch layer covering the display screen, buttons, a trackball, or a touchpad mounted on the computer device casing, or an external keyboard, touchpad, or mouse.
[0131] Those skilled in the art will understand that Figure 9 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.
[0132] In one embodiment, a computer device is provided, including a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the fisheye image compression method, fisheye video stream compression method, or panoramic video generation method of the above embodiments.
[0133] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored, which, when executed by a processor, implements the fisheye image compression method, fisheye video stream compression method, or panoramic video generation method of the above embodiments.
[0134] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the methods described above. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, or optical storage, etc. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc.
[0135] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0136] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the invention patent. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this patent application should be determined by the appended claims.
Claims
1. A method for compressing fisheye images, characterized in that, The method includes: Obtain the positioning information of the rendering area at the decoding end; determine the fisheye rendering area on the corresponding fisheye image based on the positioning information, wherein the area on the fisheye image other than the fisheye rendering area is the non-rendering area of the fisheye; the fisheye rendering area is located on at least two fisheye images. The fisheye image is compressed to obtain a compressed image; wherein the compression ratio of the fisheye rendering area in the compressed image is less than the compression ratio of the fisheye non-rendering area, and / or the fisheye rendering area is not compressed; Determining the fisheye rendering region on the corresponding fisheye image based on the positioning information includes: At least points are collected at the boundaries of the rendering area to obtain a first two-dimensional point set; Based on the positioning information of the rendering area, the first two-dimensional point set is projected onto the spherical coordinate system to obtain a three-dimensional point set; The three-dimensional point set is projected onto the fisheye image corresponding to the positioning information to obtain a second two-dimensional point set; the second two-dimensional point set is located on at least two fisheye images. The fisheye rendering region of the fisheye image is determined based on the second two-dimensional point set.
2. The method according to claim 1, characterized in that, Determining the fisheye rendering region on the corresponding fisheye image based on the positioning information includes: Points are collected at equal intervals on the boundary of the rendering area to obtain a first two-dimensional point set; Based on the positioning information of the rendering area, the first two-dimensional point set is projected onto the spherical coordinate system to obtain a three-dimensional point set; The three-dimensional point set is projected onto the fisheye image corresponding to the positioning information to obtain the second two-dimensional point set; The fisheye rendering region of the fisheye image is determined based on the second two-dimensional point set.
3. The method according to claim 2, characterized in that, Determining the fisheye rendering region of the fisheye image based on the second two-dimensional point set includes: Determine whether the second two-dimensional point set is a closed point set based on the Euclidean distance between the first and last points in the second two-dimensional point set; When the second two-dimensional point set is a closed point set, the internal region defined by the closed boundary obtained by connecting the points in the second two-dimensional point set in sequence is used as the fisheye rendering region of the fisheye image. When the second two-dimensional point set is not a closed point set, a closed second two-dimensional point set is constructed, and the internal region defined by the closed boundary obtained by connecting the points in the constructed second two-dimensional point set in sequence is used as the fisheye rendering region of the fisheye image.
4. The method according to claim 3, characterized in that, The construction of the closed second two-dimensional point set includes: Additional point sets are obtained by collecting points at equal intervals on the field-of-view boundary of the fisheye lens in the fisheye image; The additional point set is merged with the second two-dimensional point set to obtain a closed second two-dimensional point set.
5. The method according to claim 1, characterized in that, The process of compressing the fisheye image to obtain a compressed image includes: The area of the compressed image is determined based on the preset compression ratio and the resolution of the fisheye image; The fisheye image is downsampled according to the first compression ratio to obtain a fisheye thumbnail; When the total number of pixels in the fisheye thumbnail and the fisheye rendering area is less than or equal to the total number of pixels in the compressed image, the pixels in the fisheye rendering area and the pixels in the fisheye thumbnail are stored in the compressed image.
6. The method according to claim 5, characterized in that, The method further includes: When the total number of pixels in the fisheye thumbnail and the fisheye rendering region is greater than the total number of pixels in the compressed image, the fisheye rendering region is compressed using a second compression ratio, and the pixels in the compressed fisheye rendering region and the pixels in the fisheye thumbnail are stored in the compressed image; wherein, the second compression ratio is less than the first compression ratio; the preset compression ratio is related to the first compression ratio and the second compression ratio.
7. The method according to claim 5 or 6, characterized in that, Methods for storing pixels in the fisheye rendering region into a compressed image include: The fisheye rendering area is used to extract pixels sequentially in a preset direction, and the extracted pixels are stored in the compressed image in the extraction order; the preset direction includes rows or columns.
8. The method according to claim 7, characterized in that, In the compressed image, the last pixel of the nth row / column of the fisheye rendering area is immediately followed by the first pixel of the (n+1)th row / column.
9. The method according to claim 1, characterized in that, The process of compressing the fisheye image to obtain a compressed image includes: At least one downsampling mapping table is generated, and each downsampling mapping table records the mapping relationship between the non-rendered fisheye region and the compressed image, as well as the mapping relationship between the rendered fisheye region and the compressed image. The image is remapped according to the downsampling mapping tables for the non-rendered fisheye region and the rendered fisheye region to obtain a compressed image.
10. A method for compressing fisheye video streams, characterized in that, include: Acquire fisheye video stream; The fisheye image compression method as described in any one of claims 1-9 is used to compress each frame of the fisheye video stream to obtain a compressed image of each frame of the fisheye image. The video stream is compressed based on the compressed images of each frame of the fisheye video stream to obtain a compressed fisheye video stream.
11. A method for generating panoramic video, characterized in that, include: Obtain a compressed fisheye video stream; wherein the compressed fisheye video stream is obtained by processing the fisheye video stream compression method as described in claim 10; Based on the compressed fisheye video stream, a compressed image corresponding to multiple frames of the original fisheye image is obtained; The compressed image is restored to obtain the original fisheye image; By stitching together the original fisheye images, a panoramic video is obtained.
12. The method according to claim 11, characterized in that, The process of restoring the compressed image to obtain the original fisheye image includes: The compressed image is analyzed to obtain the original positioning information of each pixel in the fisheye rendering area in the original fisheye image; The compressed image is decoded based on the original positioning information to recover the original fisheye image.
13. The method according to claim 11, characterized in that, The process of restoring the compressed image to obtain the original fisheye image includes: Obtain the inverse mapping table generated from the downsampling mapping table; The original fisheye image is recovered by remapping the compressed image using the inverse mapping table.
14. A fisheye video stream compression device, characterized in that, The device includes: The information transmission module is used to obtain the positioning information of the rendering area at the decoding end; A compression region determination module is used to determine the fisheye rendering region on the corresponding fisheye image based on the positioning information, wherein the area on the fisheye image other than the fisheye rendering region is the non-rendered fisheye region; the fisheye rendering region is located on at least two fisheye images; the step of determining the fisheye rendering region on the corresponding fisheye image based on the positioning information includes: collecting points at least on the boundary of the rendering region to obtain a first two-dimensional point set; projecting the first two-dimensional point set onto a spherical coordinate system based on the positioning information of the rendering region to obtain a three-dimensional point set; projecting the three-dimensional point set onto the fisheye image corresponding to the positioning information to obtain a second two-dimensional point set; the second two-dimensional point set is located on at least two fisheye images; and determining the fisheye rendering region of the fisheye image based on the second two-dimensional point set. A compression module is used to compress the fisheye image to obtain a compressed image; wherein the compression ratio of the fisheye rendering area in the compressed image is less than the compression ratio of the fisheye non-rendering area, and / or the fisheye rendering area is not compressed.
15. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 13.
16. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 13.