Multi-view video compression method, electronic equipment, storage medium and product

By extracting the initial key view and sequence view in the multi-view video data and determining the residual information based on the reconstruction key view, the problem of low compression efficiency of multi-view video data in the prior art is solved, and more efficient video data compression is achieved.

CN120017849APending Publication Date: 2025-05-16PEKING UNIV SHENZHEN GRADUATE SCHOOL
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510094928.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-21
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

The existing compression method of multi-view video data is low when processing multi-view video data involving six degrees of freedom.

Method used

By obtaining the multi-viewpoint video data to be compressed, the initial key view and the initial sequence view in the target picture group are extracted, the initial key view is processed using the main encoder to obtain the reconstruction key view, and the residual information of each initial sequence view is determined based on the reconstruction key view.

Benefits of technology

The compression efficiency of multi-view video data is improved, and the amount of data that needs to be stored and transmitted is reduced by extracting and processing key views and sequence views in the data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120017849A_ABST
    Figure CN120017849A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-view video compression method, electronic equipment, a storage medium and a product, and relates to the technical field of multi-view videos, and the multi-view video compression method comprises the following steps: obtaining to-be-compressed multi-view video data, and extracting a target picture group in the to-be-compressed multi-view video data; an initial key view and each initial sequence view included in the target picture group are extracted, and the initial key view is a first frame view corresponding to a first viewpoint in the target picture group; and processing the initial key view through a main encoder to obtain a reconstructed key view, and determining residual information correspondingly matched with each initial sequence view based on the reconstructed key view. According to the invention, the technical effect of improving the video compression efficiency is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of multi-viewpoint video technology, and in particular to a multi-viewpoint video compression method, electronic equipment, storage medium and computer program product. Background Art

[0002] As multi-view video is increasingly used in users' daily lives, it has greatly improved users' visual experience. With the continuous development of multi-view video technology, the amount of multi-view video data is also growing rapidly.

[0003] In related technologies, technicians usually use a three-dimensional data compression method based on deep learning to compress multi-viewpoint video data to eliminate spatial redundancy between different viewpoints, thereby reducing the data volume of the compressed multi-viewpoint video data.

[0004] However, the above-mentioned multi-viewpoint video data compression method can usually only be used on binocular video data or light field video data with small parallax between viewpoints. Once the video data is multi-viewpoint video data involving six degrees of freedom, the problem of low video compression efficiency is likely to occur. Summary of the invention

[0005] The main purpose of this application is to provide a multi-viewpoint video compression method, electronic device, storage medium and computer program product, aiming to solve the technical problem of low video compression efficiency in related technologies.

[0006] To achieve the above objectives, the present application proposes a multi-view video compression method, comprising:

[0007] Acquire multi-view video data to be compressed, and extract a target picture group from the multi-view video data to be compressed;

[0008] Extracting an initial key view and initial sequence views contained in the target picture group, wherein the initial key view is a first frame view corresponding to a first viewpoint in the target picture group;

[0009] The initial key view is processed by a main encoder to obtain a reconstructed key view, and residual information corresponding to each of the initial sequence views is determined based on the reconstructed key view.

[0010] In one embodiment, the initial sequence view includes an initial time domain sequence view, an initial space domain sequence view, and an initial time and space domain sequence view;

[0011] The step of determining the residual information corresponding to each of the initial sequence views based on the reconstructed key views comprises:

[0012] Determine first residual information of the initial time-domain sequence view matching according to the reconstructed key view, wherein the time-domain sequence view is other frame views corresponding to the first viewpoint except the first frame view;

[0013] Determine a reference spatial sequence view that matches the initial spatial sequence view according to the reconstructed key view, and determine second residual information according to the reference spatial sequence view, wherein the spatial sequence view is a view corresponding to other viewpoints at the same time point as the initial key view;

[0014] A target encoder matching the initial spatiotemporal domain sequence view is determined, and the initial spatiotemporal domain sequence view is processed by the target encoder to determine third residual information.

[0015] In one embodiment, the step of determining the first residual information of the initial time domain sequence view matching according to the reconstructed key view includes:

[0016] Performing optical flow prediction processing on the reconstructed key view and the initial time-domain sequence view through a time-domain encoder to determine a first prediction frame that matches the initial time-domain sequence view;

[0017] The initial time-domain sequence view and the first predicted frame are compared to determine first residual information matching the initial time-domain sequence view.

[0018] In one embodiment, the step of determining the second residual information according to the reference spatial sequence view further includes:

[0019] Determining a reference spatial domain depth view corresponding to the reference spatial domain sequence view;

[0020] Performing fusion prediction processing on the reference spatial sequence view and the reference spatial depth view through a spatial encoder to determine a second prediction frame that matches the initial spatial sequence view;

[0021] The second predicted frame is subjected to optical flow prediction compensation processing to obtain a third predicted frame, and the initial spatial sequence view is compared with the third predicted frame to determine second residual information matching the initial spatial sequence view.

[0022] In one embodiment, after the step of performing optical flow prediction compensation processing on the second predicted frame to obtain a third predicted frame, the method further includes:

[0023] Inputting the third predicted frame into a filter network in the spatial domain encoder, so as to obtain a fourth predicted frame by smoothing the third predicted frame through the filter network;

[0024] Determine second residual information of the initial spatial sequence view matching according to the fourth predicted frame.

[0025] In one embodiment, the step of determining a target encoder for view matching of the initial spatiotemporal domain sequence further includes:

[0026] Performing optical flow prediction processing through the time domain encoder and the reconstructed key view to determine a fourth prediction frame matched with the initial spatiotemporal domain sequence view;

[0027] Performing hybrid prediction processing on the initial spatiotemporal sequence view through a spatial domain encoder to determine a fifth prediction frame matched by the initial spatiotemporal sequence view;

[0028] Determine a minimum distortion prediction frame according to the fourth prediction frame and the fifth prediction frame;

[0029] In a case where the minimum distortion prediction frame is determined to be the fourth prediction frame, determining the time domain encoder as a target encoder for view matching of the initial spatiotemporal domain sequence;

[0030] In a case where the minimum distortion prediction frame is determined to be the fifth prediction frame, the spatial domain encoder is determined to be the target encoder.

[0031] In one embodiment, after the step of determining the residual information corresponding to each of the initial sequence views based on the reconstructed key views, the method further includes:

[0032] Receiving each of the residual information;

[0033] Restoring the reconstructed sequence views that match the respective initial sequence views according to the respective residual information;

[0034] Based on the spatiotemporal structure corresponding to the multi-view video data to be compressed, the reconstructed sequence views are integrated to restore the multi-view video data to be compressed.

[0035] In addition, to achieve the above-mentioned objectives, the present application also proposes an electronic device, which includes: a main encoder, a spatial domain encoder, a time domain encoder, a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein the computer program is configured to implement the steps of the multi-viewpoint video compression method as described above.

[0036] In addition, to achieve the above-mentioned purpose, the present application also proposes a storage medium, which is a computer-readable storage medium, and a computer program is stored on the storage medium. When the computer program is executed by a processor, the steps of the multi-viewpoint video compression method described above are implemented.

[0037] In addition, to achieve the above-mentioned purpose, the present application also provides a computer program product, which includes a computer program, and when the computer program is executed by a processor, the steps of the multi-view video compression method described above are implemented.

[0038] The multi-view video compression method provided in the embodiment of the present application obtains multi-view video data to be compressed, and extracts a target picture group in the multi-view video data to be compressed; extracts an initial key view and initial sequence views contained in the target picture group, wherein the initial key view is a first frame view corresponding to a first viewpoint in the target picture group; processes the initial key view through a main encoder to obtain a reconstructed key view, and determines the corresponding matching residual information of each of the initial sequence views based on the reconstructed key view.

[0039] In this way, the present application solves the technical problem of low video compression efficiency in the related art, that is, the present application splits the multi-view video data into multiple target picture groups, thereby extracting the initial key view contained in the target picture group and other initial sequence views except the initial key view, and then processes the initial key view to obtain a reconstructed key view as a reference benchmark, and extracts the residual information corresponding to each initial sequence view based on the reconstructed key view, so that the electronic device can extract the residual information with a smaller data amount from the multi-view video data, so that the electronic device only needs to store and transmit the residual information with a smaller data amount to complete the storage and transmission operations of the multi-view video data, thereby achieving the technical effect of improving the video compression efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0040] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present application and, together with the description, serve to explain the principles of the present application.

[0041] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0042] Figure 1 A flowchart of a first embodiment of a multi-view video compression method of the present application is provided;

[0043] Figure 2 This is a flowchart of an encoder process involved in an embodiment of a multi-view video compression method of the present application;

[0044] Figure 3 This is a hybrid prediction processing flow chart involved in an embodiment of a multi-view video compression method of the present application;

[0045] Figure 4 A schematic diagram of a filter model structure involved in an embodiment of a multi-view video compression method of the present application;

[0046] Figure 5 Schematic diagram of the device structure of the hardware operating environment involved in the multi-view video compression method in the embodiment of the present application.

[0047] The purpose, features and advantages of this application will be further described in conjunction with the embodiments and with reference to the accompanying drawings. DETAILED DESCRIPTION

[0048] It should be understood that the specific embodiments described herein are only used to explain the technical solutions of the present application and are not used to limit the present application.

[0049] In order to better understand the technical solution of the present application, a detailed description will be given below in conjunction with the accompanying drawings and specific implementation methods.

[0050] In this embodiment, for the convenience of description, the following is described with an electronic device having a video compression model configured therein, or a mobile terminal, a data storage control terminal, a PC or other terminal connected to an electronic control unit of the electronic device as the execution subject. The video compression model is provided with a main encoder, a time domain encoder and a space domain encoder.

[0051] Based on the above electronic device, the overall concept of the multi-view video compression method of the present application is proposed here.

[0052] As multi-view video is increasingly used in users' daily lives, it has greatly improved the user's visual experience. With the continuous development of multi-view video technology, the amount of multi-view video data is also growing rapidly. In related technologies, technicians usually use a three-dimensional data compression method based on deep learning to compress multi-view video data to eliminate spatial redundancy between different viewpoints, thereby reducing the amount of compressed multi-view video data. However, the above-mentioned multi-view video data compression method can usually only be used on binocular video data or light field video data with small parallax between viewpoints. Once the video data is multi-view video data involving six degrees of freedom, it is easy to have low video compression efficiency.

[0053] In view of the above phenomenon, the present application provides a multi-view video compression method, comprising: obtaining multi-view video data to be compressed, and extracting a target picture group in the multi-view video data to be compressed; extracting an initial key view and initial sequence views contained in the target picture group, wherein the initial key view is a first frame view corresponding to a first viewpoint in the target picture group; processing the initial key view by a main encoder to obtain a reconstructed key view, and determining corresponding matching residual information of each of the initial sequence views based on the reconstructed key view.

[0054] In this way, the present application solves the technical problem of low video compression efficiency in the related art, that is, the present application splits the multi-view video data into multiple target picture groups, thereby extracting the initial key view contained in the target picture group and other initial sequence views except the initial key view, and then processes the initial key view to obtain a reconstructed key view as a reference benchmark, and extracts the residual information corresponding to each initial sequence view based on the reconstructed key view, so that the electronic device can extract the residual information with a smaller data amount from the multi-view video data, so that the electronic device only needs to store and transmit the residual information with a smaller data amount to complete the storage and transmission operations of the multi-view video data, thereby achieving the technical effect of improving the video compression efficiency.

[0055] Based on the overall concept of the multi-view video compression method of the present application, the embodiment of the present application provides a multi-view video compression method, referring to Figure 1 , Figure 1 This is a flowchart of a first embodiment of a multi-view video compression method of the present application. In this embodiment, the multi-view video compression method includes steps S10 to S30:

[0056] Step S10: obtaining multi-view video data to be compressed, and extracting a target picture group from the multi-view video data to be compressed;

[0057] Step S20: extracting an initial key view and initial sequence views contained in the target picture group, wherein the initial key view is a first frame view corresponding to a first viewpoint in the target picture group;

[0058] Step S30: Processing the initial key view by a main encoder to obtain a reconstructed key view, and determining the residual information corresponding to each of the initial sequence views based on the reconstructed key view.

[0059] In this embodiment, when the electronic device needs to compress multi-viewpoint video data with six degrees of freedom, the electronic device first receives the multi-viewpoint video data to be compressed, and extracts a target picture group consisting of multiple sequence views corresponding to multiple different viewpoints in the multi-viewpoint video data to be compressed. After that, the electronic device splits the target picture group to determine the first frame sequence view corresponding to the first viewpoint in the target picture group as an initial key view used as a reference. At the same time, the electronic device extracts each initial time domain sequence view, each initial spatial domain sequence view, and each initial time and space domain sequence view in the target picture group based on the initial key view. Finally, the electronic device separates the initial key view, each initial time domain sequence view, each initial spatial domain sequence view, and each initial time and space domain sequence view. The figure is input into the video compression model configured by itself, and the video compression model obtains the reconstructed key view of the initial key view through the main encoder configured by itself. At the same time, the video compression model processes each initial time domain sequence view based on the reconstructed key view through the time domain encoder configured by itself, and obtains the first residual information that matches each initial time domain sequence view. At the same time, the video compression model processes each initial spatial domain sequence view based on the reconstructed key view through the spatial domain encoder configured by itself, and obtains the second residual information that matches each initial spatial domain sequence view. At the same time, the video compression model processes each initial time and space domain sequence view through the time domain encoder or the spatial domain encoder to obtain the third residual information that matches each initial time and space domain sequence view.

[0060] For example, see Figure 2 , Figure 2 This is a flowchart of an encoder process involved in an embodiment of a multi-view video compression method of the present application, such as Figure 2 As shown, when the electronic device is running, if it is necessary to compress the multi-view video data to be compressed with six degrees of freedom, the electronic device first receives the multi-view video data to be compressed, and then extracts the target picture set GOP composed of each sequence perspective contained in the first N frames of video data in the multi-view video data to be compressed. After that, the electronic device splits the target picture set GOP to extract the first frame view under the first viewpoint contained in the target picture set GOP to determine it as the initial key view k0, and the electronic device then screens in the target picture set GOP based on the initial key view k0 to use other views except the initial key view k0 under the first viewpoint as each initial time domain sequence view f that is temporally connected. c , and the sequence views with the same timestamp as the initial key view k0 but different viewpoints are used as the initial spatial sequence views v with spatial connections t , and the initial key view k0, each initial time domain sequence view f contained in the marked picture set GOP c , each initial spatial sequence view vt The views other than the initial spatiotemporal domain sequence views are determined. Finally, the electronic device inputs the initial key view k0 into the video compression model configured by itself, and the video compression model encodes and reconstructs the initial key view k0 through the main encoder to obtain the reconstructed key view At the same time, the video compression model converts each initial time domain sequence view f c Input to the time domain encoder configured by itself, and the time domain encoder is used to reconstruct the key view As a benchmark, each initial time domain sequence view f c Processing is performed to extract each initial time domain sequence view f c The first residual information of each match, at the same time, the video compression model converts each initial spatial sequence view v t Input to the spatial encoder configured by itself, and the spatial encoder is used to reconstruct the key view As the benchmark, each initial spatial sequence view v t Processing is performed to extract each initial spatial sequence view v t The second residual information matches each other. At the same time, the video compression model selects the target encoder that matches each initial spatiotemporal sequence view in the spatial domain encoder and the temporal domain encoder, and inputs each initial spatiotemporal sequence view into the target encoder that matches each other, so as to extract the third residual information corresponding to each initial spatiotemporal sequence view through the target encoder.

[0061] In a feasible implementation, the initial sequence view includes an initial time domain sequence view, an initial space domain sequence view, and an initial space-time domain sequence view; the step of "determining the residual information corresponding to each of the initial sequence views based on the reconstructed key view" in the above step S30 may specifically include steps S301 to S303:

[0062] Step S301: determining first residual information of the initial time-domain sequence view matching according to the reconstructed key view, wherein the time-domain sequence view is other frame views corresponding to the first viewpoint except the first frame view;

[0063] Step S302: determining a reference spatial sequence view that matches the initial spatial sequence view according to the reconstructed key view, and determining second residual information according to the reference spatial sequence view, wherein the spatial sequence view is a view corresponding to other viewpoints at the same time point as the initial key view;

[0064] Step S303: determining a target encoder that matches the initial spatiotemporal sequence view, and processing the initial spatiotemporal sequence view by the target encoder to determine third residual information.

[0065] For example, the electronic device generates a reconstruction key view After that, firstly transform each initial time domain sequence view f c Input to the time domain encoder configured by itself, and the time domain encoder is used to reconstruct the key view As a benchmark, each initial time domain sequence view f c Processing is performed to extract each initial time domain sequence view f c The first residual information of each match, at the same time, the video compression model converts each initial spatial sequence view v t Input to the spatial encoder configured by itself, and the spatial encoder is used to reconstruct the key view As a benchmark, determine the initial spatial sequence view v t View of the reference spatial sequence of each match The spatial encoder then follows the reference spatial sequence view For each initial spatial sequence view v t Processing is performed to extract each initial spatial sequence view v t Finally, the video compression model selects the target encoders that match each initial spatiotemporal sequence view in the spatial domain encoder and the temporal domain encoder, and inputs each initial spatiotemporal sequence view into the target encoders that match each initial spatiotemporal sequence view, so as to extract the third residual information corresponding to each initial spatiotemporal sequence view through the target encoder.

[0066] In a feasible implementation manner, the step of “determining the first residual information of the initial time domain sequence view matching according to the reconstructed key view” in the above step S301 may specifically include steps S3011 to S3012:

[0067] Step S3011: performing optical flow prediction processing on the reconstructed key view and the initial time domain sequence view through a time domain encoder to determine a first prediction frame matched by the initial time domain sequence view;

[0068] Step S3012: Compare the initial time domain sequence view with the first predicted frame to determine first residual information matching the initial time domain sequence view.

[0069] For example, the video compression model generates a reconstruction key view After that, the key view will be reconstructed further and the above-mentioned initial time domain sequence views f c Input to the time domain encoder configured by itself, the time domain encoder determines each initial time domain sequence view f c Reconstruct the time domain sequence view of the corresponding previous frame Optical flow prediction module within the temporal encoder to reconstruct key views Reconstruct the time domain sequence view with the previous frame Perform optical flow prediction for the reference image to determine the initial time domain sequence view f c The initial optical flow information m of each match t The optical flow prediction module then uses an optical flow encoder, a quantizer, an entropy encoder, and an optical flow decoder to calculate the initial optical flow information m t Perform encoding reconstruction to obtain reconstructed optical flow information The optical flow prediction module then reconstructs the optical flow information For the initial time domain sequence view f c Matching the previous frame to reconstruct the time domain sequence view Perform optical flow compensation to obtain each initial time domain sequence view f c The first prediction frame that matches each other, after that, the time domain encoder converts each first prediction frame and each matching initial time domain sequence view f c Compare and determine the initial time domain sequence view f c The first residual information corresponding to each other.

[0070] It should be noted that the optical flow prediction process is a technology widely used in the field of image processing and video analysis. It is used to estimate the motion of objects in an image sequence and make relevant predictions based on this. It can be understood that the specific process of optical flow prediction processing is an existing technology, so it will not be repeated here.

[0071] In a feasible implementation manner, the step of “determining the second residual information according to the reference spatial sequence view” in the above step S302 may specifically include steps S3021 to S3023:

[0072] Step S3021: Determine a reference spatial depth view corresponding to the reference spatial sequence view;

[0073] Step S3022: performing fusion prediction processing on the reference spatial sequence view and the reference spatial depth view through a spatial encoder to determine a second prediction frame matched with the initial spatial sequence view;

[0074] Step S3023: performing optical flow prediction compensation processing on the second predicted frame to obtain a third predicted frame, and comparing the initial spatial sequence view with the third predicted frame to determine second residual information matching the initial spatial sequence view.

[0075] For example, please refer to Figure 3 , Figure 3 This is a hybrid prediction processing flow chart of an embodiment of a multi-view video compression method of the present application, such as Figure 3As shown, after extracting each first residual information, the video compression model can also first determine each reference spatial sequence view The corresponding reference spatial depth view d t-1 , then the video compression model converts each reference spatial sequence view Each reference airspace depth view d t-1 Input to the hybrid prediction module configured in the spatial encoder, the hybrid prediction module uses each reference spatial sequence view Each reference airspace depth view d t-1 As a benchmark, for each initial spatial sequence view v t Perform homography transformation to generate each initial spatial sequence view v t The second predicted frames are matched respectively. At this time, due to the occlusion problem between the reference viewpoint and the first viewpoint, the second predicted frame usually has a hole phenomenon. Finally, the hybrid prediction module uses the optical flow prediction method to compensate the second predicted frame, thereby obtaining each initial spatial sequence view v t The spatial encoder then converts each third predicted frame and its matching initial spatial sequence view v t Compare to determine the initial spatial sequence view v t The second residual information of each is matched.

[0076] In a feasible implementation manner, after the step of "performing optical flow prediction compensation processing on the second predicted frame to obtain a third predicted frame" in the above step S3023, the multi-view video compression method of the present application may further include steps A10 to A20:

[0077] Step A10: inputting the third prediction frame into a filter network in the spatial domain encoder, so as to obtain a fourth prediction frame by smoothing the third prediction frame through the filter network;

[0078] Step A20: Determine second residual information of the initial spatial sequence view matching according to the fourth predicted frame.

[0079] In this embodiment, please refer to Figure 4 , Figure 4 FIG. 1 is a schematic diagram of a filter model structure involved in an embodiment of a multi-view video compression method of the present application, such as Figure 4As shown, after the hybrid prediction module in the above-mentioned spatial encoder obtains the third prediction frame, the hole filling strategy based on the optical flow prediction processing will cause the third prediction frame to have obvious boundaries in some regions, thereby causing the third prediction frame to be distorted. Therefore, the spatial encoder inputs each third prediction frame into the filtering network configured by itself, and the filtering network smoothes each third prediction frame to smooth the holes by relying on the spatial attention mechanism to obtain each fourth prediction frame, which is then used by the spatial encoder according to each fourth prediction frame and each initial spatial sequence view v t The above-mentioned second residual information is obtained.

[0080] In a feasible implementation manner, the step of “determining a target encoder for matching the initial spatiotemporal sequence views” in the above step S303 may specifically include steps S3031 to S3035:

[0081] Step S3031: performing optical flow prediction processing through the time domain encoder and the reconstructed key view to determine a fourth predicted frame matched with the initial spatiotemporal domain sequence view;

[0082] Step S3032: performing hybrid prediction processing on the initial spatiotemporal sequence view through a spatial domain encoder to determine a fifth prediction frame that matches the initial spatiotemporal sequence view;

[0083] Step S3033: determining a minimum distortion prediction frame according to the fourth prediction frame and the fifth prediction frame;

[0084] Step S3034: when it is determined that the minimum distortion prediction frame is the fourth prediction frame, determining the time domain encoder as the target encoder for the initial spatiotemporal domain sequence view matching;

[0085] Step S3035: when it is determined that the minimum distortion prediction frame is the fifth prediction frame, the spatial domain encoder is determined as the target encoder.

[0086] In this embodiment, after extracting each second residual information, the video compression model first inputs each initial spatiotemporal sequence view into the time domain encoder, and the time domain encoder then processes each initial spatiotemporal sequence view to determine the time domain prediction image P that matches each initial spatiotemporal sequence view according to the optical flow prediction method. f At the same time, the video compression model inputs each initial spatiotemporal sequence view into the spatial encoder, and the spatial encoder further processes each initial spatiotemporal sequence view to perform hybrid prediction based on the reference spatiotemporal sequence view and the reference spatiotemporal depth view corresponding to each initial spatiotemporal sequence view, thereby constructing a spatial prediction map P matching each initial spatiotemporal sequence view. v , then, the video compression model is based on each time domain prediction graph Pf And the spatial prediction map P v , reconstruct the reconstructed time domain sequence diagram matching each initial spatiotemporal sequence view and reconstruct the spatial sequence diagram The video compression model then matches the temporal prediction graph P of each initial spatiotemporal sequence view f And the spatial prediction map P v , and the matching reconstructed time domain sequence diagram Reconstructing Spatial Sequence Diagram Compare and determine the target prediction graph P with the least distortion t :

[0087] Among them, F min Indicates the selection of the prediction graph with the smallest distortion, f mse represents the MSE function;

[0088] Finally, if the video compression model determines the target prediction map P t The time domain prediction map P obtained by the time domain encoder f , then the time domain encoder is determined as the target encoder; similarly, if the video compression model determines the target prediction graph P t The spatial domain prediction map P obtained by the spatial domain encoder v , the spatial domain encoder is determined as the target encoder.

[0089] In this embodiment, when the electronic device needs to compress multi-viewpoint video data with six degrees of freedom, the electronic device first receives the multi-viewpoint video data to be compressed, and extracts a target picture group consisting of multiple sequence views corresponding to multiple different viewpoints in the multi-viewpoint video data to be compressed. After that, the electronic device splits the target picture group to determine the first frame sequence view corresponding to the first viewpoint in the target picture group as an initial key view used as a reference. At the same time, the electronic device extracts each initial time domain sequence view, each initial spatial domain sequence view, and each initial time and space domain sequence view in the target picture group based on the initial key view. Finally, the electronic device separates the initial key view, each initial time domain sequence view, each initial spatial domain sequence view, and each initial time and space domain sequence view. The figure is input into the video compression model configured by itself, and the video compression model obtains the reconstructed key view of the initial key view through the main encoder configured by itself. At the same time, the video compression model processes each initial time domain sequence view based on the reconstructed key view through the time domain encoder configured by itself, and obtains the first residual information that matches each initial time domain sequence view. At the same time, the video compression model processes each initial spatial domain sequence view based on the reconstructed key view through the spatial domain encoder configured by itself, and obtains the second residual information that matches each initial spatial domain sequence view. At the same time, the video compression model processes each initial time and space domain sequence view through the time domain encoder or the spatial domain encoder to obtain the third residual information that matches each initial time and space domain sequence view.

[0090] In this way, the present application solves the technical problem of low video compression efficiency in the related art, that is, the present application splits the multi-view video data into multiple target picture groups, thereby extracting the initial key view contained in the target picture group and other initial sequence views except the initial key view, and then processes the initial key view to obtain a reconstructed key view as a reference benchmark, and extracts the residual information corresponding to each initial sequence view based on the reconstructed key view, so that the electronic device can extract the residual information with a smaller data amount from the multi-view video data, so that the electronic device only needs to store and transmit the residual information with a smaller data amount to complete the storage and transmission operations of the multi-view video data, thereby achieving the technical effect of improving the video compression efficiency.

[0091] Based on the first embodiment of the present application, a second embodiment of the present application is proposed. In the second embodiment of the present application, the same or similar contents as those of the above embodiments can be referred to the above description and will not be described in detail later. On this basis, after the above step S30, the multi-viewpoint video compression method of the present application can also include steps B10 to B30:

[0092] Step B10: receiving each residual information;

[0093] Step B20: restoring the reconstructed sequence views that match the initial sequence views according to the residual information;

[0094] Step B30: Based on the spatiotemporal structure corresponding to the multi-view video data to be compressed, the reconstructed sequence views are integrated to restore the multi-view video data to be compressed.

[0095] In this embodiment, when the electronic device needs to decompress the compressed residual information to restore the multi-view video data, it first receives the residual information obtained by compressing the multi-view video data. Then, the electronic device inputs the residual information into a preset video decompression model. The video decompression model inputs the first residual information included in the residual information into a time domain decoder, and inputs the second residual information into a spatial domain decoder. The video decompression model then restores the first residual information through the time domain decoder to obtain the same residual information as the initial time domain sequence view f. c Respectively matched reconstructed time domain series views At the same time, the video decompression model further restores each second residual information through the spatial domain decoder to obtain the same initial spatial domain sequence view v t Reconstructed spatial sequence views of their respective matches At the same time, the video decompression model restores the third residual information through a time domain decoder or a spatial domain decoder to obtain each reconstructed time and space domain sequence view. At the same time, the video decompression model restores the reconstructed key view through a main decoder to obtain an initial key view. Finally, the electronic device restores the initial key view, each reconstructed spatial domain sequence view Each reconstructed time domain sequence view Each reconstructed spatiotemporal sequence view is transformed according to the spatiotemporal characteristics of the multi-view video data to restore the multi-view video data.

[0096] In this way, through the video decompression model, only the residual data with a smaller data volume can be processed, so as to restore the multi-viewpoint video data.

[0097] The present application provides an electronic device, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the multi-viewpoint video compression method in the above-mentioned embodiment 1.

[0098] Reference below Figure 5, which shows a schematic diagram of the structure of an electronic device suitable for implementing the embodiments of the present application. The electronic devices in the embodiments of the present application may include but are not limited to mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Portable Application Descriptions), PMPs (Portable Media Players), vehicle-mounted terminals (such as vehicle-mounted navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. Figure 5 The electronic device shown is merely an example and should not bring any limitation to the functions and scope of use of the embodiments of the present application.

[0099] like Figure 5 As shown, the electronic device may include a processing device 1001 (e.g., a central processing unit, a graphics processor, etc.), which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM: Read Only Memory) 1002 or a program loaded from a storage device 1003 to a random access memory (RAM: Random Access Memory) 1004. In RAM1004, various programs and data required for the operation of the electronic device are also stored. The processing device 1001, ROM1002, and RAM1004 are connected to each other through a bus 1005. An input / output (I / O) interface 1006 is also connected to the bus. Generally, the following systems can be connected to the I / O interface 1006: an input device 1007 including, for example, a touch screen, a touch pad, a keyboard, a mouse, an image sensor, a microphone, an accelerometer, a gyroscope, etc.; an output device 1008 including, for example, a liquid crystal display (LCD: Liquid Crystal Display), a speaker, a vibrator, etc.; a storage device 1003 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 1009. The communication device 1009 can allow the electronic device to communicate with other devices wirelessly or by wire to exchange data. Although the figure shows an electronic device with various systems, it should be understood that it is not required to implement or have all the systems shown. More or fewer systems can be implemented or have alternatively.

[0100] In particular, according to the embodiments disclosed in the present application, the process described above with reference to the flowchart can be implemented as a computer software program. For example, the embodiments disclosed in the present application include a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program includes a program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network through a communication device, or installed from a storage device 1003, or installed from a ROM 1002. When the computer program is executed by the processing device 1001, the above-mentioned functions defined in the method of the embodiment disclosed in the present application are executed.

[0101] The electronic device provided by the present application adopts the multi-view video compression method in the above embodiment, which can solve the technical problem of low video compression efficiency in the related art. Compared with the prior art, the beneficial effects of the electronic device provided by the present application are the same as the beneficial effects of the multi-view video compression method provided by the above embodiment, and other technical features in the electronic device are the same as the features disclosed in the method of the previous embodiment, which will not be repeated here.

[0102] It should be understood that the various parts disclosed in this application can be implemented by hardware, software, firmware or a combination thereof. In the description of the above embodiments, specific features, structures, materials or characteristics can be combined in any one or more embodiments or examples in a suitable manner.

[0103] The above is only a specific implementation of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art who is familiar with the present technical field can easily think of changes or substitutions within the technical scope disclosed in the present application, which should be included in the protection scope of the present application. Therefore, the protection scope of the present application should be based on the protection scope of the claims.

[0104] The present application provides a computer-readable storage medium having computer-readable program instructions (ie, computer programs) stored thereon, wherein the computer-readable program instructions are used to execute the multi-view video compression method in the above-mentioned embodiment.

[0105] The computer-readable storage medium provided in the present application may be, for example, a USB flash drive, but is not limited to electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems or devices, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM: Random Access Memory), a read-only memory (ROM: Read Only Memory), an erasable programmable read-only memory (EPROM: Erasable Programmable Read Only Memory or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM: CD-Read Only Memory), an optical storage device, a magnetic storage device, or any suitable combination of the above. In this embodiment, the computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in combination with an instruction execution system or device. The program code contained on the computer-readable storage medium may be transmitted using any appropriate medium, including but not limited to: wires, optical cables, RF (Radio Frequency: Radio Frequency), etc., or any suitable combination of the above.

[0106] The computer-readable storage medium may be included in the electronic device, or may exist independently without being installed in the electronic device.

[0107] The computer-readable storage medium carries one or more programs. When the one or more programs are executed by an electronic device, the electronic device: obtains multi-view video data to be compressed, and extracts a target picture group in the multi-view video data to be compressed; extracts an initial key view and initial sequence views contained in the target picture group, wherein the initial key view is a first frame view corresponding to a first viewpoint in the target picture group; processes the initial key view through a main encoder to obtain a reconstructed key view, and determines the corresponding matching residual information of each of the initial sequence views based on the reconstructed key view.

[0108] Computer program code for performing the operations of the present application may be written in one or more programming languages ​​or a combination thereof, including object-oriented programming languages ​​such as Java, Smalltalk, C++, and conventional procedural programming languages ​​such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a separate software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0109] The flow chart and block diagram in the accompanying drawings illustrate the possible architecture, function and operation of the system, method and computer program product according to various embodiments of the present application. In this regard, each square box in the flow chart or block diagram can represent a module, a program segment or a part of a code, and the module, the program segment or a part of the code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the square box can also occur in a sequence different from that marked in the accompanying drawings. For example, two square boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each square box in the block diagram and / or flow chart, and the combination of the square boxes in the block diagram and / or flow chart can be implemented with a dedicated hardware-based system that performs a specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.

[0110] The modules involved in the embodiments described in this application may be implemented by software or hardware, wherein the name of the module does not constitute a limitation on the unit itself in some cases.

[0111] The readable storage medium provided by the present application is a computer-readable storage medium, which stores computer-readable program instructions (i.e., computer programs) for executing the above-mentioned multi-view video compression method, and can solve the technical problem of low video compression efficiency in the related art. Compared with the prior art, the beneficial effects of the computer-readable storage medium provided by the present application are the same as the beneficial effects of the multi-view video compression method provided by the above-mentioned embodiment, and will not be described in detail here.

[0112] The present application also provides a computer program product, including a computer program, which implements the steps of the multi-view video compression method as described above when the computer program is executed by a processor.

[0113] The computer program product provided by the present application can solve the technical problem of low video compression efficiency in the related art. Compared with the prior art, the beneficial effects of the computer program product provided by the present application are the same as the beneficial effects of the multi-viewpoint video compression method provided by the above embodiment, which will not be repeated here.

[0114] The above descriptions are only some embodiments of the present application, and are not intended to limit the patent scope of the present application. All equivalent structural changes made using the contents of the present application specification and drawings under the technical concept of the present application, or direct / indirect applications in other related technical fields are included in the patent protection scope of the present application.

Claims

1. A method for compressing multi-viewpoint video, characterized in that: The multi-view video compression method comprises: Acquire multi-view video data to be compressed, and extract a target picture group from the multi-view video data to be compressed; Extracting an initial key view and initial sequence views contained in the target picture group, wherein the initial key view is a first frame view corresponding to a first viewpoint in the target picture group; The initial key view is processed by a main encoder to obtain a reconstructed key view, and residual information corresponding to each of the initial sequence views is determined based on the reconstructed key view.

2. The multi-view video compression method according to claim 1, characterized in that: The initial sequence view includes an initial time domain sequence view, an initial space domain sequence view and an initial time and space domain sequence view; The step of determining the residual information corresponding to each of the initial sequence views based on the reconstructed key views comprises: Determine first residual information of the initial time-domain sequence view matching according to the reconstructed key view, wherein the time-domain sequence view is other frame views corresponding to the first viewpoint except the first frame view; Determine a reference spatial sequence view that matches the initial spatial sequence view according to the reconstructed key view, and determine second residual information according to the reference spatial sequence view, wherein the spatial sequence view is a view corresponding to other viewpoints at the same time point as the initial key view; A target encoder matching the initial spatiotemporal domain sequence view is determined, and the initial spatiotemporal domain sequence view is processed by the target encoder to determine third residual information.

3. The multi-view video compression method according to claim 2, characterized in that: The step of determining the first residual information of the initial time domain sequence view matching according to the reconstructed key view comprises: Performing optical flow prediction processing on the reconstructed key view and the initial time-domain sequence view through a time-domain encoder to determine a first prediction frame that matches the initial time-domain sequence view; The initial time-domain sequence view and the first predicted frame are compared to determine first residual information matching the initial time-domain sequence view.

4. The multi-view video compression method according to claim 2, characterized in that: The step of determining the second residual information according to the reference spatial sequence view further includes: Determining a reference spatial domain depth view corresponding to the reference spatial domain sequence view; Performing fusion prediction processing on the reference spatial sequence view and the reference spatial depth view through a spatial encoder to determine a second prediction frame that matches the initial spatial sequence view; The second predicted frame is subjected to optical flow prediction compensation processing to obtain a third predicted frame, and the initial spatial sequence view is compared with the third predicted frame to determine second residual information matching the initial spatial sequence view.

5. The multi-view video compression method according to claim 4, characterized in that: After the step of performing optical flow prediction compensation processing on the second prediction frame to obtain a third prediction frame, the method further includes: Inputting the third predicted frame into a filter network in the spatial domain encoder, so as to obtain a fourth predicted frame by smoothing the third predicted frame through the filter network; Determine second residual information of the initial spatial sequence view matching according to the fourth predicted frame.

6. The multi-view video compression method according to claim 2, characterized in that: The step of determining a target encoder for matching the initial spatiotemporal sequence view further includes: Performing optical flow prediction processing through the time domain encoder and the reconstructed key view to determine a fourth prediction frame matched with the initial spatiotemporal domain sequence view; Performing hybrid prediction processing on the initial spatiotemporal sequence view through a spatial domain encoder to determine a fifth prediction frame matched by the initial spatiotemporal sequence view; Determine a minimum distortion prediction frame according to the fourth prediction frame and the fifth prediction frame; In a case where the minimum distortion prediction frame is determined to be the fourth prediction frame, determining the time domain encoder as a target encoder for view matching of the initial spatiotemporal domain sequence; In a case where the minimum distortion prediction frame is determined to be the fifth prediction frame, the spatial domain encoder is determined to be the target encoder.

7. The multi-view video compression method according to claim 1, wherein: After the step of determining the residual information corresponding to each of the initial sequence views based on the reconstructed key views, the method further includes: Receiving each of the residual information; Restoring the reconstructed sequence views that match the respective initial sequence views according to the respective residual information; Based on the spatiotemporal structure corresponding to the multi-view video data to be compressed, the reconstructed sequence views are integrated to restore the multi-view video data to be compressed.

8. An electronic device, characterized in that: The device includes: a main encoder, a spatial domain encoder, a temporal domain encoder, a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the computer program is configured to implement the steps of the multi-view video compression method as described in any one of claims 1 to 7.

9. A storage medium, characterized in that: The storage medium is a computer-readable storage medium, and a computer program is stored on the storage medium. When the computer program is executed by a processor, the steps of the multi-view video compression method according to any one of claims 1 to 7 are implemented.

10. A computer program product, characterized in that The computer program product comprises a computer program, and when the computer program is executed by a processor, the steps of the multi-view video compression method according to any one of claims 1 to 7 are implemented.