Underwater multi-view video prediction and compression method, device and system
Through sparse viewing 3D Gaussian splash reconstruction, virtual viewing angle reference frames are generated, and weighted fusion is combined with space-time reference frames, which solves the problem that underwater multi-view video prediction and compression methods in the prior art cannot effectively utilize the correlation between multi-view videos, and achieves efficient underwater multi-view video compression.
Patent Information
- Application Number
- CN202510499337.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-21
- Publication Date
- 2025-06-24
AI Technical Summary
The existing underwater multi-view video prediction and compression methods cannot effectively utilize the correlation between multi-view videos when processing high-speed motion videos, resulting in information redundancy and high computing overhead, and cannot achieve efficient video compression.
The sparse viewing angle 3D Gaussian splash reconstruction technology is used to generate virtual viewing angle supplementary reference frames, and weighted fusion is combined with time domain and airspace reference frames to predict and compress the target frame.
The information density of the time-domain reference frame is enhanced through virtual reference frames, breaking through the perspective limitations of physical cameras, improving the accuracy of target frame prediction, achieving efficient compression of underwater multi-view videos, and reducing storage overhead.
Smart Images

Figure CN120201203A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of underwater computer vision, and more specifically, relates to an underwater multi-view video prediction and compression method, device, and system. Background Art
[0002] When using an underwater camera to collect underwater motion video information, it is often necessary to deploy an underwater multi-camera array to monitor the entire process of high-speed motion. When using a small number of low-frame-rate cameras for collection, due to the limited field of view of the cameras, the amount of information obtained is limited, the entire process of underwater motion cannot be covered, and motion blur will occur when recording high-speed motion. When using multiple high-frame-rate cameras, although a complete and clear underwater motion process can be collected, a large amount of video data will be generated every time a motion process is recorded, resulting in a large storage overhead for the data. To solve this problem, an underwater multi-view video prediction and compression method can be used to efficiently compress the data of multiple cameras, so as to save the complete underwater motion video under the premise of lower storage overhead.
[0003] Currently, traditional multi-view video prediction and compression methods mainly use the redundant information between time and views for inter-frame prediction to improve the compression efficiency, ignoring the correlation between views of multi-view videos, and the information is still highly redundant. At the same time, the number of reference frames selected by general time-domain prediction and compression methods is limited, and insufficient information can be provided for prediction coding, resulting in insufficient accuracy of target frame prediction. In addition, although some newly emerged methods using 3DGS or neural radiance fields to assist compression can reduce the information redundancy between views to a certain extent, their high computational overhead and time overhead hinder the further improvement of compression coding performance. Summary of the Invention
[0004] In view of the above defects or improvement requirements of the prior art, the present invention provides an underwater multi-view video prediction and compression method, device, and system, aiming to improve the compression performance of underwater multi-view videos.
[0005] To achieve the above object, the present invention provides an underwater multi-view video prediction and compression method, including:
[0006] Obtain n key frames before the current moment t and adjacent to the current moment t in each underwater video stream of N synchronized underwater multi-view video streams as time-domain reference frames {K1, K2,..., K n};
[0007] For the time-domain reference frames {K1, K2,..., K n}Perform sparse-view 3D Gaussian splash reconstruction to obtain a Gaussian point set R; perform differentiable splash rendering on the Gaussian point set R to synthesize n new views with different perspectives from the time-domain reference frame, and use the n new views as the virtual reference frames {Q1, Q2, …, Q n};
[0008] Based on the n target frames in each underwater video stream and the frames in other underwater video streams at the same time, construct the spatial-domain reference frames {P1, P2, …, P n}; where the n target frames are the frames at the current time t and the subsequent n - 1 predicted times adjacent to it in each underwater video stream;
[0009] Calculate the motion vectors between the time-domain reference frame and the n target frames, and use the time-domain reference frame plus the corresponding motion vectors as the predicted frame I temporal ; Calculate the perspective offsets between the spatial-domain reference frame and the n target frames, and use the spatial-domain reference frame plus the corresponding perspective offsets as the predicted frame I spatial ; Use the virtual reference frame as the predicted frame I virtual ; Use the predicted frame I temporal 、I spatial and I virtual After weighted fusion, use it as the estimation of the n target frames in each underwater video stream
[0010] Calculate the residual R between the n target frames and their estimation and perform quantization and entropy coding on the residual R to complete compression.
[0011] Furthermore, perform sparse-view 3D Gaussian splash reconstruction on the time-domain reference frames {K1, K2, …, K n}, and obtain the Gaussian point set R, including:
[0012] Input the time-domain reference frame into a feature pyramid network to extract multi-scale features {F1, F2, …, F m}; where m is the number of scales;
[0013] Select one scale feature F m from the multi-scale features {F1, F2, …, F i to perform multi-view stereo depth estimation to obtain the corresponding depth map D i , specifically including:
[0014] According to the height H and width W of the feature F i to determine the number of depth plane cuts X, and use the feature F iWarp to each depth plane to obtain a cost volume of H*W*X;
[0015] Decompose the cost volume into three cascaded cost filters by means of three-plane decomposition; after normalizing the three cascaded cost filters, output the feature F i The probability of falling on each depth plane, and the probability constitutes the depth map D of the temporal reference frame i ;
[0016] According to the camera pose information, project the depth map D i backward in depth to obtain the Gaussian sphere center position coordinates μ of each Gaussian point cloud;
[0017] Through the MLP, decode the feature F i into other Gaussian sphere parameters of each Gaussian point cloud, and the other Gaussian sphere parameters include spherical harmonic coefficients s, Gaussian sphere radius r, color c, and transparency α; among them, the Gaussian sphere center position coordinates μ of each Gaussian point cloud and its corresponding other Gaussian sphere parameters constitute the Gaussian point set R.
[0018] Further, based on n target frames in each underwater video stream and frames in other underwater video streams at the same time, construct the spatial domain reference frames {P1, P2, …, P n} of each underwater video stream, including:
[0019] For the k-th underwater video stream, construct a view similarity matrix S by means of disparity matching k , and the element S k in the view similarity matrix S i,j represents the disparity between the i-th target frame in the k-th underwater video stream and the frame in the j-th underwater video stream at the same time; where, i ∈ {1, 2, …, n}, k, j ∈ {1, 2, …, N} and k ≠ j;
[0020] Select the top n views with higher element values in the view similarity matrix S k as the spatial domain reference frames of the k-th underwater video stream.
[0021] Further, the n key frames are the top n frames with the highest PSNR selected from multiple key frames before the current time t in each underwater video stream and adjacent to the current time t.
[0022] Further, perform disparity matching with epipolar geometry constraints on the spatial domain reference frames and the n target frames to obtain the view offset.
[0023] The present invention also provides an underwater multi-view video prediction and compression device, including a computer-readable storage medium and a processor;
[0024] The computer-readable storage medium is used to store executable instructions;
[0025] The processor is used to read the executable instructions stored in the computer-readable storage medium and execute the underwater multi-view video prediction and compression method described in any one of the above.
[0026] The present invention also provides an underwater multi-camera array acquisition system, including an underwater high-speed camera array and the underwater multi-view video prediction and compression device described above;
[0027] The underwater high-speed camera array is used to provide N channels of synchronous underwater multi-view video streams;
[0028] The underwater multi-view video prediction and compression device is used to perform video prediction and compression on the N channels of synchronous underwater multi-view video streams.
[0029] Further, the underwater high-speed camera array includes: 2 main-view cameras and N-2 auxiliary-view cameras; the main-view cameras are located at both ends of the underwater area, and the auxiliary-view cameras are located on the sides of the underwater area. The intervals of the auxiliary-view cameras are the same, and there is an overlap between the fields of view of adjacent auxiliary-view cameras.
[0030] The present invention also provides a computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, it implements the underwater multi-view video prediction and compression method described in any one of the above.
[0031] The present invention also provides a computer program product, including a computer program. When the computer program runs on a computer, it causes the computer to execute the underwater multi-view video prediction and compression method described in any one of the above.
[0032] Generally speaking, through the above technical solutions conceived by the present invention, the following beneficial effects can be achieved:
[0033] (1) The present invention proposes an efficient compression framework based on sparse-view 3D Gaussian splash (3DGS) reconstruction and spatio-temporal joint prediction. Specifically, the virtual views generated by 3D Gaussian splash reconstruction are used to supplement the reference frames, that is, the virtual reference frames. The newly added virtual views dynamically enhance the information density of the temporal reference frames, provide sufficient information for predictive coding, break through the physical camera view limitation, and improve the accuracy of target frame prediction. Based on the virtual reference frames, temporal reference frames, and spatial reference frames for target frame prediction, the time information, spatial information (representing the correlation between different views of multi-view video), and the newly added virtual views of the reference frames are fully utilized during the prediction process, enabling significant compression of data frames and improving the compression performance of underwater multi-view video.
[0034] (2) Preferably, in the sparse view 3DGS reconstruction method of the present invention, multi-view stereo matching (MVS) is combined with 3DGS to restore a three-dimensional scene from a small number of multi-view images, greatly reducing the computational cost and time cost; through the method of three-plane decomposition, the cost volume is decomposed into a cascade of three two-dimensional cost filters, reducing the computational cost and time cost brought by originally using three-dimensional convolution.
[0035] Generally speaking, the compression method based on sparse 3DGS (3D Gaussian splash) reconstruction proposed by the present invention generates adaptive Gaussian parameters through feature pyramid-guided MVS (multi-view stereo matching) depth estimation. The improved cascade cost volume structure compresses the memory occupancy of depth estimation. Combined with the channel attention feature pyramid network, the GPU utilization rate in the sparse reconstruction stage is significantly improved; combined with the joint prediction of three types of reference frames in time, space, and virtual space, it can better adapt to the prediction and compression of underwater multi-view videos. Description of the Drawings
[0036] Figure 1 It is a flowchart of the underwater multi-view video prediction and compression method in the embodiment of the present invention;
[0037] Figure 2 It is a top view of the arrangement of the underwater multi-camera array in the embodiment of the present invention;
[0038] Figure 3 It is a side view of the arrangement of the underwater multi-camera array in the embodiment of the present invention;
[0039] Figure 4 It is a schematic diagram of the sparse view 3DGS new view synthesis method in the embodiment of the present invention. Detailed Embodiments
[0040] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention. In addition, the technical features involved in the various embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.
[0041] Embodiment 1
[0042] As Figure 1 shown, the embodiment of the present invention provides an underwater multi-view video prediction and compression method, which mainly includes:
[0043] S1. Obtain n key frames in each underwater video stream of N synchronous underwater multi-view video streams as time-domain reference frames {K1, K2,..., K n}; Among them, the N-channel synchronous underwater multi-view video stream is collected by an N-channel high-speed camera array, and the N-channel high-speed camera array fully covers the underwater area; the selected n key frames are the n key frames before the current time t in this underwater video stream and adjacent to the current time t.
[0044] S2. Perform sparse view 3DGS (3D Gaussian splash) reconstruction on the time-domain reference frames {K1, K2,..., K n} of each underwater video stream to obtain a Gaussian point set R; and perform differentiable splash rendering on the Gaussian point set R to synthesize n new-view images with different views from the time-domain reference frames {K1, K2,..., K n}, and use the new rendered views as the virtual reference frames {Q1, Q2,..., Q n} of this underwater video stream.
[0045] S3. For each underwater video stream, construct a view similarity matrix using the method of disparity matching, where k represents the k-th underwater video stream, k ∈ {1, 2,..., N}; the element S k in the view similarity matrix S k i,j represents the disparity between the i-th target frame in the k-th underwater video stream and the frame in the j-th underwater video stream at the same time; among them, i ∈ {1, 2,..., n}, j ∈ {1, 2,..., N}, and k ≠ j, and the n target frames are the frames at the current time t and the subsequent n - 1 predicted times adjacent to it in each underwater video stream; select the top n views with higher element values in the view similarity matrix S k (the frames in the n underwater video streams at the same time corresponding to the top n values with higher element values in the view similarity matrix S k ) as the spatial-domain reference frames {P1, P2,..., P n} of this underwater video stream;
[0046] S4. Calculate the motion vectors between the time-domain reference frames {K1, K2,..., K n} and the n target frames, where the calculation method of the motion vectors is the prior art; add the corresponding motion vectors to the time-domain reference frames {K1, K2,..., K n} as the predicted frames I temporal of the n target frames. Calculate the view offsets between the spatial-domain reference frames {P1, P2,..., P n} and the n target frames, and add the corresponding view offsets to the spatial-domain reference frames {P1, P2,..., P n} as the predicted frames I spatial of the n target frames; the virtual reference frames {Q1, Q2,..., Q n}Directly serve as the predicted frame I of n target frames virtual ; For each underwater video stream, the predicted frame I temporal , the predicted frame I spatial , the predicted frame I virtual are weighted and fused to serve as the estimation of n target frames in this underwater video stream
[0047] S5. Calculate the residual R between the n target frames and their estimation . Quantize and entropy code the residual R, and store the residual to complete the compression.
[0048] Preferably, in S1, the N-channel high-speed camera array includes 2 main-view cameras and N - 2 auxiliary-view cameras. The main-view cameras are located at both ends of the underwater area, and the auxiliary-view cameras are located on the sides of the underwater area; the intervals of the auxiliary-view cameras are the same, and there is an overlap between the fields of view of adjacent auxiliary-view cameras, ensuring that their three-dimensional scenes can be stitched; the main-view cameras and the auxiliary-view cameras can fully cover the underwater area. As Figure 2 and Figure 3 shown, in the embodiment of the present invention, the underwater multi-camera array includes 2 main-view cameras (#21, #22) at the pool ends and 20 auxiliary-view cameras (#1 to #20) on the pool sides; the 20 auxiliary-view cameras are divided into 10 groups, each group consists of two cameras, one above water and one underwater, and a group of cameras is arranged every 5m to achieve full coverage of a 50m pool. There is a 5% - 10% overlapping area between the images captured by adjacent auxiliary-view cameras. In the embodiment of the present invention, the resolution of each channel is ≥1920×1080, and the frame rate is ≥100fps. In practical applications, the number of the camera array can be flexibly adjusted.
[0049] In S1, a key frame selector is used to select n key frames. The larger the number n of key frames, the richer the information obtained, and the more accurate the predicted target frames. Correspondingly, the data storage amount is larger. In practical applications, it is selected according to specific application requirements. Preferably, the first n frames with the highest PSNR are selected from multiple key frames before the current moment t in this underwater video stream and adjacent to the current moment t as the temporal reference frames {K1, K2, …, K n}.
[0050] Preferably, in S2, as Figure 4 shown, perform sparse-view 3DGS reconstruction on the temporal reference frames {K1, K2, …, K n} of each underwater video stream to obtain the Gaussian point set R, including:
[0051] S21. For the temporal reference frames {K I , K2, …, K n}Input to the feature pyramid network to extract multi-scale features {F1, F2, …, F m} of the multi-view images, which are used to characterize the view correlation between the time-domain reference frames; where m is the number of scales.
[0052] S22. Select one scale feature F from the multi-scale features {F1, F2, …, F m} i and perform MVS (Multi-View Stereo) depth estimation, which includes the following steps:
[0053] S221. Determine the number X of depth plane cuts according to the height H and width W of the feature F i , and warp the feature F i onto each depth plane. Each depth plane is a feature map of H*W, obtaining a cost volume CostVolume of H*W*X;
[0054] S222. Decompose the cost volume CostVolume by means of three-plane decomposition into three cascaded cost filters to achieve the cascading of multiple cost filters; and after normalizing the three cascaded cost filters, output the probabilities of the feature F i falling on each depth plane. The probabilities of each depth plane constitute the depth map D n of the time-domain reference frames {K1, K2, …, K i} (the two-dimensional representation of the scene);
[0055] S223. According to the camera pose information, perform depth back-projection on the depth map D i to obtain the Gaussian sphere center position coordinates μ of each Gaussian point cloud;
[0056] Among them, for the selection of the scale feature F in the multi-scale features {F1, F2, …, F m}, if a larger data compression ratio is required, select the feature with a lower resolution from the multi-scale features {F1, F2, …, F i}; if higher image quality is required, select the feature with a higher resolution from the multi-scale features {F1, F2, …, F m}. m}
[0057] S23. Decode the feature F i into other Gaussian sphere parameters of each Gaussian point cloud through the MLP, including the spherical harmonic coefficient s, the Gaussian sphere radius r, the color c, and the transparency α; the Gaussian sphere center position coordinates μ and other Gaussian sphere parameters of each Gaussian point cloud constitute the Gaussian point set R, that is, each point in the Gaussian point set R contains μ, s, r, c, α.
[0058] In S4, the predicted frame I of the i-th target frametemporal_i is the i-th time-domain reference frame K i plus its corresponding motion vector; the I of the i-th target frame spatial_i is the i-th spatial-domain reference frame P i plus its corresponding perspective offset. The perspective offset (disparity vector) between the target frame and the spatial-domain reference frame is obtained by performing epipolar geometry constraint-based disparity matching between the target frame and the spatial-domain reference frame.
[0059] In S5, the i-th target frame I i and its preliminary estimate The residual between them is R i :
[0060] In the spatio-temporal reference frame set dynamic construction mechanism in the embodiments of the present invention, in water, the fluctuation range of the peak signal-to-noise ratio of the compression coding performance index is reduced, and the stability is improved compared with the fixed reference frame scheme. Further, when encoding the three types of time-space-virtual reference frames, the motion vector, perspective deviation, and virtual reference frame are aggregated, and the entropy coding is combined to reduce the encoding and decoding delay and the computational energy consumption.
[0061] Embodiment 2
[0062] The embodiments of the present invention provide an underwater multi-view video prediction and compression device, including a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the steps of the underwater multi-view video prediction and compression method in Embodiment 1 above are implemented.
[0063] The related technical solutions are the same as above and will not be elaborated here.
[0064] Embodiment 3 The embodiments of the present invention provide a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the underwater multi-view video prediction and compression method in Embodiment 1 above are implemented.
[0065] Specifically, the memory may include a high-speed random access memory, and may also include a non-volatile memory, such as a hard disk, a memory, a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, at least one magnetic disk storage device, a flash memory device, or other volatile solid-state storage devices.
[0066] The related technical solutions are the same as above and will not be elaborated here.
[0067] Embodiment 4
[0068] An embodiment of the present invention provides a computer program product, including a computer program. When the computer program runs on a computer, it causes the computer to execute the steps of the underwater multi-view video prediction and compression method in Embodiment 1 above.
[0069] The related technical solutions are the same as above and will not be elaborated here.
[0070] Embodiment 5
[0071] An embodiment of the present invention provides an underwater multi-camera array acquisition system, including: an underwater high-speed camera array and the underwater multi-view video prediction and compression device provided in Embodiment 2.
[0072] The underwater high-speed camera array is used to provide N channels of synchronous underwater multi-view video streams;
[0073] The underwater multi-view video prediction and compression device provided in Embodiment 2 is used to implement underwater multi-view video prediction and compression based on N channels of synchronous underwater multi-view video streams.
[0074] The related technical solutions are the same as above and will not be elaborated here.
[0075] Those skilled in the art can easily understand that the above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, and improvements made within the spirit and principles of the present invention should be included in the protection scope of the present invention.
Claims
1. A method for underwater multi-view video prediction and compression, characterized in that: include: Obtain n key frames before the current time t and adjacent to the current time t in each underwater video stream of N synchronized underwater multi-view video streams as temporal reference frames {K1, K2, …, K n }; For the time domain reference frames {K1, K2, ..., K n }Perform sparse perspective 3D Gaussian splash reconstruction to obtain a Gaussian point set R; perform differentiable splash rendering on the Gaussian point set R to synthesize n new views different from the perspective of the temporal reference frame, and use the n new views as virtual reference frames {Q1, Q2, …, Q n }; Based on the n target frames in each underwater video stream and the frames in other underwater video streams at the same time, a spatial reference frame {P1, P2, …, P n }; wherein the n target frames are frames at the current time t and the next n-1 predicted time moments adjacent to the current time t in each underwater video stream; Calculate the motion vector between the temporal reference frame and n target frames, and use the temporal reference frame plus the corresponding motion vector as the prediction frame I temporal ; Calculate the viewing angle offset between the spatial reference frame and the n target frames, and use the spatial reference frame plus the corresponding viewing angle offset as the prediction frame I spatial ; Using the virtual reference frame as the prediction frame I virtual ; The predicted frame I temporal ,I spatial and I virtual After weighted fusion, it is used as the estimation of n target frames in each underwater video stream Calculate the n target frames and their estimated The residual R between them is quantized and entropy encoded to complete the compression.
2. The underwater multi-view video prediction and compression method according to claim 1, characterized in that: For the time domain reference frames {K1, K2, ..., K n }Perform sparse view 3D Gaussian splash reconstruction to obtain the Gaussian point set R, including: The temporal reference frame is input into the feature pyramid network to extract multi-scale features {F 1, F 2, …,F m }; where m is the scale number; Select the multi-scale feature {F 1, F 2, …,F m A scale feature F in i Perform multi-view stereo depth estimation to obtain the corresponding depth map D i , including: According to the feature F i The height H and width W determine the number of depth plane cuts X, and the feature F i Warp to each depth plane to get the cost volume of H*W*X; The cost volume is decomposed into three cascaded cost filters by three-plane decomposition; after the three cascaded cost filters are normalized, the feature F is output. i The probability of falling on each depth plane constitutes the depth map D of the temporal reference frame. i ; According to the camera pose information, the depth map D i Perform depth back projection to obtain the coordinates μ of the center position of the Gaussian sphere of each Gaussian point cloud; Through MLP, the feature F i Decoded into other Gaussian sphere parameters of each Gaussian point cloud, the other Gaussian sphere parameters include spherical harmonic coefficient s, Gaussian sphere radius r, color c, and transparency α; wherein the Gaussian sphere center position coordinates μ of each Gaussian point cloud and its corresponding other Gaussian sphere parameters constitute the Gaussian point set R.
3. The underwater multi-view video prediction and compression method according to claim 1, characterized in that: Based on the n target frames in each underwater video stream and the frames in other underwater video streams at the same time, a spatial reference frame {P1, P2, …, P n },include: For the k-th underwater video stream, the view similarity matrix S is constructed by disparity matching. k , the perspective similarity matrix S k The element S in k i,j represents the disparity between the i-th target frame in the k-th underwater video stream and the frame in the j-th underwater video stream at the same time; where i∈{1,2,…,n}, k,j∈{1,2,…,N} and k≠j; Select the perspective similarity matrix S k The first n perspectives with higher element values are used as the spatial reference frames of the kth underwater video stream.
4. The underwater multi-view video prediction and compression method according to any one of claims 1 to 3, characterized in that: The n key frames are the first n frames with the highest PSNR selected from a plurality of key frames before the current time t and adjacent to the current time t in each underwater video stream.
5. The underwater multi-view video prediction and compression method according to claim 4, characterized in that: The spatial reference frame is subjected to epipolar geometrically constrained disparity matching with n target frames to obtain the viewing angle offset.
6. An underwater multi-view video prediction and compression device, characterized in that: comprising a computer readable storage medium and a processor; The computer-readable storage medium is used to store executable instructions; The processor is used to read the executable instructions stored in the computer-readable storage medium to execute the underwater multi-view video prediction and compression method described in any one of claims 1-5.
7. An underwater multi-camera array acquisition system, characterized in that: It comprises an underwater high-speed camera array and the underwater multi-view video prediction and compression device as claimed in claim 6; The underwater high-speed camera array is used to provide N-channel synchronous underwater multi-view video streams; The underwater multi-view video prediction and compression device is used to perform video prediction and compression on the N-channel synchronous underwater multi-view video streams.
8. The underwater multi-camera array acquisition system according to claim 7, characterized in that: The underwater high-speed camera array includes: 2 main-view cameras and N-2 auxiliary-view cameras; the main-view cameras are located at both ends of the underwater area, and the auxiliary-view cameras are located on the sides of the underwater area. The auxiliary-view cameras are spaced uniformly, and there is overlap between the fields of view of adjacent auxiliary-view cameras.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the underwater multi-view video prediction and compression method as described in any one of claims 1-5 is implemented.
10. A computer program product, characterized in that It comprises a computer program, which, when running on a computer, enables the computer to execute the underwater multi-view video prediction and compression method according to any one of claims 1 to 5.