Coding method and apparatus for multi-view video, electronic device and medium
By establishing spatial correspondences between viewpoint coordinate systems and pre-defined encoding rules, the problem of difficulty in utilizing common content between viewpoints in multi-viewpoint video encoding is solved, thereby improving encoding efficiency and accuracy and ensuring the efficient generation of encoded information.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- CHINA TELECOM CLOUD TECH CO LTD
- Filing Date
- 2025-11-19
- Publication Date
- 2026-06-04
AI Technical Summary
Existing multi-view video coding standards fail to effectively utilize the spatial relationships between different viewpoints, making it difficult to accurately capture similar parts between video frames captured from different viewpoints when the camera positions differ too much. This makes it difficult to utilize common content between viewpoints, thus affecting coding efficiency.
By obtaining the installation positions of the main viewpoint and auxiliary viewpoint cameras, a coordinate system for the main viewpoint and auxiliary viewpoint is established to determine the spatial correspondence. The main viewpoint video frame is converted into a reference viewpoint video frame in the auxiliary viewpoint coordinate system using rotation and translation matrices, and then encoded using preset encoding rules to fill in the missing details.
It improves the efficiency and accuracy of multi-view video coding, ensures the efficient generation of encoded information, makes full use of common content between viewpoints, and reduces the transmission of redundant information.
Smart Images

Figure CN2025136114_04062026_PF_FP_ABST
Abstract
Description
A method, apparatus, electronic device, and medium for encoding multi-view video.
[0001] Cross-reference to related applications
[0002] This application claims priority to Chinese Patent Application No. 202411706807.7, filed on November 26, 2024, entitled "A method, apparatus, electronic device and medium for encoding multi-view video", the entire contents of which are incorporated herein by reference. Technical Field
[0003] This application relates to the field of image processing, and in particular to a method, apparatus, electronic device, and medium for encoding multi-view video. Background Technology
[0004] With the development of technologies such as virtual reality (VR), augmented reality (AR), and panoramic video, multi-view content is growing exponentially. Specifically, it allows for the capture of video of the same scene from multiple viewpoints, such as in a studio or using a camera array. The captured multi-view video is then transmitted over a network and presented on various devices, increasing viewer immersion and interactivity.
[0005] However, the exponential increase in the amount of video data from multiple sources has brought great challenges to transmission. Existing multi-view video coding standards cannot effectively utilize the spatial relationships between different viewpoints. When the camera positions differ too much, it is difficult to accurately obtain the similar parts between video frames captured from different viewpoints. The common content between viewpoints is difficult to utilize, which seriously affects the efficiency of multi-view video coding. Summary of the Invention
[0006] In view of the above problems, some embodiments of this application propose a method, apparatus, electronic device and medium for encoding multi-view video.
[0007] In a first aspect of this application, a method for encoding multi-view video is provided, characterized in that the multi-view video includes a main viewpoint video captured by a main viewpoint camera and an auxiliary viewpoint video captured by an auxiliary viewpoint camera, the method comprising:
[0008] Obtain the installation positions of the main viewpoint camera and the auxiliary viewpoint camera;
[0009] Establish a main viewpoint coordinate system based on the installation position of the main viewpoint camera, and establish an auxiliary viewpoint coordinate system based on the installation position of the auxiliary viewpoint camera;
[0010] The spatial correspondence between the main viewpoint coordinate system and the auxiliary viewpoint coordinate system is determined based on the installation positions of the main viewpoint camera and the auxiliary viewpoint camera.
[0011] Obtain the main viewpoint video frame sequence from the main viewpoint video and the auxiliary viewpoint video sequence from the auxiliary viewpoint video;
[0012] Based on the spatial correspondence, the main viewpoint video frame sequence in the main viewpoint coordinate system will be converted into a reference viewpoint video frame sequence in the auxiliary viewpoint coordinate system;
[0013] The main viewpoint video frame sequence and the auxiliary viewpoint video frame sequence are encoded using preset encoding rules and reference viewpoint video frame sequences to obtain the encoding information of multi-viewpoint video.
[0014] In some embodiments, a main viewpoint camera and an auxiliary viewpoint camera are positioned in a preset scene, and the spatial correspondence includes a rotational correspondence and / or a translational correspondence. Determining the spatial correspondence between the main viewpoint coordinate system and the auxiliary viewpoint coordinate system based on the installation positions of the main viewpoint camera and the auxiliary viewpoint camera includes:
[0015] Select at least one common feature point in the preset scene;
[0016] The main viewpoint feature coordinates in the main viewpoint coordinate system are determined based on the spatial relationship between the installation position of the main viewpoint camera and the common feature points.
[0017] The auxiliary viewpoint feature coordinates in the auxiliary viewpoint coordinate system are determined based on the spatial relationship between the installation position of the auxiliary viewpoint camera and the common feature points.
[0018] Use the feature coordinates of the main viewpoint and the feature coordinates of the auxiliary viewpoint as matching point pairs of common feature points;
[0019] Determine the rotational and / or translational correspondences between the main viewpoint coordinate system and the auxiliary viewpoint coordinate system based on the matching point pairs.
[0020] In some embodiments, converting a sequence of main viewpoint video frames in the main viewpoint coordinate system into a sequence of reference viewpoint video frames in the auxiliary viewpoint coordinate system according to spatial correspondence includes:
[0021] Obtain at least one main viewpoint video frame from the main viewpoint video frame sequence in the main viewpoint coordinate system;
[0022] Based on the rotation correspondence, the main viewpoint video frame is converted into a main viewpoint rotated video frame.
[0023] Based on the translation correspondence, the main viewpoint rotated video frame is converted into a reference viewpoint video frame in the auxiliary viewpoint coordinate system;
[0024] At least one reference viewpoint video frame is combined in chronological order to form a reference viewpoint video frame sequence in the auxiliary viewpoint coordinate system.
[0025] In some embodiments, the preset encoding rules include inter-frame encoding rules and inter-viewpoint encoding rules. The preset encoding rules and a reference viewpoint video frame sequence are used to encode the main viewpoint video frame sequence and the auxiliary viewpoint video frame sequence to obtain the encoding information of the multi-viewpoint video, including:
[0026] The main viewpoint video frame sequence is encoded using inter-frame coding rules to obtain the main viewpoint video coding information;
[0027] The auxiliary viewpoint video frame sequence is encoded using inter-frame coding rules, inter-viewpoint coding rules, and reference viewpoint video frame sequences to obtain auxiliary viewpoint video coding information;
[0028] The encoding information for multi-view video is determined based on the primary viewpoint video encoding information and the secondary viewpoint video encoding information.
[0029] In some embodiments, the auxiliary viewpoint video frame sequence is encoded using inter-frame coding rules, inter-viewpoint coding rules, and a reference viewpoint video frame sequence to obtain auxiliary viewpoint video coding information, including:
[0030] The auxiliary viewpoint video frame sequence is encoded using inter-frame coding rules to obtain the first coded information of the auxiliary viewpoint;
[0031] The auxiliary viewpoint video frame sequence is encoded using inter-viewpoint coding rules and a reference viewpoint video frame sequence to obtain the second coding information of the auxiliary viewpoint.
[0032] The auxiliary viewpoint video coding information is determined based on the first and second auxiliary viewpoint coding information.
[0033] In some embodiments, the main viewpoint video and the secondary viewpoint video are captured within a preset time period. After converting the main viewpoint video frame sequence in the main viewpoint coordinate system into a reference viewpoint video frame sequence in the secondary viewpoint coordinate system according to spatial correspondence, the method further includes:
[0034] Obtain any reference viewpoint video frame from the reference viewpoint video frame sequence;
[0035] The region to be detected in the reference viewpoint video frame is determined based on the spatial correspondence.
[0036] A sliding detection window is set in the area to be detected; the size of the sliding detection window is determined by the distance between the installation positions of the main viewpoint camera and the auxiliary viewpoint camera; the sliding detection window includes several pixels to be detected in the reference viewpoint video frame;
[0037] During the sliding detection window's movement across the detection area, the grayscale values of the pixels to be detected in the sliding detection window and the median grayscale values of several pixels to be detected in the sliding detection window are obtained.
[0038] Determine the grayscale difference between the grayscale value of any pixel to be detected in the sliding detection window and the median grayscale value;
[0039] If the grayscale difference is greater than the preset threshold, the pixel corresponding to the median grayscale value will be used to replace the pixel to be detected.
[0040] After the sliding detection window traverses the area to be detected, the reference viewpoint video frames formed by replacing the pixels to be detected are combined to form a new reference viewpoint video frame sequence.
[0041] In some embodiments, during the sliding detection window's movement across the detection area, acquiring the grayscale values of the pixels to be detected within the sliding detection window and the median grayscale value of a plurality of pixels to be detected within the sliding detection window includes:
[0042] As the sliding detection window slides across the area to be detected, the grayscale values of several pixels to be detected in the sliding detection window are acquired in real time.
[0043] Sort the gray values of several pixels to be detected in ascending order;
[0044] The median of the sorted gray values is used as the median gray value of several pixels to be detected.
[0045] In a second aspect of this application, a multi-view video encoding apparatus is also provided, characterized in that the multi-view video includes a main view video captured by a main view camera and an auxiliary view video captured by an auxiliary view camera, and the apparatus includes:
[0046] The installation location acquisition module is used to acquire the installation location of the main viewpoint camera and the installation location of the auxiliary viewpoint camera.
[0047] The coordinate system establishment module is used to establish a main viewpoint coordinate system based on the installation position of the main viewpoint camera and an auxiliary viewpoint coordinate system based on the installation position of the auxiliary viewpoint camera.
[0048] The spatial correspondence determination module is used to determine the spatial correspondence between the main viewpoint coordinate system and the auxiliary viewpoint coordinate system based on the installation positions of the main viewpoint camera and the auxiliary viewpoint camera.
[0049] The video frame sequence acquisition module is used to acquire the main viewpoint video frame sequence in the main viewpoint video and the auxiliary viewpoint video frame sequence in the auxiliary viewpoint video;
[0050] The spatial transformation module is used to convert the main viewpoint video frame sequence in the main viewpoint coordinate system into a reference viewpoint video frame sequence in the auxiliary viewpoint coordinate system according to the spatial correspondence.
[0051] The encoding module is used to encode the main viewpoint video frame sequence and the auxiliary viewpoint video frame sequence using preset encoding rules and reference viewpoint video frame sequences to obtain the encoding information of multi-viewpoint video.
[0052] In a third aspect of this application, an electronic device is also provided, characterized in that it includes a processor, a memory, and a computer program stored in the memory and capable of running on the processor, wherein the computer program, when executed by the processor, implements the multi-view video encoding method described above.
[0053] In a fourth aspect of this application, a computer-readable storage medium is also provided, characterized in that a computer program is stored on the computer-readable storage medium, and when the computer program is executed by a processor, it implements the above-described multi-view video encoding method.
[0054] Some embodiments of this application have the following advantages:
[0055] In some embodiments of this application, the installation positions of the main viewpoint camera and the auxiliary viewpoint camera are obtained. A main viewpoint coordinate system is established based on the installation position of the main viewpoint camera, and an auxiliary viewpoint coordinate system is established based on the installation position of the auxiliary viewpoint camera. The spatial correspondence between the main viewpoint coordinate system and the auxiliary viewpoint coordinate system is determined based on the installation positions of the main viewpoint camera and the auxiliary viewpoint camera. The main viewpoint video frame sequence in the main viewpoint video and the auxiliary viewpoint video frame sequence in the auxiliary viewpoint video are obtained. Based on the spatial correspondence, the main viewpoint video frame sequence in the main viewpoint coordinate system is converted into a reference viewpoint video frame sequence in the auxiliary viewpoint coordinate system. The main viewpoint video frame sequence and the auxiliary viewpoint video frame sequence are encoded using a preset encoding rule and the reference viewpoint video frame sequence to obtain the encoding information of the multi-viewpoint video. This application first converts the main viewpoint video frame sequence into a reference viewpoint video frame sequence in an auxiliary viewpoint coordinate system. Then, it encodes both the main viewpoint and auxiliary viewpoint video frame sequences using preset encoding rules and the reference viewpoint video frame sequence. This solves the problem of unutilizing common content between viewpoints and improves encoding efficiency. It ensures that the encoded information for multi-viewpoint video can be generated efficiently and accurately. Attached Figure Description
[0056] To more clearly illustrate the technical solutions in some embodiments of this application or in the prior art, the accompanying drawings used in the description of some embodiments or in the prior art will be briefly introduced below.
[0057] Figure 1 is a schematic flowchart of a multi-viewpoint independent encoding and decoding provided in some embodiments of this application;
[0058] Figure 2 is a schematic diagram of a multi-view video coding architecture provided in some embodiments of this application;
[0059] Figure 3 is a flowchart of the steps of a multi-view video encoding method provided in some embodiments of this application;
[0060] Figure 4 is a schematic diagram of the spatial relationship transformation process provided in some embodiments of this application;
[0061] Figure 5 is a schematic diagram of the framework of multi-view coding based on spatial relationship transformation provided in some embodiments of this application;
[0062] Figure 6 is a schematic diagram of different viewpoint camera shooting angles provided in some embodiments of this application;
[0063] Figure 7 is a schematic diagram illustrating the loss of video frame detail information provided in some embodiments of this application;
[0064] Figure 8 is a schematic diagram of the framework of a multi-view video coding method combining spatial relationship transformation and detail filling provided in some embodiments of this application;
[0065] Figure 9 is a schematic diagram of the structure of a multi-view video encoding device provided in some embodiments of this application. Detailed Implementation
[0066] To make the objectives, technical solutions, and advantages of some embodiments of this application clearer, the various embodiments of this application will be described in detail below with reference to the accompanying drawings. However, those skilled in the art will understand that many technical details are presented in the various embodiments of this application to facilitate a better understanding of the application. However, the technical solutions claimed in this application can be implemented even without these technical details and various changes and modifications based on the following embodiments. The division of some embodiments below is for ease of description and should not constitute any limitation on the specific implementation of this application. Some embodiments can be combined with and referenced to each other without contradiction.
[0067] With the development of technologies such as Virtual Reality (VR), Augmented Reality (AR), and panoramic video, multi-view content is growing exponentially. Specifically, video of the same scene can be captured from multiple viewpoints, such as in a studio or using a camera array. The captured multi-view video is then transmitted over a network and presented on terminals in various ways, increasing viewer immersion and interactivity. However, the exponential increase in the amount of data from multiple video streams presents significant challenges to transmission. Currently, a large amount of multi-view video content does not use dedicated encoding standards; instead, it is encoded separately for each viewpoint using common standards (such as H.264 / AVC and H.265 / HEVC).
[0068] It should be noted that multi-view coding technology can be a technique that utilizes redundant information between different cameras or viewpoints to achieve efficient compression and encoding. In some embodiments of this application, multi-view video can be encoded using multi-view coding technology.
[0069] HEVC (High Efficiency Video Coding) is a next-generation video coding standard, mainly including modules such as transform, quantization, entropy coding, and intra-frame and inter-frame prediction. Compared with previous international standards, HEVC introduces new coding techniques in each module, improving compression efficiency. In some embodiments of this application, intra-frame coding or inter-frame coding can be performed with reference to the principles of HEVC.
[0070] MV-HEVC (MultiView-HEVC, High-Efficiency Video Coding): This is a high-level syntax extension of HEVC and can be seen as an implementation of multiview coding on the high-efficiency video coding standard. In some embodiments of this application, encoding can also be performed with reference to the principles of MV-HEVC.
[0071] Referring to Figure 1, a flowchart illustrating a multi-viewpoint independent encoding and decoding process provided in some embodiments of this application is shown. For the video of each viewpoint from viewpoint 1, viewpoint 2 to viewpoint n, the video is fed into a general encoder and a general decoder, respectively. Then, the decoding results of the video of each viewpoint are combined for multi-viewpoint video playback.
[0072] Since multi-view video targets the same scene, video frames from different viewpoints at the same time share many commonalities. Therefore, to leverage the relationships between multiple viewpoints, MV-HEVC (MultiView-HEVC) can be used for efficient multi-view video encoding. The MV-HEVC standard introduces the concept of a viewpoint, typically represented by View 0 (primary view), View 1 (secondary view), and others. The primary view is encoded using the basic HEVC rule, while each frame of the secondary viewpoints is encoded on top of the basic HEVC, with an additional inter-viewpoint reference frame.
[0073] Referring to Figure 2, a schematic diagram of a multi-view video coding architecture provided by some embodiments of this application is shown. Figure 2 illustrates a three-view MV-HEVC coding architecture, where viewpoint 0 is the main view. The main view uses HEVC coding, which can consider inter-frame correlation on the time axis, while the auxiliary view considers both inter-view and temporal correlation. For example, at time T2 of viewpoint 2, the B-frame considers both the P-frame at time T0 of viewpoint 2 and the B-frame at time T4 of viewpoint 2, as well as the B-frame at time T2 of viewpoint 0, utilizing the inter-view relationships.
[0074] However, there are some problems with MV-HEVC:
[0075] First, MV-HEVC does not utilize the spatial relationships between different viewpoints. Taking Figure 2 as an example, when performing inter-frame coding for viewpoint 2, it utilizes both the temporal correlation of frames and the correlation between viewpoints. For instance, when coding the B-frame at time T2 of viewpoint 2, it considers both the P-frame at time T0 and the B-frame at time T4 of viewpoint 2, as well as the B-frame at time T2 of viewpoint 0, thus utilizing the relationships between viewpoints. However, when utilizing the correlation between viewpoints, spatial transformation is not performed. For example, when inter-frame coding the B-frame at time T2 of viewpoint 2, the B-frame at time T2 of viewpoint 0 is used, and encoding is performed by calculating the similarity between the B-frame at time T2 of viewpoint 0 and the B-frame at time T2 of viewpoint 2.
[0076] Although multiple viewpoints are focused on the same scene, due to the different positions and angles of the cameras, the content in the video frames captured by multiple viewpoints often has fixed rotation and translation. When the rotation and translation are large, it is difficult to obtain the similar parts between frames by motion estimation and compensation techniques alone, which will greatly limit the encoding effect.
[0077] Secondly, video frames transformed based on spatial relationships in MV-HEVC may suffer from detail loss. Due to the different camera positions and angles at different viewpoints, the content captured by each viewpoint varies. Therefore, when a video frame from one viewpoint is transformed to another, some details may be lost. This can lead to significant pixel differences in the areas of lost detail between the two frames, affecting the encoding efficiency between viewpoints.
[0078] Therefore, some embodiments of this application propose a multi-view video encoding method. When encoding multi-view video, the spatial relationships between multiple viewpoints can be utilized. Based on the spatial positional relationships of the cameras, rotation and translation matrices are calculated, and then the video frame of the main viewpoint is converted to the secondary viewpoint based on these two matrices. When performing inter-frame encoding on the secondary viewpoint, its similarity to the converted main viewpoint video frame is utilized to avoid the problem of difficulty in obtaining similar parts caused by large differences in camera positions, thereby improving the efficiency of multi-view video encoding.
[0079] In addition, some embodiments of this application address the issue of low coding efficiency between viewpoints due to the loss of detail information in the converted video frames. By setting a sliding window and a discrimination threshold to fill in the lost detail information, the coding efficiency between viewpoints is improved. On the other hand, based on the influence of different rotations and translations on the position of the lost detail, the search area is adaptively searched to improve the efficiency of detail filling.
[0080] Referring to FIG3, a flowchart of the steps of a multi-view video encoding method provided in some embodiments of this application is shown. The multi-view video includes a main view video captured by a main view camera and an auxiliary view video captured by an auxiliary view camera.
[0081] It should be noted that multi-view video (MVV) refers to video of the same event or scene captured from multiple different perspectives. MVV can provide a richer visual experience, allowing users to view the same event or scene from different angles, enhancing immersion and interactivity. In some embodiments of this application, multi-view video needs to be encoded.
[0082] A master view camera is a camera used in a multi-view video system to capture the primary perspective. It is typically located in the center of the scene, providing the most dominant viewpoint.
[0083] First-viewpoint video refers to video captured by a first-viewpoint camera. First-viewpoint video typically contains the main perspective of the scene, providing the most important visual information.
[0084] In some embodiments of this application, a main viewpoint camera can be used to capture main viewpoint video.
[0085] A secondary viewpoint camera is a camera used in a multi-view video system to capture supplementary perspectives. Secondary viewpoint cameras are typically located around the primary viewpoint camera, providing different viewpoints.
[0086] Auxiliary viewpoint video refers to video captured by an auxiliary viewpoint camera. Auxiliary viewpoint video typically includes an auxiliary perspective of the scene, providing richer visual information.
[0087] In some embodiments of this application, auxiliary viewpoint video can be acquired using an auxiliary viewpoint camera.
[0088] The method specifically includes the following steps:
[0089] S101: Obtain the installation position of the main viewpoint camera and the installation position of the auxiliary viewpoint camera;
[0090] In some embodiments of this application, the installation positions of the main viewpoint camera and the auxiliary viewpoint camera can be obtained. The installation positions can be in the form of three-dimensional coordinates. The installation positions may also include the orientation information of the main viewpoint camera and the auxiliary viewpoint camera. There may be one or more auxiliary viewpoint cameras.
[0091] S102: Establish a main viewpoint coordinate system based on the installation position of the main viewpoint camera, and establish an auxiliary viewpoint coordinate system based on the installation position of the auxiliary viewpoint camera.
[0092] In some embodiments of this application, a main viewpoint coordinate system can be established based on the installation position of the main viewpoint camera, and an auxiliary viewpoint coordinate system can be established based on the installation position of the auxiliary viewpoint camera. That is, a main viewpoint coordinate system based on the main viewpoint camera and an auxiliary viewpoint coordinate system based on the auxiliary viewpoint camera can be constructed based on the three-dimensional coordinates and orientation information of the main viewpoint camera and the auxiliary viewpoint camera.
[0093] S103: Determine the spatial correspondence between the main viewpoint coordinate system and the auxiliary viewpoint coordinate system based on the installation positions of the main viewpoint camera and the auxiliary viewpoint camera.
[0094] In some embodiments of this application, the spatial correspondence between the main viewpoint coordinate system and the auxiliary viewpoint coordinate system can be determined based on the installation positions of the main viewpoint camera and the auxiliary viewpoint camera.
[0095] In practical implementation, camera calibration can be used to obtain the internal parameters (such as focal length, principal point position, etc.) and external parameters (such as rotation and translation matrices) of each camera. Based on the spatial position of the cameras, the rotation and translation matrices between the main viewpoint camera and each auxiliary viewpoint camera are calculated. These rotation and translation matrices are then used as the spatial correspondence between the main viewpoint coordinate system and the auxiliary viewpoint coordinate systems.
[0096] S104: Obtain the main viewpoint video frame sequence in the main viewpoint video and the auxiliary viewpoint video frame sequence in the auxiliary viewpoint video.
[0097] In some embodiments of this application, a sequence of main viewpoint video frames in a main viewpoint video and a sequence of auxiliary viewpoint video frames in an auxiliary viewpoint video can be obtained. The main viewpoint video and the auxiliary viewpoint video can be captured at the same time for the same scene.
[0098] In a specific implementation, the main viewpoint video can be converted into a main viewpoint video frame sequence, and the secondary viewpoint video can be converted into a secondary viewpoint video frame sequence. The main viewpoint video frame sequence and the secondary viewpoint video frame sequence are the main viewpoint video frames and secondary viewpoint video frames corresponding to the same time series.
[0099] S105: Based on the spatial correspondence, the main viewpoint video frame sequence in the main viewpoint coordinate system will be converted into a reference viewpoint video frame sequence in the auxiliary viewpoint coordinate system.
[0100] In some embodiments of this application, a sequence of main viewpoint video frames in the main viewpoint coordinate system can be converted into a sequence of reference viewpoint video frames in the auxiliary viewpoint coordinate system according to spatial correspondence.
[0101] In practical implementation, the reference viewpoint video frame sequence can be understood as a video frame sequence consisting of the main viewpoint video frame in the main viewpoint coordinate system and the projected video frame in the auxiliary viewpoint coordinate system, corresponding to the same time series.
[0102] S106: Encode the main viewpoint video frame sequence and the auxiliary viewpoint video frame sequence using preset encoding rules and reference viewpoint video frame sequences to obtain the encoding information of multi-viewpoint video.
[0103] In some embodiments of this application, a preset encoding rule and a reference viewpoint video frame sequence can be used to encode the main viewpoint video frame sequence and the auxiliary viewpoint video frame sequence to obtain the encoding information of the multi-viewpoint video. The preset encoding rule can be a multi-viewpoint encoding rule based on spatial relationship transformation.
[0104] In practical implementation, when performing inter-frame coding on the secondary viewpoint, the similarity between the secondary viewpoint video frame and the reference viewpoint video frame at the same time can be used for coding, avoiding the problem of difficulty in obtaining similar parts caused by large differences in camera positions.
[0105] In some embodiments of this application, the installation positions of the main viewpoint camera and the auxiliary viewpoint camera are obtained. A main viewpoint coordinate system is established based on the installation position of the main viewpoint camera, and an auxiliary viewpoint coordinate system is established based on the installation position of the auxiliary viewpoint camera. The spatial correspondence between the main viewpoint coordinate system and the auxiliary viewpoint coordinate system is determined based on the installation positions of the main viewpoint camera and the auxiliary viewpoint camera. The main viewpoint video frame sequence in the main viewpoint video and the auxiliary viewpoint video frame sequence in the auxiliary viewpoint video are obtained. Based on the spatial correspondence, the main viewpoint video frame sequence in the main viewpoint coordinate system is converted into a reference viewpoint video frame sequence in the auxiliary viewpoint coordinate system. The main viewpoint video frame sequence and the auxiliary viewpoint video frame sequence are encoded using a preset encoding rule and the reference viewpoint video frame sequence to obtain the encoding information of the multi-viewpoint video. This application first converts the main viewpoint video frame sequence into a reference viewpoint video frame sequence in an auxiliary viewpoint coordinate system. Then, it encodes both the main viewpoint and auxiliary viewpoint video frame sequences using preset encoding rules and the reference viewpoint video frame sequence. This solves the problem of unutilizing common content between viewpoints and improves encoding efficiency. It ensures that the encoded information for multi-viewpoint video can be generated efficiently and accurately.
[0106] In some embodiments of this application, the main viewpoint camera and the auxiliary viewpoint camera are set in a preset scene, and the spatial correspondence includes a rotational correspondence and / or a translational correspondence. Step S103 further includes the following sub-steps:
[0107] S11: Select at least one common feature point in the preset scene;
[0108] S12: Determine the main viewpoint feature coordinates of the common feature points in the main viewpoint coordinate system based on the spatial relationship between the installation position of the main viewpoint camera and the common feature points;
[0109] S13: Determine the auxiliary viewpoint feature coordinates of the common feature points in the auxiliary viewpoint coordinate system based on the spatial relationship between the installation position of the auxiliary viewpoint camera and the common feature points;
[0110] S14: Use the feature coordinates of the main viewpoint and the feature coordinates of the auxiliary viewpoint as a matching point pair of common feature points;
[0111] S15: Determine the rotational and / or translational correspondences between the main viewpoint coordinate system and the auxiliary viewpoint coordinate system based on the matching point pairs.
[0112] In some embodiments of this application, at least one common feature point can be selected in a preset scene. Then, the main viewpoint feature coordinates of the common feature point in the main viewpoint coordinate system are determined according to the spatial relationship between the installation position of the main viewpoint camera and the common feature point. Then, the auxiliary viewpoint feature coordinates of the common feature point in the auxiliary viewpoint coordinate system are determined according to the spatial relationship between the installation position of the auxiliary viewpoint camera and the common feature point.
[0113] Then, the feature coordinates of the main viewpoint and the feature coordinates of the auxiliary viewpoint can be used as matching point pairs of common feature points. Based on the matching point pairs, the rotational correspondence and / or translational correspondence between the main viewpoint coordinate system and the auxiliary viewpoint coordinate system can be determined.
[0114] In the specific implementation, the n viewpoints can be represented by View 0, View 1, ..., View n, and the corresponding camera positions are P0, P1, ..., Pn, respectively. n .
[0115] Let View 0 be the primary viewpoint, and the other viewpoints be secondary viewpoints. Based on the spatial position of the cameras, obtain the rotation matrix R between the primary viewpoint camera and each secondary viewpoint camera through camera calibration and other methods. 0,i and T 0,i Translation matrix, R 0,i ,T 0,i =Camera_Calibration(P0,P i ), where i represents View i, i = 1, 2, ..., n.
[0116] This application obtains the rotation and translation matrices between the main viewpoint camera and each auxiliary viewpoint camera through camera calibration and other methods, ensuring that video frames can be accurately converted and encoded, thus improving the efficiency and quality of multi-view video encoding. By selecting common feature points and determining their feature coordinates in different coordinate systems, the accuracy and reliability of matching point pairs are ensured, improving the stability and precision of the system. By determining the rotation and translation correspondences, the accuracy of video frames can be accurately converted and encoded, reducing the transmission of redundant information and improving encoding efficiency and bandwidth utilization.
[0117] In some embodiments of this application, step S105 further includes the following sub-steps:
[0118] S21: Obtain at least one main viewpoint video frame from the main viewpoint video frame sequence in the main viewpoint coordinate system;
[0119] S22: Convert the main viewpoint video frame into a main viewpoint rotated video frame according to the rotation correspondence;
[0120] S23: Convert the main viewpoint rotated video frame into a reference viewpoint video frame in the auxiliary viewpoint coordinate system according to the translation correspondence;
[0121] S24: Combine at least one reference viewpoint video frame in chronological order to form a reference viewpoint video frame sequence in the auxiliary viewpoint coordinate system.
[0122] In some embodiments of this application, at least one main viewpoint video frame can be obtained from the main viewpoint video frame sequence in the main viewpoint coordinate system. The main viewpoint video frame is then converted into a main viewpoint rotated video frame according to a rotation correspondence. The main viewpoint rotated video frame is then converted into a reference viewpoint video frame in the auxiliary viewpoint coordinate system according to a translation correspondence. At least one reference viewpoint video frame is then combined in chronological order to form a reference viewpoint video frame sequence in the auxiliary viewpoint coordinate system. The rotation correspondence can be a rotation matrix, and the translation correspondence can be a translation matrix.
[0123] In the specific implementation, it can be based on R 0,i and T 0,i Frame M captured from the main viewpoint at time t t The video frames are switched to various secondary viewpoints, with the frames switched to View i using... It means that among them Where i represents View i, i = 1, 2, ..., n.
[0124] In specific implementations, spatial relationship transformation can be performed with reference to Figure 4. Figure 4 shows a schematic diagram of the spatial relationship transformation process provided in some embodiments of this application. When using inter-viewpoint relationships for inter-frame coding, let the video frame of View i at time t be... The traditional approach in related technologies is to utilize the main viewpoint frame M. t and auxiliary viewpoint frames The inter-frame relationship between them. However, due to the large difference in camera position, the main viewpoint frame M... t and auxiliary viewpoint frames Some identical content in a video frame may be located far apart in relative positions. Furthermore, in practical applications, due to real-time requirements, complex full-search algorithms are generally not used. Therefore, under the limitation of the search range, motion estimation may fall into local optima, and inter-frame prediction cannot fully utilize the common information between different viewpoints. However, by using the main viewpoint frame M... t Convert to reference viewpoint Then, the main viewpoint frame M is displayed. t Feature reference view frame and The same content is located relatively close to each other in two video frames, which allows for full utilization of common information between viewpoints and improves coding efficiency.
[0125] This application converts the main viewpoint video frame into a reference viewpoint video frame through rotation and translation correspondences, ensuring accurate conversion and encoding of the video frame and improving the efficiency and quality of multi-viewpoint video coding. By converting the main viewpoint frame into a reference viewpoint frame, it ensures that the relative positional distance of the same content in the two video frames is small, making full use of the common information between viewpoints and improving coding efficiency and prediction accuracy.
[0126] In some embodiments of this application, the preset coding rules include inter-frame coding rules and inter-viewpoint coding rules, and step S106 further includes the following sub-steps:
[0127] S31: Encode the main viewpoint video frame sequence using inter-frame coding rules to obtain the main viewpoint video coding information;
[0128] S32: Encode the auxiliary viewpoint video frame sequence using inter-frame coding rules, inter-viewpoint coding rules, and reference viewpoint video frame sequence to obtain auxiliary viewpoint video coding information;
[0129] S33: Determine the encoding information of the multi-view video based on the main viewpoint video encoding information and the auxiliary viewpoint video encoding information.
[0130] In some embodiments of this application, the main viewpoint video frame sequence can be encoded using inter-frame coding rules to obtain main viewpoint video coding information. The auxiliary viewpoint video frame sequence can be encoded using inter-frame coding rules, inter-viewpoint coding rules, and reference viewpoint video frame sequence to obtain auxiliary viewpoint video coding information. The coding information of the multi-viewpoint video is determined based on the main viewpoint video coding information and the auxiliary viewpoint video coding information.
[0131] In a specific implementation, multi-view coding based on spatial relationship transformation can be performed with reference to Figure 5. Figure 5 shows a schematic diagram of the framework for multi-view coding based on spatial relationship transformation provided in some embodiments of this application. In Figure 5, the main view corresponding to the main viewpoint View 0 can be encoded using HEVC, considering only the inter-frame correlation on the time axis. Similar to MV-HEVC, the auxiliary views corresponding to the secondary viewpoints View 1 and View 2 consider both inter-viewpoint and temporal correlations during inter-frame coding. When utilizing inter-viewpoint relationships, this application first transforms each frame in the main viewpoint View 0 to the corresponding viewpoint based on the rotation and translation matrices between different viewpoints; then, it uses the similarity between the transformed reference viewpoint frame and the auxiliary viewpoint frame to perform inter-viewpoint coding.
[0132] In some embodiments of this application, step S32 further includes the following sub-steps:
[0133] S41: The auxiliary viewpoint video frame sequence is encoded using inter-frame coding rules to obtain the first coding information of the auxiliary viewpoint;
[0134] S42: The auxiliary viewpoint video frame sequence is encoded using the inter-viewpoint coding rules and the reference viewpoint video frame sequence to obtain the auxiliary viewpoint second coding information;
[0135] S43: Determine the auxiliary viewpoint video coding information based on the auxiliary viewpoint first coding information and the auxiliary viewpoint second coding information.
[0136] In some embodiments of this application, the auxiliary viewpoint video frame sequence can be encoded using inter-frame coding rules to obtain the first auxiliary viewpoint coding information. Then, the auxiliary viewpoint video frame sequence can be encoded using inter-viewpoint coding rules and the reference viewpoint video frame sequence to obtain the second auxiliary viewpoint coding information. Finally, the auxiliary viewpoint video coding information can be determined based on the first auxiliary viewpoint coding information and the second auxiliary viewpoint coding information.
[0137] In the specific implementation, referring to Figure 5, taking the B-frame at time T2 of viewpoint 2 as an example, the B-frame at time T2 of viewpoint 2 utilizes the P-frame at time T0 and the B-frame at time T4 of viewpoint 2 on the time axis. Between viewpoints, it utilizes the B2 frame, which represents the spatial correspondence between viewpoints, transitioning from viewpoint 0 to viewpoint 2, where B2 = R. 0,2 ×B+T 0,2 .
[0138] In some embodiments of this application, the main viewpoint video and the auxiliary viewpoint video are acquired within a preset time period. After converting the main viewpoint video frame sequence in the main viewpoint coordinate system into a reference viewpoint video frame sequence in the auxiliary viewpoint coordinate system according to the spatial correspondence, the method further includes the following steps:
[0139] S51: Obtain any reference viewpoint video frame from the reference viewpoint video frame sequence;
[0140] S52: Determine the region to be detected in the reference viewpoint video frame based on spatial correspondence;
[0141] S53: Set a sliding detection window in the area to be detected; the size of the sliding detection window is determined by the distance between the installation positions of the main viewpoint camera and the auxiliary viewpoint camera; the sliding detection window includes several pixels to be detected in the reference viewpoint video frame;
[0142] S54: During the sliding detection window's movement in the detection area, obtain the grayscale value of the pixel to be detected in the sliding detection window and the median grayscale value of several pixels to be detected in the sliding detection window.
[0143] S55: Determine the grayscale difference between the grayscale value of any pixel to be detected in the sliding detection window and the median grayscale value;
[0144] S56: If the grayscale difference is greater than the preset threshold, the pixel corresponding to the median grayscale value is used to replace the pixel to be detected.
[0145] S57: After traversing the area to be detected by the sliding detection window, the reference viewpoint video frames formed by replacing the pixels to be detected will be used to form a new reference viewpoint video frame sequence.
[0146] In some embodiments of this application, any reference viewpoint video frame in a reference viewpoint video frame sequence is obtained, and the region to be detected in the reference viewpoint video frame is determined according to a spatial correspondence. The spatial correspondence can be the rotation correspondence and translation correspondence described above.
[0147] Then, a sliding detection window can be set in the area to be detected. The size of the sliding detection window is determined by the distance between the installation positions of the main viewpoint camera and the auxiliary viewpoint camera. Specifically, the greater the distance between the installation positions of the main viewpoint camera and the auxiliary viewpoint camera, the larger the size of the sliding detection window; conversely, the smaller the distance, the smaller the size of the sliding detection window.
[0148] The sliding detection window includes several pixels to be detected in the reference viewpoint video frame. As the sliding detection window slides in the area to be detected, the grayscale value of the pixel to be detected in the sliding detection window and the median grayscale value of several pixels to be detected in the sliding detection window can be obtained. The grayscale difference between the grayscale value of any pixel to be detected in the sliding detection window and the median grayscale value can be determined. If the grayscale difference is greater than a preset threshold, the pixel to be detected is replaced by the pixel corresponding to the median grayscale value.
[0149] After the sliding detection window traverses the area to be detected, the reference viewpoint video frames formed by replacing the pixels to be detected can be combined to form a new reference viewpoint video frame sequence.
[0150] In specific implementations, the angles of cameras from different viewpoints can be obtained by referring to Figure 6. Figure 6 shows a schematic diagram of the shooting angles of cameras from different viewpoints provided in some embodiments of this application. Since the shooting angles of cameras from different viewpoints are different, the content captured by each viewpoint is different. For example, the camera of the secondary viewpoint View 2 is placed on the right side of the target object, while the camera of the primary viewpoint View 0 is placed in the middle of the target object. The video frames under View 2 can capture more information about the right side of the target object, while the video frames under View 0 may lose some detailed information about the right side of the target object. Therefore, when the video frames of the primary viewpoint View 0 are switched to the secondary viewpoint View 2, there may be a loss of detailed information.
[0151] In a specific implementation, a schematic diagram of missing detailed information can be obtained by referring to Figure 7. Figure 7 shows a schematic diagram of the loss of video frame detailed information provided by some embodiments of this application. The area within the black frame in Figure 7 shows that the detailed information in the image is missing.
[0152] In the specific implementation, in the converted main viewpoint video frame, the grayscale values of the parts with lost detail information differ significantly from the grayscale values of their neighboring pixels. Furthermore, when the main viewpoint video frame is converted to the right, the lost information tends to concentrate on the right side, and vice versa. Therefore, this patent proposes an adaptive region sliding window detail loss detection method when detecting the lost information. First, candidate regions are calculated based on the rotation and translation matrices, i.e., Area = computeArea(R... 0,i ,T 0,i This can be understood as using the right half of the video frame as the adaptive region if the auxiliary viewpoint is to the right of the main viewpoint. By selecting candidate regions, we can avoid searching for lost details across the entire image frame, thus improving retrieval efficiency.
[0153] Then, a k×k sliding window is set up, which slides across the candidate region to traverse the entire region. During the sliding process, the median gray value of each pixel is compared with the median gray value of the pixels within the window. If the difference between a pixel and the median gray value of the pixels within the window is greater than a set threshold, the pixel is considered a missing part, and the pixel with the median value is used to fill the missing part. By filling in the details, the missing parts of the converted image can be replaced by the median of adjacent pixels, which is closer to the pixels at the corresponding positions of the secondary viewpoints. This allows full utilization of the redundancy between frames from different viewpoints, improving coding efficiency.
[0154] In some embodiments of this application, step S54 includes the following sub-steps:
[0155] S61: During the sliding detection window's movement across the detection area, the grayscale values of several pixels to be detected in the sliding detection window are acquired in real time.
[0156] S62: Sort the gray values of several pixels to be detected in ascending order;
[0157] S63: Use the median of the sorted gray values as the median gray value of several pixels to be detected.
[0158] In some embodiments of this application, as the sliding detection window slides in the area to be detected, the gray values of several pixels to be detected in the sliding detection window can be acquired in real time, the gray values of several pixels to be detected can be sorted in ascending order, and the median of the sorted gray values can be used as the median of the gray values of several pixels to be detected.
[0159] In the implementation, the grayscale values of all pixels within each window can be extracted. These extracted grayscale values are then sorted. The median of the sorted grayscale values is calculated based on the window size. For odd-sized windows, the median is the middle value after sorting; for even-sized windows, the median can be the average of the two middle values.
[0160] Referring to Figure 8, a schematic diagram of the framework of a multi-view video coding method combining spatial relationship transformation and detail filling provided in some embodiments of this application is shown. The specific details are as follows:
[0161] The multi-view video coding method combining spatial relationship transformation and detail filling mainly consists of two parts: the first part is spatial relationship transformation, which first obtains the spatial relationship between cameras and then performs viewpoint transformation based on the spatial relationship; the second part is adaptive sliding window detail filling, which sets a sliding window for the transformed video frame and adapts the sliding area based on the spatial relationship to improve the efficiency of detail filling.
[0162] Referring to Figure 9, a schematic diagram of a multi-view video encoding device provided in some embodiments of this application is shown. The multi-view video includes a main viewpoint video captured by a main viewpoint camera and an auxiliary viewpoint video captured by an auxiliary viewpoint camera. The device includes:
[0163] The installation location acquisition module 901 is used to acquire the installation location of the main viewpoint camera and the installation location of the auxiliary viewpoint camera.
[0164] The coordinate system establishment module 902 is used to establish a main viewpoint coordinate system based on the installation position of the main viewpoint camera and an auxiliary viewpoint coordinate system based on the installation position of the auxiliary viewpoint camera.
[0165] The spatial correspondence determination module 903 is used to determine the spatial correspondence between the main viewpoint coordinate system and the auxiliary viewpoint coordinate system based on the installation positions of the main viewpoint camera and the auxiliary viewpoint camera.
[0166] The video frame sequence acquisition module 904 is used to acquire the main viewpoint video frame sequence in the main viewpoint video and the auxiliary viewpoint video frame sequence in the auxiliary viewpoint video;
[0167] The spatial transformation module 905 is used to convert the main viewpoint video frame sequence in the main viewpoint coordinate system into a reference viewpoint video frame sequence in the auxiliary viewpoint coordinate system according to the spatial correspondence.
[0168] The encoding module 906 is used to encode the main viewpoint video frame sequence and the auxiliary viewpoint video frame sequence using preset encoding rules and reference viewpoint video frame sequences to obtain the encoding information of multi-viewpoint video.
[0169] In some embodiments, a main viewpoint camera and an auxiliary viewpoint camera are positioned in a preset scene, and the spatial correspondence includes a rotational correspondence and / or a translational correspondence. The spatial correspondence determination module 903 includes:
[0170] The common feature point acquisition module is used to select at least one common feature point in a preset scene;
[0171] The feature point main view coordinate acquisition module is used to determine the main view feature coordinates of the common feature points in the main view coordinate system based on the spatial relationship between the installation position of the main view camera and the common feature points.
[0172] The auxiliary viewpoint coordinate acquisition module is used to determine the auxiliary viewpoint feature coordinates of the common feature points in the auxiliary viewpoint coordinate system based on the spatial relationship between the installation position of the auxiliary viewpoint camera and the common feature points.
[0173] The matching point pair acquisition module is used to use the feature coordinates of the main viewpoint and the feature coordinates of the auxiliary viewpoint as matching point pairs of common feature points;
[0174] The rotation and translation relationship acquisition module is used to determine the rotational and / or translational correspondence between the main viewpoint coordinate system and the auxiliary viewpoint coordinate system based on the matching point pairs.
[0175] In some embodiments, the video frame sequence acquisition module 904 includes:
[0176] The main viewpoint video frame acquisition module is used to acquire at least one main viewpoint video frame from the main viewpoint video frame sequence in the main viewpoint coordinate system.
[0177] The main viewpoint video frame rotation module is used to convert the main viewpoint video frame into a main viewpoint rotated video frame according to the rotation correspondence.
[0178] The main viewpoint video frame translation module is used to convert the main viewpoint rotated video frame into a reference viewpoint video frame in the auxiliary viewpoint coordinate system according to the translation correspondence.
[0179] The video frame sequence acquisition submodule is used to combine at least one reference viewpoint video frame in chronological order to form a reference viewpoint video frame sequence in the auxiliary viewpoint coordinate system.
[0180] In some embodiments, the preset encoding rules include inter-frame encoding rules and inter-viewpoint encoding rules. The encoding module 906 includes:
[0181] The main encoding information acquisition module is used to encode the main viewpoint video frame sequence using inter-frame coding rules to obtain the main viewpoint video encoding information;
[0182] The auxiliary coding information acquisition module is used to encode the auxiliary viewpoint video frame sequence using inter-frame coding rules, inter-viewpoint coding rules, and reference viewpoint video frame sequence to obtain auxiliary viewpoint video coding information.
[0183] The video encoding information acquisition module is used to determine the encoding information of multi-view video based on the main viewpoint video encoding information and the auxiliary viewpoint video encoding information.
[0184] In some embodiments, the auxiliary encoding information acquisition module includes:
[0185] The first encoding information acquisition module is used to encode the auxiliary viewpoint video frame sequence using inter-frame coding rules to obtain the first encoding information of the auxiliary viewpoint.
[0186] The second encoding information acquisition module is used to encode the auxiliary viewpoint video frame sequence using inter-viewpoint encoding rules and a reference viewpoint video frame sequence to obtain the auxiliary viewpoint second encoding information.
[0187] The auxiliary encoding information acquisition submodule is used to determine the auxiliary viewpoint video encoding information based on the auxiliary viewpoint first encoding information and the auxiliary viewpoint second encoding information.
[0188] In some embodiments, the main viewpoint video and the secondary viewpoint video are captured within a preset time period. After converting the main viewpoint video frame sequence in the main viewpoint coordinate system into a reference viewpoint video frame sequence in the secondary viewpoint coordinate system according to spatial correspondence, the apparatus further includes:
[0189] The reference viewpoint video frame acquisition module is used to acquire any reference viewpoint video frame in the reference viewpoint video frame sequence.
[0190] The region to be detected module is used to determine the region to be detected in the reference viewpoint video frame based on spatial correspondence.
[0191] The sliding window acquisition module is used to set a sliding detection window in the area to be detected; the size of the sliding detection window is determined by the distance between the installation positions of the main viewpoint camera and the auxiliary viewpoint camera; the sliding detection window includes several pixels to be detected in the reference viewpoint video frame;
[0192] The grayscale value acquisition module is used to acquire the grayscale value of the pixel to be detected in the sliding detection window and the median grayscale value of several pixels to be detected in the sliding detection window during the sliding detection window's movement in the detection area.
[0193] The grayscale comparison module is used to determine the grayscale difference between the grayscale value of any pixel to be detected in the sliding detection window and the median grayscale value.
[0194] The pixel replacement module is used to replace the pixel to be detected with the pixel corresponding to the median gray value if the gray-level difference is greater than a preset threshold.
[0195] The reference viewpoint video frame optimization module is used to assemble a new reference viewpoint video frame sequence by replacing the pixels to be detected with the reference viewpoint video frames formed after the sliding detection window traverses the area to be detected.
[0196] In some embodiments, the grayscale value acquisition module includes:
[0197] The grayscale value acquisition submodule is used to acquire the grayscale values of several pixels to be detected in the sliding detection window in real time as the sliding detection window slides in the area to be detected.
[0198] The grayscale value sorting module is used to sort the grayscale values of several pixels to be detected in ascending order.
[0199] The grayscale median acquisition module is used to take the median of sorted grayscale values as the grayscale median of several pixels to be detected.
[0200] Some embodiments of this application also provide an electronic device, characterized in that it includes a processor, a memory, and a computer program stored in the memory and capable of running on the processor, wherein the computer program, when executed by the processor, implements the multi-view video encoding method described above.
[0201] Some embodiments of this application also provide a computer-readable storage medium, characterized in that a computer program is stored on the computer-readable storage medium, and when the computer program is executed by a processor, it implements the multi-view video encoding method described above.
[0202] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties. Furthermore, the collection, use and processing of the relevant data must comply with the relevant laws, regulations and standards of the relevant countries and regions, and corresponding operation portals are provided for users to choose to authorize or refuse.
[0203] Some embodiments in this specification are described in a progressive manner, and some embodiments focus on the differences from other embodiments. For the same or similar parts between some embodiments, please refer to each other.
[0204] Those skilled in the art will understand that some embodiments of this application can be provided as methods, apparatus, or computer program products. Therefore, some embodiments of this application can take the form of entirely hardware embodiments, entirely software embodiments, or embodiments combining software and hardware aspects. Furthermore, some embodiments of this application can take the form of computer program products implemented on one or more computer-usable storage media containing computer-usable program code, including but not limited to disk storage, CD-ROM (Compact Disc Read-Only Memory), optical storage, etc.
[0205] Some embodiments of this application are described with reference to flowchart illustrations and / or block diagrams of methods, terminal devices (systems), and computer program products according to some embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing terminal device to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing terminal device, create means for implementing the functions specified in one or more blocks of the flowchart illustrations and / or one or more blocks of the block diagrams.
[0206] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing terminal device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means that implement the functions specified in one or more flowcharts and / or one or more block diagrams.
[0207] These computer program instructions may also be loaded onto a computer or other programmable data processing terminal equipment to cause a series of operational steps to be performed on the computer or other programmable terminal equipment to produce a computer-implemented process, such that the instructions, which execute on the computer or other programmable terminal equipment, provide steps for implementing the functions specified in one or more flowcharts and / or one or more block diagrams.
[0208] Although some embodiments of this application have been described, those skilled in the art, upon learning the basic inventive concept, can make further changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including some embodiments as well as all changes and modifications falling within the scope of some embodiments of this application.
[0209] Finally, it should be noted that in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or terminal device that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or terminal device. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or terminal device that includes the aforementioned element.
[0210] The above provides a detailed description of a multi-view video encoding method, apparatus, electronic device, and medium. Specific examples have been used to illustrate the principles and implementation methods of this application. The descriptions of some embodiments above are only for the purpose of helping to understand the method and core ideas of this application. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of this application. Therefore, the content of this specification should not be construed as a limitation of this application.
Claims
1. A method for encoding multi-view video, characterized in that, The multi-viewpoint video includes a main-viewpoint video captured by a main-viewpoint camera and an auxiliary-viewpoint video captured by an auxiliary-viewpoint camera. The method includes: Obtain the installation position of the main viewpoint camera and the installation position of the auxiliary viewpoint camera; A main viewpoint coordinate system is established based on the installation position of the main viewpoint camera, and an auxiliary viewpoint coordinate system is established based on the installation position of the auxiliary viewpoint camera. The spatial correspondence between the main viewpoint coordinate system and the auxiliary viewpoint coordinate system is determined based on the installation positions of the main viewpoint camera and the auxiliary viewpoint camera. Obtain the main viewpoint video frame sequence in the main viewpoint video and the auxiliary viewpoint video frame sequence in the auxiliary viewpoint video; Based on the spatial correspondence, the main viewpoint video frame sequence in the main viewpoint coordinate system is converted into a reference viewpoint video frame sequence in the auxiliary viewpoint coordinate system; The main viewpoint video frame sequence and the auxiliary viewpoint video frame sequence are encoded using preset encoding rules and the reference viewpoint video frame sequence to obtain the encoding information of the multi-viewpoint video.
2. The method according to claim 1, characterized in that, The main viewpoint camera and the auxiliary viewpoint camera are set in a preset scene. The spatial correspondence includes a rotational correspondence and / or a translational correspondence. Determining the spatial correspondence between the main viewpoint coordinate system and the auxiliary viewpoint coordinate system based on the installation positions of the main viewpoint camera and the auxiliary viewpoint camera includes: Select at least one common feature point in the preset scene; The main viewpoint feature coordinates of the common feature points in the main viewpoint coordinate system are determined based on the spatial relationship between the installation position of the main viewpoint camera and the common feature points. The auxiliary viewpoint feature coordinates of the common feature points in the auxiliary viewpoint coordinate system are determined based on the spatial relationship between the installation position of the auxiliary viewpoint camera and the common feature points. The main viewpoint feature coordinates and the auxiliary viewpoint feature coordinates are used as the matching point pairs of the common feature points; The rotational and / or translational correspondences between the main viewpoint coordinate system and the auxiliary viewpoint coordinate system are determined based on the matching point pairs.
3. The method according to claim 2, characterized in that, The step of converting the main viewpoint video frame sequence in the main viewpoint coordinate system into a reference viewpoint video frame sequence in the auxiliary viewpoint coordinate system according to the spatial correspondence includes: At least one main viewpoint video frame is obtained from the main viewpoint video frame sequence in the main viewpoint coordinate system; The main viewpoint video frame is converted into a main viewpoint rotated video frame according to the rotation correspondence. Based on the translation correspondence, the main viewpoint rotated video frame is converted into a reference viewpoint video frame in the auxiliary viewpoint coordinate system; At least one of the reference viewpoint video frames is combined in chronological order to form the reference viewpoint video frame sequence in the auxiliary viewpoint coordinate system.
4. The method according to claim 3, characterized in that, The preset encoding rules include inter-frame encoding rules and inter-viewpoint encoding rules. The step of encoding the main viewpoint video frame sequence and the auxiliary viewpoint video frame sequence using the preset encoding rules and the reference viewpoint video frame sequence to obtain the encoding information of the multi-viewpoint video includes: The main viewpoint video frame sequence is encoded using the aforementioned inter-frame coding rules to obtain main viewpoint video coding information; The auxiliary viewpoint video frame sequence is encoded using the inter-frame coding rules, the inter-viewpoint coding rules, and the reference viewpoint video frame sequence to obtain auxiliary viewpoint video coding information. The encoding information of the multi-view video is determined based on the main viewpoint video encoding information and the auxiliary viewpoint video encoding information.
5. The method according to claim 4, characterized in that, The step of encoding the auxiliary viewpoint video frame sequence using the inter-frame coding rules, the inter-viewpoint coding rules, and the reference viewpoint video frame sequence to obtain auxiliary viewpoint video coding information includes: The auxiliary viewpoint video frame sequence is encoded using the inter-frame coding rules to obtain the first coding information of the auxiliary viewpoint; The auxiliary viewpoint video frame sequence is encoded using the inter-viewpoint coding rules and the reference viewpoint video frame sequence to obtain the second coding information of the auxiliary viewpoint; The auxiliary viewpoint video encoding information is determined based on the first encoding information and the second encoding information of the auxiliary viewpoint.
6. The method according to claim 1, characterized in that, The main viewpoint video and the auxiliary viewpoint video are captured within a preset time period. After converting the main viewpoint video frame sequence in the main viewpoint coordinate system into a reference viewpoint video frame sequence in the auxiliary viewpoint coordinate system according to the spatial correspondence, the method further includes: Obtain any one of the reference viewpoint video frames in the reference viewpoint video frame sequence; The region to be detected in the reference viewpoint video frame is determined based on the spatial correspondence; A sliding detection window is set in the area to be detected; the size of the sliding detection window is determined by the distance between the installation positions of the main viewpoint camera and the auxiliary viewpoint camera; the sliding detection window includes several pixels to be detected in the reference viewpoint video frame; During the process of the sliding detection window sliding in the area to be detected, the grayscale value of the pixel to be detected in the sliding detection window and the median of the grayscale values of a plurality of pixels to be detected in the sliding detection window are obtained. Determine the grayscale difference between the grayscale value of any pixel to be detected in the sliding detection window and the median grayscale value; If the grayscale difference is greater than a preset threshold, the pixel to be detected is replaced by the pixel corresponding to the median grayscale value. After the sliding detection window traverses the region to be detected, the reference viewpoint video frames formed by replacing the pixels to be detected are used to form a new reference viewpoint video frame sequence.
7. The method according to claim 6, characterized in that, During the process of the sliding detection window sliding in the area to be detected, acquiring the grayscale value of the pixel to be detected in the sliding detection window and the median of the grayscale values of a plurality of pixels to be detected in the sliding detection window includes: During the sliding detection window's movement across the detection area, the grayscale values of several pixels to be detected within the sliding detection window are acquired in real time. The gray values of the pixels to be detected are sorted in ascending order; The median of the sorted grayscale values is used as the median grayscale value of the pixels to be detected.
8. A multi-view video encoding device, characterized in that, The multi-viewpoint video includes a main-viewpoint video captured by a main-viewpoint camera and an auxiliary-viewpoint video captured by an auxiliary-viewpoint camera. The device includes: The installation location acquisition module is used to acquire the installation location of the main viewpoint camera and the installation location of the auxiliary viewpoint camera; The coordinate system establishment module is used to establish a main viewpoint coordinate system based on the installation position of the main viewpoint camera and to establish an auxiliary viewpoint coordinate system based on the installation position of the auxiliary viewpoint camera. The spatial correspondence determination module is used to determine the spatial correspondence between the main viewpoint coordinate system and the auxiliary viewpoint coordinate system based on the installation positions of the main viewpoint camera and the auxiliary viewpoint camera. The video frame sequence acquisition module is used to acquire the main viewpoint video frame sequence in the main viewpoint video and the auxiliary viewpoint video frame sequence in the auxiliary viewpoint video; A spatial transformation module is used to convert the main viewpoint video frame sequence in the main viewpoint coordinate system into a reference viewpoint video frame sequence in the auxiliary viewpoint coordinate system according to the spatial correspondence. The encoding module is used to encode the main viewpoint video frame sequence and the auxiliary viewpoint video frame sequence using preset encoding rules and the reference viewpoint video frame sequence to obtain the encoding information of the multi-viewpoint video.
9. An electronic device, characterized in that, It includes a processor, a memory, and a computer program stored in the memory and capable of running on the processor, wherein the computer program, when executed by the processor, implements the encoding method for multi-view video as described in any one of claims 1-7.
10. A computer-readable storage medium, characterized in that, A computer program is stored on the computer-readable storage medium, which, when executed by a processor, implements the encoding method for multi-view video as described in any one of claims 1-7.