Image coding method and device, storage medium and computer program product
By acquiring multiple depth maps and encoding them based on viewpoint and pixel information, the problem of low depth map encoding efficiency in existing technologies is solved, achieving more efficient and accurate encoding results and improving the quality of 3D full-scale video.
Patent Information
- Application Number
- CN202511462881.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-13
- Publication Date
- 2026-02-06
AI Technical Summary
In existing technologies, the efficiency and accuracy of determining reference depth maps when encoding depth maps are poor, and the spatial correlation between depth maps from different viewpoints at the same time is ignored.
By acquiring multiple depth maps at different times, a target depth map and a reference depth map are determined based on the acquisition viewpoint of each depth map. The pixel information of the reference depth map is then used for encoding processing, including determining motion vectors and matrix operations to improve encoding accuracy.
It improves the efficiency and accuracy of depth map encoding, solves the problem of low efficiency in existing technologies, and ensures the quality of generated 3D full-scale video.
Smart Images

Figure CN121486564A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of image processing, and in particular to an image encoding method, device, storage medium and computer program product. BACKGROUND
[0002] In generating a three dimensions (3D) real video, it is usually required to encode a plurality of texture maps and depth maps of different time instants of multiple perspectives collected, and send the encoded texture maps and depth maps to a video generation device for processing to generate the 3D real video. At present, in the related art, when encoding a depth map of a secondary perspective at a time instant, a reference depth map with the same perspective as the depth map is determined from a depth map collected at a previous time instant before the time instant, and the depth map is encoded and processed according to related information of the reference depth map. However, this way of determining the reference depth map is extremely inefficient, and this way also ignores the spatial correlation of depth maps of different perspectives at the same time instant, which results in inaccurate determination of the reference depth map. SUMMARY
[0003] To solve the above technical problems, the embodiments of the present application provide an image encoding method, device, storage medium and computer program product, which solve the problem of low efficiency and low accuracy in determining a reference depth map when encoding a depth map in the related art.
[0004] To achieve the above object, the technical scheme of the embodiments of the present application is as follows: An image encoding method, the method comprising: obtaining a plurality of depth map sets of different time instants for a target scene; wherein different depth maps in each depth map set correspond to different collection perspectives; determining a target depth map and a reference depth map corresponding to the target depth map from each depth map set based on the collection perspective of each depth map in each depth map set; encoding and processing the target depth map based on pixel information of the reference depth map to obtain a target encoded map.
[0005] In the above scheme, the target depth map and the reference depth map corresponding to the target depth map are determined from each depth map set based on the collection perspective of each depth map in each depth map set, comprising: determining a depth map with the collection perspective as a first perspective in each depth map set as the target depth map; If the number of the target depth maps is one, determining, from each of the depth map sets, a depth map with the second view angle as the reference depth map of the target depth map; wherein the second view angle is a view angle other than the first view angle among the plurality of view angles.
[0006] In the above scheme, the method further comprises: If the number of the target depth maps is a plurality, sorting the plurality of target depth maps according to a target sorting strategy; From each of the depth map sets, determining a reference depth map corresponding to a target depth map with a first sorting order, as a depth map with the second view angle; For a target depth map with an i-th sorting order, determining a candidate depth map adjacent to the target depth map with the i-th sorting order from each of the depth map sets; wherein i is an integer greater than 1; From the candidate depth map, determining a reference depth map corresponding to the target depth map with the i-th sorting order.
[0007] In the above scheme, the encoding processing of the target depth map based on the pixel information of the reference depth map comprises: Determining a motion vector corresponding to the target depth map based on the pixel information; Encoding processing the target depth map based on the motion vector to obtain the target encoded map.
[0008] In the above scheme, the determining of the motion vector corresponding to the target depth map based on the pixel information comprises: Obtaining first pose information of a first image acquisition device corresponding to the target depth map; Obtaining attribute information and second pose information of a second image acquisition device corresponding to the reference depth map of the target depth map; Determining the motion vector based on the pixel information, the attribute information, the first pose information and the second pose information.
[0009] In the above scheme, the determining of the motion vector based on the pixel information, the attribute information, the first pose information and the second pose information comprises: Determining a first matrix and a second matrix based on the first pose information and the second pose information; Constructing a third matrix based on the attribute information; Determining a fourth matrix based on the first matrix, the third matrix and the pixel information; Determining the motion vector based on the fourth matrix and the second matrix.
[0010] In the above scheme, determining the first matrix and the second matrix based on the first pose information and the second pose information includes: The first matrix is determined based on the first rotation matrix and the second rotation matrix; The second matrix is determined based on the first translation matrix, the second translation matrix, and the first rotation matrix; wherein the first pose information includes the first rotation matrix and the first translation matrix; and the second pose information includes the second rotation matrix and the second translation matrix.
[0011] An image encoding device includes: a processor and a memory for storing a computer program capable of running on the processor; The processor is used to execute the steps of the above method when running a computer program.
[0012] A storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the above method.
[0013] A computer program product includes a computer program that, when executed by a processor, implements the steps of the above-described method.
[0014] The image encoding method, device, storage medium, and computer program product provided in this application embodiment can acquire multiple depth maps at different times for a target scene; different depth maps in each depth map set correspond to different acquisition perspectives; based on the acquisition perspective of each depth map in each depth map set, a target depth map and a reference depth map corresponding to the target depth map are determined from each depth map set; based on the pixel information of the reference depth map, the target depth map is encoded to obtain a target encoded map; thus, for a depth map set at a certain time, the target depth map can be determined from the depth map set based on the acquisition perspective of each depth map in the depth map set. The target depth map (i.e., the sub-view depth map) is encoded, and then a reference depth map corresponding to the target depth map is determined from the depth map set at that moment. Then, the target depth map is encoded based on the pixel information of the reference depth map. That is, when encoding the target depth map at a certain moment, the corresponding reference depth map is directly determined from the depth map set at that moment, instead of determining the reference depth map from the depth map acquired at the previous moment as in related technologies. This solves the problem of poor efficiency and low accuracy in determining the reference depth map when encoding depth maps in related technologies, thereby improving the encoding efficiency of the target depth map. Attached Figure Description
[0015] Figure 1 A flowchart illustrating an image encoding method provided in an embodiment of this application; Figure 2This is a schematic diagram illustrating the determination of a reference depth map in an image encoding method provided in an embodiment of this application; Figure 3 A schematic diagram of the structure of an image encoding device provided for an embodiment of this application; Figure 4 This is a schematic diagram of the structure of an image encoding device provided for an embodiment of this application. Detailed Implementation
[0016] The technical solutions in the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings.
[0017] It should be understood that the phrases "embodiments of this application" or "foreign embodiments" throughout the specification mean that a specific feature, structure, or characteristic related to an embodiment is included in at least one embodiment of this application. Therefore, "embodiments of this application" or "in the foreign embodiments" appearing throughout the specification do not necessarily refer to the same embodiment. Furthermore, these specific features, structures, or characteristics can be combined in any suitable manner in one or more embodiments. In the various embodiments of this application, the sequence numbers of the above-described processes do not imply a sequential order of execution; the execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application. The sequence numbers of the above-described embodiments are merely descriptive and do not represent the superiority or inferiority of the embodiments.
[0018] Unless otherwise specified, any step in the embodiments of this application performed by the electronic device may be executed by the processor of the electronic device. It is also worth noting that the embodiments of this application do not limit the order in which the electronic device performs the following steps. Furthermore, the methods used to process data in different embodiments may be the same or different methods. It should also be noted that any step in the embodiments of this application can be executed independently by the electronic device; that is, when the electronic device performs any step in the following embodiments, it may not depend on the execution of other steps.
[0019] It should be understood that the specific embodiments described herein are for illustrative purposes only and are not intended to limit the scope of this application.
[0020] This application provides an image encoding method, referring to... Figure 1 As shown, the method may include the following steps: Step 101: Obtain multiple depth maps at different times for the target scene.
[0021] In each depth map set, different depth maps correspond to different acquisition perspectives.
[0022] In the embodiments of this application, each moment corresponds to a depth map set; each depth map set includes at least two depth maps; the acquisition viewpoint may refer to the acquisition angle of the image acquisition device; for the depth map set at the same moment, different image acquisition devices correspond to different depth maps.
[0023] In this embodiment, multiple image acquisition devices can be evenly distributed on a circular ring. This allows for the acquisition of depth maps from various angles of the target scene, thereby improving the playback effect of the generated 3D full-realistic video. It should be noted that the multiple image acquisition devices can also be distributed on an arc, a straight line, a rectangle, or a square; no specific limitation is made here.
[0024] In one feasible approach, multiple depth maps at the same time can be acquired directly from image acquisition devices deployed at multiple acquisition angles. These image acquisition devices can be depth cameras (such as binocular depth cameras, time-of-flight (TOF) depth cameras, or structured light cameras).
[0025] In another feasible approach, multiple texture maps at the same moment can be acquired first from image acquisition devices deployed at multiple acquisition angles. These texture maps can then be processed to convert them into multiple depth maps. In this case, the image acquisition devices can be ordinary two-dimensional (2D) cameras.
[0026] It should be noted that a depth atlas can be considered as a Group of Pictures (GOP).
[0027] Step 102: Based on the acquisition viewpoint of each depth map in each depth map set, determine the target depth map and the corresponding reference depth map from each depth map set.
[0028] In this embodiment, the acquisition perspective may include a first perspective and a second perspective. The first perspective is a secondary perspective, and the second perspective is a primary perspective. The image acquisition device for the primary perspective can be referred to as a Master device, and the image acquisition device for the secondary perspective can be referred to as a Slave device.
[0029] It should be noted that there can be multiple first-viewpoints, but only one second-viewpoint. If a depth map is considered as a GOP, then the depth map acquired by the image acquisition device of the second viewpoint (i.e., the main viewpoint) can be an IDR frame, and the depth map acquired by the image acquisition device of the first viewpoint can be a P frame.
[0030] In the embodiments of this application, for each depth map set at any given time, a target depth map that needs to be encoded can be determined from each depth map set based on the acquisition viewpoint of each depth map. Then, based on the number of target depth maps, different filtering strategies can be used to determine the reference depth map corresponding to the target depth map from the depth map set.
[0031] It should be noted that one target depth map corresponds to one reference depth map, and the target depth map and the reference depth map are different.
[0032] Step 103: Based on the pixel information of the reference depth map, encode the target depth map to obtain the target encoded map.
[0033] In this embodiment, the pixel information of the reference depth map can be the depth value corresponding to each pixel in the reference depth map. Specifically, the target depth map can be encoded based on the depth value corresponding to each pixel in the reference map to obtain a target encoded map.
[0034] In other embodiments of this application, step 102 described above can be implemented in the following ways: Step 102a: Determine the depth map in each depth map set whose viewing angle is the first viewpoint as the target depth map.
[0035] In the embodiments of this application, for the depth map set at each moment, the depth map with the acquisition viewpoint as the first viewpoint (i.e., the secondary viewpoint) can be determined as the target depth map that needs to be encoded. That is, in the depth map set at each moment, all depth maps except the depth map of the main viewpoint are target depth maps.
[0036] It should be noted that the depth map of the main view (i.e., the second view) at a certain moment is based on the depth map of the main view at the previous moment. The method of determining the reference depth map corresponding to the depth map of the main view and the method of encoding the depth map of the main view are existing technologies and will not be elaborated here.
[0037] In this embodiment of the application, step 102b or steps 102c-102f can be performed after step 102a.
[0038] Step 102b: If there is only one target depth map, determine the depth map with the second perspective from each depth map set as the reference depth map.
[0039] The second perspective is the perspective other than the first perspective among multiple acquisition perspectives.
[0040] In this embodiment of the application, after determining the target depth map, the number of target depth maps can be determined. If there is only one target depth map in the depth map set, the depth map of the second view (i.e. the main view) is directly determined as the reference depth map.
[0041] Step 102c: If there are multiple target depth maps, sort the multiple target depth maps according to the target sorting strategy.
[0042] In the embodiments of this application, the target ranking strategy can be determined based on the deployment method of multiple image acquisition devices; different deployment methods correspond to different target ranking strategies.
[0043] In one feasible approach, if multiple image acquisition devices are deployed in a ring shape, the target sorting strategy can be to determine the target slave device from the slave cameras adjacent to the master device, and then sort all slave devices in a clockwise / counterclockwise direction starting from the target slave device. Subsequently, the multiple target depth maps are sorted according to the order of the slave devices. Specifically, if the target slave device is located clockwise from the master device, all slave devices can be sorted clockwise; conversely, if the target slave device is located counterclockwise from the master device, all slave devices can be sorted counterclockwise. It should be noted that the depth map at the starting point of the sorting is the depth map ranked first.
[0044] In another possible approach, if multiple image acquisition devices are deployed in a straight line, the target sorting strategy can be to first determine the target Slave device from the Slave devices adjacent to the Master device. If there is only one Slave device adjacent to the Master device, it means that the Master device is deployed at both ends of the line. In this case, starting from the target Slave device, all Slave cameras are sorted along the direction away from the Master device. If there are two Slave devices adjacent to the Master device, one of them is randomly selected as the target Slave device. Starting from the target Slave device, all Slave cameras are sorted from the other Slave device adjacent to the Master device. Then, the multiple target depth maps are sorted according to the arrangement order of all Slave devices.
[0045] For example, if multiple image acquisition devices are deployed in a linear fashion in the order of first slave device, second slave device, master device, third slave device, and fourth slave device, and the target slave device is the third slave device, then the sorting result after sorting all slave devices is third slave device, fourth slave device, first slave device, and second slave device. Then, the multiple target depth maps are sorted according to the arrangement order of all slave devices, and the sorting result is: third target depth map, fourth target depth map, first target depth map, and second target depth map.
[0046] It should be noted that the first-ranked depth map is usually the depth map adjacent to the depth map of the master device's main view.
[0047] Step 102d: From each depth map set, determine the depth map with the second perspective as the reference depth map corresponding to the target depth map ranked first.
[0048] In the embodiments of this application, such as Figure 2 As shown, for each depth map set (i.e., the depth map set at each time moment), the depth map with the acquisition viewpoint as the main viewpoint (i.e., the depth map acquired by the Master device) can be directly determined as the reference depth map corresponding to the target depth map ranked first.
[0049] Step 102e: For the target depth map sorted i, determine the candidate depth maps adjacent to the target depth map sorted i from each depth map set.
[0050] Where i is an integer greater than 1.
[0051] In this embodiment of the application, if there are multiple target depth maps, then for the target depth map ranked i in the depth map set other than the target depth map ranked first, the depth map adjacent to the target depth map ranked i in the depth map set can be determined from the depth map set, that is, the candidate depth map.
[0052] Specifically, for the target depth map ranked second in each depth map set, the depth maps adjacent to the target depth map ranked second in each depth map set can be determined, that is, the target depth map ranked first and the target depth map ranked third are candidate depth maps. For the target depth map ranked third in each depth map set, the target depth map ranked second and the target depth map ranked fourth in each depth map set can be determined as candidate depth maps, and so on, until the candidate depth map corresponding to the target depth map ranked last is determined.
[0053] Step 102f: Determine the reference depth map corresponding to the sorted target depth map from the candidate depth maps.
[0054] In the embodiments of this application, such as Figure 2 As shown, for the target depth map ranked i, a reference depth map can be determined from the candidate depth maps whose ranking position is between the target depth maps ranked i. Specifically, for the target depth map ranked second, the target depth map ranked first can be determined from the target depth maps ranked first and third as the reference depth map; for the target depth map ranked third, the target depth map ranked second can be determined from the target depth maps ranked second and fourth as the reference depth map, and so on, until the reference depth map corresponding to the target depth map ranked last is determined.
[0055] It should be noted that parallel processing can be used to simultaneously determine the reference depth map corresponding to each target depth map.
[0056] In other embodiments of this application, step 103 described above can be implemented in the following ways: Step 103a: Determine the motion vector corresponding to the target depth map based on pixel information.
[0057] In this embodiment, pixel information may include the depth value corresponding to the pixel. Specifically, for each target depth map, the motion vector (MV) corresponding to each target depth map can be determined based on the depth value corresponding to each pixel in the reference depth map corresponding to each target depth map.
[0058] It should be noted that one target depth map corresponds to one motion vector.
[0059] In the embodiments of this application, step 103a can be implemented by steps 103a1-103a3.
[0060] Step 103a1: Obtain the first pose information of the first image acquisition device corresponding to the target depth map.
[0061] In this embodiment, the first pose information may include a first rotation matrix and a first translation matrix.
[0062] It should be noted that if the first image acquisition device is a camera, then the first pose information can refer to the camera's extrinsic parameters.
[0063] Step 103a2: Obtain the attribute information and second pose information of the second image acquisition device for the reference depth map corresponding to the target depth map.
[0064] In this embodiment, the attribute information may include the focal length, distortion coefficient, pixel size, and optical center coordinates of the second image acquisition device; the second pose information may include a second rotation matrix and a second translation matrix. The focal length includes a first focal length in the x-direction and a second focal length in the y-direction; the optical center coordinates include a first optical center coordinate in the x-direction and a second optical center coordinate in the y-direction.
[0065] It should be noted that if the second image acquisition device is a camera, the attribute information can refer to the camera's intrinsic parameters, and the second pose information can refer to the camera's extrinsic parameters.
[0066] Step 103a3: Determine the motion vector based on pixel information, attribute information, first pose information and second pose information.
[0067] In this embodiment of the application, the depth value, attribute information, first pose information and second pose information corresponding to each pixel in the reference depth map can be calculated to obtain the motion vector corresponding to the target depth map.
[0068] In the embodiments of this application, step 103a3 can be implemented by steps 103a31-103a34.
[0069] Step 103a31: Determine the first matrix and the second matrix based on the first pose information and the second pose information.
[0070] In this embodiment of the application, the first pose information and the second pose information can be calculated to obtain the first matrix and the second matrix.
[0071] In the embodiments of this application, step 103a3 can be implemented in the following ways.
[0072] A1. Determine the first matrix based on the first rotation matrix and the second rotation matrix.
[0073] The first pose information includes a first rotation matrix and a first translation matrix; the second pose information includes a second rotation matrix and a second translation matrix.
[0074] In the embodiments of this application, the first rotation matrix and the second rotation matrix can be operated according to the following formula (1) to obtain the first matrix.
[0075] Formula (1) Where R represents the first matrix, Let Rc denote the first rotation matrix and Rc denote the second rotation matrix.
[0076] A2. Determine the second matrix based on the first translation matrix, the second translation matrix, and the first rotation matrix.
[0077] In this embodiment, the first translation matrix and the second translation matrix can be subtracted first, and then the matrix obtained by subtraction can be multiplied by the first rotation matrix to obtain the second matrix. It should be noted that the calculation formula for the second matrix is as shown in formula (2) below: Formula (2) Where T represents the second matrix, Let Tc represent the first translation matrix and Tc represent the second translation matrix.
[0078] Step 103a32: Construct a third matrix based on attribute information.
[0079] In this embodiment, the first focal length in the x-direction, the second focal length in the y-direction, the first optical center coordinates in the x-direction, and the second optical center coordinates in the y-direction of the second image acquisition device can be obtained from the attribute information. Then, a third matrix S can be constructed based on the first focal length, the second focal length, the first optical center coordinates, and the second optical center coordinates. .
[0080] Where fx represents the first focal length, fy represents the second focal length, cx represents the first optical center coordinates, and cy represents the second optical center coordinates.
[0081] Step 103a33: Determine the fourth matrix based on the first matrix, the third matrix, and the pixel information.
[0082] In this embodiment of the application, the depth value corresponding to each pixel in the reference depth map, the first matrix, and the third matrix can be multiplied to obtain the fourth matrix.
[0083] It should be noted that the formula for calculating the fourth matrix is as shown in formula (3) below: Formula (3) Where E represents the fourth matrix, S represents the third matrix, R represents the first matrix, and d represents the depth value corresponding to each pixel in the reference depth map.
[0084] Step 103a34: Determine the motion vector based on the fourth matrix and the second matrix.
[0085] In this embodiment of the application, the fourth matrix and the second matrix can be added together according to the following formula (4) to obtain the motion vector.
[0086] Formula (4) Where mv represents the motion vector, E represents the fourth matrix, and T represents the second matrix.
[0087] In this embodiment, compared to related technologies that determine motion vectors by referring to the position information of pixels in the depth map, this application considers not only the characteristics of the depth map (i.e., the depth value corresponding to the pixel in the depth map) but also the attribute information and pose information of the image acquisition device itself when determining the motion vector. This not only significantly improves the accuracy of the motion vector corresponding to the determined target depth map, but also solves the problem in related technologies where the characteristics of the depth map are ignored, resulting in ghosting in the decoded depth map, which affects the generation of 3D full-life video.
[0088] Step 103b: Encode the target depth map based on the motion vector to obtain the target encoded map.
[0089] In this embodiment of the application, the motion vector residual (MVD) and the index of the reference frame can be determined based on the determined motion vector mv. Then, the target depth map can be encoded based on the motion vector residual MVD and the index of the reference frame to obtain the target encoded map.
[0090] It should be noted that the process of encoding the target depth map based on motion vector mv is an existing technology, which is only briefly explained here and will not be elaborated on in detail.
[0091] In this embodiment of the application, by incorporating the correlation between the depth value corresponding to the pixel in the depth map and the depth map at the same time into the depth map encoding process, it is possible to save more depth map information without changing the bit rate, thereby solving the problem of ghosting in the synthesized virtual viewpoint image when the client uses the decoded depth map to synthesize virtual viewpoint.
[0092] It should be noted that the descriptions of the same steps and contents as in other embodiments in this embodiment can be found in the descriptions in other embodiments, and will not be repeated here.
[0093] The image encoding method provided in this application, for a depth map set at a certain moment, can first determine the target depth map (i.e., the sub-view depth map) to be encoded from the depth map set based on the acquisition viewpoint of each depth map in the depth map set, and then determine the reference depth map corresponding to the target depth map from the depth map set at that moment. Then, the target depth map is encoded based on the pixel information of the reference depth map. That is, when encoding the target depth map at a certain moment, the corresponding reference depth map is directly determined from the depth map set at that moment, instead of determining the reference depth map from the depth map acquired at the previous moment as in related technologies. This solves the problem of poor efficiency and low accuracy in determining the reference depth map when encoding depth maps in related technologies, thereby improving the encoding efficiency of the target depth map.
[0094] Based on the foregoing embodiments, embodiments of this application provide an image encoding apparatus that can be applied to... Figure 1 In the image encoding method provided in the corresponding embodiment, refer to Figure 3 As shown, the image encoding device 2 may include: an acquisition unit 21, a determination unit 22, and a processing unit 23, wherein: The acquisition unit 21 is used to acquire multiple depth maps at different times for the target scene; wherein, different depth maps in each depth map set correspond to different acquisition perspectives; The determining unit 22 is used to determine the target depth map and the reference depth map corresponding to the target depth map from each depth map set based on the acquisition view of each depth map in each depth map set; The processing unit 23 is used to encode the target depth map based on the pixel information of the reference depth map to obtain the target encoded map.
[0095] In other embodiments of this application, the determining unit 22 is further configured to perform the following steps: The depth map acquired from the first-person perspective in each depth map set is identified as the target depth map. If there is only one target depth map, the depth map with the second perspective from each depth map set is determined as the reference depth map; where the second perspective is the perspective other than the first perspective among the multiple acquisition perspectives.
[0096] In other embodiments of this application, the determining unit 22 is further configured to perform the following steps: If there are multiple target depth maps, sort them according to the target sorting strategy; From each depth map set, the depth map with the second perspective is identified as the reference depth map corresponding to the target depth map ranked first. For the target depth map sorted i, determine the candidate depth maps adjacent to the target depth map sorted i from each depth map set; where i is an integer greater than 1. Determine the reference depth map corresponding to the target depth map ranked i from the candidate depth maps.
[0097] In other embodiments of this application, the processing unit 23 is further configured to perform the following steps: Determine the motion vector corresponding to the target depth map based on pixel information; The target depth map is encoded based on the motion vectors to obtain the target encoded map.
[0098] In other embodiments of this application, the processing unit 23 is further configured to perform the following steps: Obtain the first pose information of the first image acquisition device corresponding to the target depth map; Obtain the attribute information and second pose information of the second image acquisition device for the reference depth map corresponding to the target depth map; The motion vector is determined based on pixel information, attribute information, first pose information, and second pose information.
[0099] In other embodiments of this application, the processing unit 23 is further configured to perform the following steps: Based on the first pose information and the second pose information, determine the first matrix and the second matrix; Construct a third matrix based on attribute information; The fourth matrix is determined based on the first matrix, the third matrix, and pixel information; The motion vector is determined based on the fourth matrix and the second matrix.
[0100] In other embodiments of this application, the processing unit 23 is further configured to perform the following steps: The first matrix is determined based on the first rotation matrix and the second rotation matrix; A second matrix is determined based on the first translation matrix, the second translation matrix, and the first rotation matrix; wherein, the first pose information includes the first rotation matrix and the first translation matrix; and the second pose information includes the second rotation matrix and the second translation matrix.
[0101] It should be noted that the specific implementation process of the steps performed by each unit in the embodiments of this application can be referred to Figure 1 The implementation process of the image encoding method provided in the corresponding embodiments will not be described in detail here.
[0102] The image encoding apparatus provided in the embodiments of this application, for a depth map set at a certain moment, can first determine the target depth map (i.e., the sub-view depth map) to be encoded from the depth map set based on the acquisition viewpoint of each depth map in the depth map set, and then determine the reference depth map corresponding to the target depth map from the depth map set at that moment. Then, the target depth map is encoded based on the pixel information of the reference depth map. That is, when encoding the target depth map at a certain moment, the corresponding reference depth map is directly determined from the depth map set at that moment, instead of determining the reference depth map from the depth map acquired at the previous moment as in related technologies. This solves the problem of poor efficiency and low accuracy in determining the reference depth map when encoding depth maps in related technologies, thereby improving the encoding efficiency of the target depth map.
[0103] Based on the foregoing embodiments, embodiments of this application provide an image encoding device that can be applied to... Figure 1 In the image encoding method provided in the corresponding embodiment, refer toFigure 4 As shown, the image encoding device 3 may include: a processor 31 and a memory 32 for storing computer programs capable of running on the processor, and the image encoding device 3 may further include a communication bus 33, wherein: Communication bus 33 is used to realize the communication connection between processor 31 and memory 32; The processor 31 is used to execute the image encoding program in the memory 32 to perform the following steps: Acquire multiple depth maps at different times for the target scene; where different depth maps in each depth map set correspond to different acquisition perspectives; Based on the acquisition viewpoint of each depth map in each depth map set, the target depth map and the corresponding reference depth map are determined from each depth map set. Based on the pixel information of the reference depth map, the target depth map is encoded to obtain the target encoded map.
[0104] In other embodiments of this application, the processor 31 is used to execute an image encoding program in the memory 32 to perform the following steps: The depth map acquired from the first-person perspective in each depth map set is identified as the target depth map. If there is only one target depth map, the depth map with the second perspective from each depth map set is determined as the reference depth map; where the second perspective is the perspective other than the first perspective among the multiple acquisition perspectives.
[0105] In other embodiments of this application, the processor 31 is used to execute an image encoding program in the memory 32 to perform the following steps: If there are multiple target depth maps, sort them according to the target sorting strategy; From each depth map set, the depth map with the second perspective is identified as the reference depth map corresponding to the target depth map ranked first. For the target depth map sorted i, determine the candidate depth maps adjacent to the target depth map sorted i from each depth map set; where i is an integer greater than 1. Determine the reference depth map corresponding to the target depth map ranked i from the candidate depth maps.
[0106] In other embodiments of this application, the processor 31 is used to execute an image encoding program in the memory 32 to perform the following steps: Determine the motion vector corresponding to the target depth map based on pixel information; The target depth map is encoded based on the motion vectors to obtain the target encoded map.
[0107] In other embodiments of this application, the processor 31 is used to execute an image encoding program in the memory 32 to perform the following steps: Obtain the first pose information of the first image acquisition device corresponding to the target depth map; Obtain the attribute information and second pose information of the second image acquisition device for the reference depth map corresponding to the target depth map; The motion vector is determined based on pixel information, attribute information, first pose information, and second pose information.
[0108] In other embodiments of this application, the processor 31 is used to execute an image encoding program in the memory 32 to perform the following steps: Based on the first pose information and the second pose information, determine the first matrix and the second matrix; Construct a third matrix based on attribute information; The fourth matrix is determined based on the first matrix, the third matrix, and pixel information; The motion vector is determined based on the fourth matrix and the second matrix.
[0109] In other embodiments of this application, the processor 31 is used to execute an image encoding program in the memory 32 to perform the following steps: The first matrix is determined based on the first rotation matrix and the second rotation matrix; A second matrix is determined based on the first translation matrix, the second translation matrix, and the first rotation matrix; wherein, the first pose information includes the first rotation matrix and the first translation matrix; and the second pose information includes the second rotation matrix and the second translation matrix.
[0110] It should be noted that a detailed description of the steps performed by the processor can be found in [reference needed]. Figure 1 The image encoding method provided in the corresponding embodiments will not be described in detail here.
[0111] The image encoding device provided in the embodiments of this application, for a depth map set at a certain moment, can first determine the target depth map (i.e., the sub-view depth map) to be encoded from the depth map set based on the acquisition viewpoint of each depth map in the depth map set, and then determine the reference depth map corresponding to the target depth map from the depth map set at that moment. Then, the target depth map is encoded based on the pixel information of the reference depth map. That is, when encoding the target depth map at a certain moment, the corresponding reference depth map is directly determined from the depth map set at that moment, instead of determining the reference depth map from the depth map acquired at the previous moment as in related technologies. This solves the problem of poor efficiency and low accuracy in determining the reference depth map when encoding depth maps in related technologies, thereby improving the encoding efficiency of the target depth map.
[0112] Based on the foregoing embodiments, embodiments of this application provide a storage medium storing a computer program thereon, which is implemented when executed by a processor. Figure 1 The corresponding embodiment provides the steps of the image encoding method.
[0113] Based on the foregoing embodiments, embodiments of this application provide a computer program product, which includes a computer program that, when executed by a processor, implements... Figure 1 The corresponding embodiment provides the steps of the image encoding method.
[0114] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of hardware embodiments, software embodiments, or embodiments combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage and optical storage) containing computer-usable program code.
[0115] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart... Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0116] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0117] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. One or more processes and / or boxes The steps of the function specified in one or more boxes.
[0118] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
Claims
1. An image encoding method, characterized in that, The method includes: Acquire multiple depth maps at different times for the target scene; where different depth maps in each depth map set correspond to different acquisition perspectives; Based on the acquisition viewpoint of each depth map in each depth map set, a target depth map and a reference depth map corresponding to the target depth map are determined from each depth map set; Based on the pixel information of the reference depth map, the target depth map is encoded to obtain the target encoded map.
2. The method according to claim 1, characterized in that, The step of determining a target depth map and a corresponding reference depth map from each depth map set based on the acquisition viewpoint of each depth map in each depth map set includes: The depth map in each depth map set whose acquisition viewpoint is the first viewpoint is determined as the target depth map; If the number of target depth maps is one, the depth map with the acquisition viewpoint as the second viewpoint is determined from each depth map set as the reference depth map; wherein, the second viewpoint is a viewpoint other than the first viewpoint among multiple acquisition viewpoints.
3. The method according to claim 2, characterized in that, The method further includes: If there are multiple target depth maps, the multiple target depth maps are sorted according to the target sorting strategy; From each depth map set, the depth map whose acquisition viewpoint is the second viewpoint is determined as the reference depth map corresponding to the target depth map ranked first. For a target depth map of sorting i, determine candidate depth maps adjacent to the target depth map of sorting i from each depth map set; where i is an integer greater than 1. From the candidate depth maps, determine the reference depth map corresponding to the sorted target depth map i.
4. The method according to claim 1, characterized in that, The step of encoding the target depth map based on the pixel information of the reference depth map to obtain a target encoded map includes: The motion vector corresponding to the target depth map is determined based on the pixel information; The target depth map is encoded based on the motion vector to obtain the target encoded map.
5. The method according to claim 4, characterized in that, Determining the motion vector corresponding to the target depth map based on the pixel information includes: Obtain the first pose information of the first image acquisition device corresponding to the target depth map; Obtain the attribute information and second pose information of the second image acquisition device for the reference depth map corresponding to the target depth map; The motion vector is determined based on the pixel information, the attribute information, the first pose information, and the second pose information.
6. The method according to claim 5, characterized in that, Determining the motion vector based on the pixel information, the attribute information, the first pose information, and the second pose information includes: Based on the first pose information and the second pose information, determine the first matrix and the second matrix; Construct a third matrix based on the attribute information; Based on the first matrix, the third matrix, and the pixel information, a fourth matrix is determined; The motion vector is determined based on the fourth matrix and the second matrix.
7. The method according to claim 6, characterized in that, The step of determining the first matrix and the second matrix based on the first pose information and the second pose information includes: The first matrix is determined based on the first rotation matrix and the second rotation matrix; The second matrix is determined based on the first translation matrix, the second translation matrix, and the first rotation matrix; wherein the first pose information includes the first rotation matrix and the first translation matrix; and the second pose information includes the second rotation matrix and the second translation matrix.
8. An image encoding device, characterized in that, include: Processor and memory used to store computer programs that can run on the processor; When the processor is used to run a computer program, it executes the steps of the image encoding method according to any one of claims 1 to 7.
9. A storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the image encoding method according to any one of claims 1 to 7.
10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the image encoding method according to any one of claims 1 to 7.