Displacement information video generation method and device, displacement information video decoding method and device and electronic equipment
The method enhances real-time processing of 3D mesh data by converting multiple frames of displacement data into a displacement video format, addressing the limitations of existing compression standards.
Patent Information
- Application Number
- CN202410061885.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-01-15
- Publication Date
- 2025-07-15
AI Technical Summary
The existing standards for compressed three-dimensional networks cannot meet real-time requirements.
By processing multi-frame displacement data, the size of the displacement video is determined and converted into displacement video, improving the flexibility and real-timeness of video encoding.
It improves the real-time and coding efficiency of three-dimensional grid compression, and meets the real-time processing needs.
Smart Images

Figure CN120321415A_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of communication technologies, and particularly relates to a method and device for generating and decoding displacement information videos, and an electronic device. Background Art
[0002] With the rapid development of multimedia technologies, relevant research results have been rapidly industrialized and have become an essential part of people's lives. 3D models have become a new generation of digital media following audio, images, and videos. 3D meshes are a commonly used representation of 3D models. Compared with traditional multimedia such as images and videos, 3D network models have stronger interactivity and realism, making them increasingly widely used in various fields such as commerce, manufacturing, construction, education, medicine, entertainment, art, and the military.
[0003] The Visual Volumetric Video-based Coding (V3C) standard provides a method for encoding and decoding various 3D media through video or image coding technologies. Specifically, before encoding, it converts 3D media content from a 3D representation to multiple 2D representations (referred to as V3C components) through projection or other means, and then uses existing video or image coding technologies to encode the 2D representations. V3C components mainly include occupancy components, geometric components, and attribute components. The occupancy component can represent which regions in the 2D representation are associated with the data in the 3D representation; the geometric component represents information related to the position of the 3D data in space, and the attribute component can provide attribute information corresponding to vertices, such as materials, textures, etc. In addition, the components also contain information on how to reconstruct the 3D model through these components, which is called atlas information.
[0004] Atlas information is used to associate all components, and additional information for reconstructing the 3D from the 2D is also included in the atlas component. The atlas consists of multiple basic units, and the basic unit is called a patch. Each patch represents a region in the available 2D components and contains the information required to project that region back into 3D space.
[0005] Video-based Dynamic Mesh Coding (VDMC) is a standard for compressing 3D meshes. Its main idea is to compress 3D meshes by using the existing V3C standard. Since 3D meshes have connection information that needs to be encoded, its specific encoding process is slightly different from that of V3C, and it is necessary to extend the syntax semantics and decoding operations at the decoding end of the V3C standard to support the decoding and reconstruction of 3D meshes.
[0006] However, the existing standards for compressing 3D networks cannot meet the real-time requirements. Summary of the Invention
[0007] An embodiment of the present application provides a method and apparatus for generating a displacement information video, a decoding method and apparatus, and an electronic device, which can solve the problem that the existing standards for compressing three-dimensional networks cannot meet the real-time requirements.
[0008] In a first aspect, a method for generating a displacement information video is provided, including:
[0009] A first electronic device processes multiple frames of displacement data to obtain multiple frames of first displacement data;
[0010] The first electronic device determines the size of the displacement video corresponding to the multiple frames of displacement data;
[0011] The first electronic device converts the multiple frames of first displacement data into a displacement video according to the size of the displacement video.
[0012] In a second aspect, a decoding method is provided, including:
[0013] A second electronic device obtains a displacement bitstream and decodes the displacement bitstream into a displacement video;
[0014] The second electronic device extracts pixel values from each frame of displacement image of the displacement video in a target order to obtain target data; the target order is the same as the order in which the first electronic device arranges the displacement images;
[0015] The second electronic device processes the extracted target data to obtain multiple frames of displacement data.
[0016] In a third aspect, a displacement information video generation apparatus is provided, which is applied to a first electronic device and includes:
[0017] A first processing module, configured to process multiple frames of displacement data to obtain multiple frames of first displacement data;
[0018] A first determination module, configured to determine the size of the displacement video corresponding to the multiple frames of displacement data;
[0019] A first conversion module, configured to convert the multiple frames of first displacement data into a displacement video according to the size of the displacement video.
[0020] In a fourth aspect, a decoding apparatus is provided, which is applied to a second electronic device and includes:
[0021] An acquisition module, configured to obtain a displacement bitstream and decode the displacement bitstream into a displacement video;
[0022] An extraction module, configured to extract pixel values from each displacement image of the displacement video in a target order to obtain target data; the target order is the same as the order in which the first electronic device arranges the displacement images;
[0023] A second processing module, configured to process the extracted target data to obtain multiple frames of displacement data.
[0024] In a fifth aspect, a first electronic device is provided, which includes a processor and a memory. The memory stores a program or instruction that can be run on the processor. When the program or instruction is executed by the processor, the steps of the method described in the first aspect are implemented.
[0025] In a sixth aspect, a first electronic device is provided, including a processor and a communication interface. The communication interface is configured to obtain multiple frames of displacement data, and the processor is configured to process the multiple frames of displacement data to obtain multiple frames of first displacement data; determine the size of the displacement video corresponding to the multiple frames of displacement data; and convert the multiple frames of first displacement data into a displacement video according to the size of the displacement video.
[0026] In a seventh aspect, a second electronic device is provided, which includes a processor and a memory. The memory stores a program or instruction that can be run on the processor. When the program or instruction is executed by the processor, the steps of the method described in the second aspect are implemented.
[0027] In an eighth aspect, a second electronic device is provided, including a processor and a communication interface. The communication interface is configured to obtain a displacement bitstream; the processor is configured to decode the displacement bitstream into a displacement video; extract pixel values from each displacement image of the displacement video in a target order to obtain target data; the target order is the same as the order in which the first electronic device arranges the displacement images; and process the extracted target data to obtain multiple frames of displacement data.
[0028] In a ninth aspect, a readable storage medium is provided. The readable storage medium stores a program or instruction. When the program or instruction is executed by a processor, the steps of the method described in the first aspect are implemented, or the steps of the method described in the second aspect are implemented.
[0029] In a tenth aspect, a wireless communication system is provided, including: a first electronic device and a second electronic device. The first electronic device can be used to execute the steps of the method described in the first aspect, and the second electronic device can be used to execute the steps of the method described in the second aspect.
[0030] In an eleventh aspect, a chip is provided, which includes a processor and a communication interface. The communication interface is coupled to the processor, and the processor is configured to run programs or instructions to implement the method described in the first aspect or the method described in the second aspect.
[0031] In a twelfth aspect, a computer program / program product is provided. The computer program / program product is stored in a storage medium and is executed by at least one processor to implement the steps of the method described in the first aspect or the method described in the second aspect.
[0032] In the embodiments of the present application, multiple frames of displacement data are converted into a displacement video, and the size of the displacement video is determined before generating the displacement video, thereby improving the flexibility and real-time performance of subsequent video coding. BRIEF DESCRIPTION OF THE DRAWINGS
[0033] Figure 1 A block diagram showing a wireless communication system to which the embodiments of the present application can be applied;
[0034] Figure 2 A flowchart showing the steps of the displacement information video generation method provided by the embodiments of the present application;
[0035] Figure 3 A schematic diagram showing the principle of the displacement information video generation method provided by the embodiments of the present application;
[0036] Figure 4 A flowchart showing the steps of the decoding method provided by the embodiments of the present application;
[0037] Figure 5 A schematic diagram showing the principle of the decoding method provided by the embodiments of the present application;
[0038] Figure 6 A schematic diagram showing the application of the displacement information video generation method provided by the embodiments of the present application in the VDMC framework;
[0039] Figure 7 A schematic diagram showing an example of mesh simplification in the VDMC framework provided by the embodiments of the present application;
[0040] Figure 8 A schematic diagram showing the principle of the sub-grid compression module in the VDMC framework provided by the embodiments of the present application;
[0041] Figure 9 A schematic diagram showing the application of the decoding method provided by the embodiments of the present application in the VDMC framework;
[0042] Figure 10 A schematic diagram showing the structure of the displacement information video generation device provided by the embodiments of the present application;
[0043] Figure 11 It is a schematic structural diagram of the decoding device provided by the embodiment of the present application;
[0044] Figure 12 It is a schematic structural diagram of the communication device provided by the embodiment of the present application;
[0045] Figure 13 It is a schematic structural diagram of the terminal provided by the embodiment of the present application;
[0046] Figure 14 It is a schematic structural diagram of the network-side device provided by the embodiment of the present application. Detailed implementation manners
[0047] Next, the technical solutions in the embodiments of the present application will be clearly described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art belong to the scope of protection of the present application.
[0048] The terms "first", "second", etc. in the present application are used to distinguish similar objects, rather than to describe a specific order or sequence. It should be understood that such terms can be interchanged under appropriate circumstances, so that the embodiments of the present application can be implemented in an order other than those illustrated or described herein, and the objects distinguished by "first" and "second" are usually of the same category, and the number of objects is not limited. For example, the first object can be one or multiple. In addition, "or" in the present application means at least one of the connected objects. For example, "A or B" covers three scenarios, namely, Scenario 1: including A and not including B; Scenario 2: including B and not including A; Scenario 3: including both A and B. The character " / " generally indicates an "or" relationship between the associated objects before and after.
[0049] The term "indicate" in the present application can be either a direct indication (or an explicit indication) or an indirect indication (or an implicit indication). Among them, a direct indication can be understood as that the sender clearly informs the receiver of specific information, operations to be performed, or request results, etc. in the sent indication; an indirect indication can be understood as that the receiver determines the corresponding information according to the indication sent by the sender, or makes a judgment and determines the operations to be performed or request results, etc. according to the judgment result.
[0050] It should be noted that the technology described in the embodiments of this application is not limited to Long Term Evolution (LTE) / LTE-Advanced (LTE-A) systems, but can also be used in other wireless communication systems, such as Code Division Multiple Access (CDMA), Time Division Multiple Access (TDMA), Frequency Division Multiple Access (FDMA), Orthogonal Frequency Division Multiple Access (OFDMA), Single-carrier Frequency-Division Multiple Access (SC-FDMA), or other systems. The terms "system" and "network" in the embodiments of this application are often used interchangeably, and the described technology can be used not only in the systems and radio technologies mentioned above, but also in other systems and radio technologies. The following description describes the New Radio (NR) system for example purposes, and the NR terms are used in most of the following descriptions, but these technologies can also be applied to systems other than the NR system, such as the 6th th Generation (6G) communication system.
[0051] Figure 1A block diagram of a wireless communication system to which embodiments of the present application can be applied is shown. The wireless communication system includes a terminal 11 and a network-side device 12. Among them, the terminal 11 can be a mobile phone, a tablet personal computer, a laptop computer, a notebook computer, a personal digital assistant (PDA), a handheld computer, a netbook, an ultra-mobile personal computer (UMPC), a mobile internet device (MID), an augmented reality (AR), a virtual reality (VR) device, a robot, a wearable device, a flight vehicle, a vehicle user equipment (VUE), a shipborne device, a pedestrian user equipment (PUE), a smart home (home appliances with wireless communication functions, such as refrigerators, TVs, washing machines or furniture, etc.), a game console, a personal computer (PC), a teller machine or a self-service machine, etc. Wearable devices include: smart watches, smart bracelets, smart earphones, smart glasses, smart jewelry (smart bracelets, smart bracelets, smart rings, smart necklaces, smart anklets, smart ankle chains, etc.), smart wristbands, smart clothing, etc. Among them, the vehicle user equipment can also be referred to as a vehicle terminal, a vehicle controller, a vehicle module, a vehicle component, a vehicle chip or a vehicle unit, etc. It should be noted that the specific type of the terminal 11 is not limited in the embodiments of the present application. The network-side device 12 can include an access network device or a core network device. Among them, the access network device can also be referred to as a radio access network (RAN) device, a radio access network function or a radio access network unit. The access network device can include a base station, a wireless local area network (WLAN) access point (AP) or a wireless fidelity (WiFi) node, etc.Among them, the base station can be referred to as Node B (NB), Evolved Node B (eNB), the next generation Node B (gNB), New Radio Node B (NR Node B), access point, Relay Base Station (RBS), Serving Base Station (SBS), Base Transceiver Station (BTS), radio base station, radio transceiver, Basic Service Set (BSS), Extended Service Set (ESS), home Node B (HNB), home evolved Node B, Transmission Reception Point (TRP), or some other suitable term in the art. As long as the same technical effects are achieved, the base station is not limited to specific technical terms. It should be noted that in the embodiments of this application, only the base station in the NR system is taken as an example for introduction, and the specific type of the base station is not limited.
[0052] The triangular mesh is currently the most common representation method for three-dimensional networks. A three-dimensional mesh can be regarded as composed of three basic elements: vertices, edges, and faces. Vertices are the most basic elements in the mesh, which define positions in a three-dimensional space. An edge is a line segment connecting two vertices in the mesh. A face can be regarded as a polygon formed by a closed path of edges. For a triangular mesh, each face is a triangle.
[0053] The information contained in the mesh is usually divided into three categories: geometric information, connectivity information, and attribute information. Geometric information is the position of each vertex of the mesh in three-dimensional space. Connectivity information describes the association relationships between the elements in the mesh, that is, the connection relationships between vertices. Attribute information is optional, and it can associate attributes to the corresponding mesh elements (such as vertex color, normal vector, etc. can be associated with mesh vertices). The mesh can also be mapped from three-dimensional space to a two-dimensional planar region using mesh parameterization. This mapping relationship is usually described by a set of parametric coordinates, called UV coordinates or texture coordinates, which are associated with mesh vertices. This two-dimensional mapping can be used to represent high-resolution attribute information, such as texture, normal vector, etc.
[0054] In almost all application fields that use 3D meshes (such as computational simulation, entertainment, medical imaging, digital cultural relics, computer design, e-commerce, etc.), as people have higher and higher requirements for the visual effects of 3D mesh models, the models are becoming more and more complex and the accuracy of the models is also getting higher and higher. Therefore, the amount of data required to represent 3D meshes has also increased correspondingly. The above problems have led to the increasing complexity of the processing, visualization, transmission, and storage of 3D meshes. 3D mesh compression can be regarded as a way to solve the above problems. It reduces the size of the model data and is beneficial to the processing, storage, and transmission of 3D meshes. Therefore, it is necessary to propose an efficient and general 3D mesh compression algorithm.
[0055] Recently, the international standardization organization (MPEG) in the field of audio and video coding and compression has begun to formulate a compression standard VDMC for 3D meshes. This standard is specified based on the existing V3C standard. The V3C standard provides a general method for compressing 3D models, and the models can be presented in the form of point clouds, meshes, panoramic videos, etc. Compatibility of the 3D mesh model compression method with this standard is helpful for the popularization and applicability of the method. Therefore, it is of great significance to optimize the 3D mesh encoding and decoding method in VDMC and combine the optimization method with the V3C standard. A possible optimization method is the optimization of the intra-frame base mesh encoding and decoding. In the existing framework, the information of the base mesh will be processed by dividing it into connection relationships, geometric information, and attribute information. For the geometric information encoding and decoding of the base mesh in 3D mesh encoding and decoding, a geometric information encoding and decoding that only depends on part of the connection relationships is proposed, which can support the parallel scheme of connection relationship and geometric information encoding and decoding. The encoding and decoding of connection relationships and geometric information can be completed only through one traversal, which can reduce the time complexity.
[0056] The following will combine the accompanying drawings and, through some embodiments and their application scenarios, elaborate in detail on the displacement information video generation method and decoding method provided by the embodiments of the present application.
[0057] As Figure 2 shown, the embodiments of the present application provide a displacement information video generation method, including:
[0058] Step 201, the first electronic device processes multiple frames of displacement data to obtain multiple frames of first displacement data;
[0059] Step 202, the first electronic device determines the size of the displacement video corresponding to the multiple frames of displacement data;
[0060] Step 203, the first electronic device converts the multiple frames of first displacement data into a displacement video according to the size of the displacement video.
[0061] Among them, the displacement data is an important part in the dynamic mesh encoding process, storing the distance information from the vertices of the subdivided base network to the nearest neighbor points of the original input network.
[0062] In at least one embodiment of the present application, the method further includes:
[0063] The first electronic device encodes the displacement video to obtain a displacement bitstream;
[0064] The first electronic device sends the displacement bitstream to the second electronic device.
[0065] It should be noted that in the embodiments of the present application, the first electronic device and the second electronic device may be different devices or the same device. When the first electronic device and the second electronic device are the same device, it can be understood that the device has both encoding and decoding functions, and the encoding and decoding functions can be set in different modules of the same device, which is not specifically limited herein.
[0066] In one implementation, as Figure 3 shown, the preprocessing module processes multiple frames of displacement data, such as prediction, transformation, quantization, etc.; after preprocessing, multiple frames of first displacement data are obtained. After preprocessing, each frame of the first displacement data needs to be input into the Packing module. Packing is a displacement image, and finally multiple frames of displacement data are converted into a displacement video. Before performing the Packing operation, the size of the displacement video needs to be determined and stored in a buffer, which is called by the Packing module. Finally, the generated displacement video is sent to a video encoder for encoding to obtain a displacement bitstream. There are various choices for the specific video encoder, such as High Efficiency Video Coding (HEVC), Vertex Chain Code (VCC), etc.
[0067] As an alternative embodiment, step 202 includes:
[0068] The first electronic device designates the size of the displacement video; for example, the user of the first electronic device directly designates the size of the displacement video;
[0069] Or, step 202 includes:
[0070] The first electronic device determines the size of the displacement video according to the size of the target frame displacement image; for example, using the size of the target frame displacement image as the size of the shifted video.
[0071] In at least one embodiment of the present application, in the implementation process of using the size of the target frame displacement image as the size of the displacement video, it is necessary to first determine the size of the target frame displacement image. In one implementation, the method further includes:
[0072] The first electronic device determines the size of the target frame displacement image according to the number of data blocks in each row of the displacement video, the size of the data blocks in the displacement video, and the number of target frame displacement data.
[0073] It should be noted that after calculating the size of the current target frame displacement image, it is used as the size of the displacement video composed of all displacement images from the current target frame to the next target frame. After determining the size of the displacement video, it is stored in the Buffer; when the next target frame is encountered, the size of the displacement video in the Buffer needs to be updated.
[0074] Wherein, the first electronic device determines the size of the target frame displacement image according to the number of data blocks in each row of the displacement video, the size of the data blocks in the displacement video, and the number of target frame displacement data, including:
[0075] Determine the width of the target frame displacement image according to the number of data blocks in each row of the displacement video and the size of the data blocks in the displacement video;
[0076] Determine the height of the target frame displacement image according to the width of the target frame displacement image and the number of target frame displacement data.
[0077] In an alternative implementation, the number of data blocks (block) in each row of the displacement video can be specified by the user or the first electronic device; in another alternative implementation, the size of the data blocks (block) in the displacement video can be specified by the user or the first electronic device.
[0078] Optionally, determining the width of the target frame displacement image according to the number of data blocks in each row of the displacement video and the size of the data blocks in the displacement video includes:
[0079] Determine the width of the target frame displacement image according to the first formula; wherein, the first formula is:
[0080] Width=VideoWidthInBlocks*VideoBlockSize
[0081] Wherein, Width represents the width of the target frame displacement image; VideoWidthInBlocks represents the number of data blocks in each row of the displacement video; VideoBlockSize represents the size of the data blocks in the displacement video.
[0082] Optionally, determining the height of the target frame displacement image according to the width of the target frame displacement image and the number of target frame displacement data includes:
[0083] Determining the height of the target frame displacement image according to the second formula; wherein, the second formula is:
[0084]
[0085] wherein, Heigth represents the height of the target frame displacement image; NumOfDisplacements represents the number of target frame displacement data; Width represents the width of the target frame displacement image; ceil represents the ceiling operation.
[0086] In at least one embodiment of the present application, step 203 includes:
[0087] The first electronic device arranges each frame of the first displacement data into a displacement image in a predetermined order according to the size of the displacement video; multiple frames of displacement images constitute the displacement video.
[0088] In one implementation, as Figure 3 shown, after determining the size of the displacement video, store it in the Buffer (when encountering the next target frame, the size of the displacement video in the Buffer needs to be updated), for the Packing module. The Packing module sorts each frame of displacement data into a displacement image in a certain order, finally converts it into a displacement video, and outputs possible relevant identifiers.
[0089] The present application does not specify a specific arrangement order, as long as the encoding and decoding ends (i.e., the second electronic device) are consistent. For example, the Morton order can be adopted within the block for arrangement, and the raster scan order can be adopted between blocks, which is not specifically limited herein.
[0090] In an optional embodiment of the present application, during the arrangement process, the method further includes:
[0091] If the number of displacement data of the current frame is less than the size of the displacement video, fill the extra pixels with a default value; the present application does not specify a specific filling value;
[0092] Or,
[0093] If the number of displacement data of the current frame is greater than the size of the displacement video, remove the extra displacement data.
[0094] Among them, the size of the displacement video can also be referred to as the number of pixels in one frame of the displacement video, and the number of pixels is equal to the width of the displacement image multiplied by the height of the displacement image.
[0095] In one implementation, before determining the target frame displacement image, it is necessary to first determine whether the current frame displacement data is the target frame displacement data. The method further includes:
[0096] Determine whether the current frame displacement data is the target frame displacement data according to the first method; the first method includes at least one of the following:
[0097] Preset the Mth frame displacement data as the target frame displacement data;
[0098] Divide multiple frames of displacement data into multiple groups of pictures (GOPs) in advance, and set the Nth frame displacement data of each GOP as the target frame displacement data;
[0099] where M and N are integers greater than or equal to 1 respectively.
[0100] For example, use the first frame displacement data as the target frame displacement data; or divide multiple frames of displacement data into multiple GOPs, and use the first frame displacement data of each GOP as the target frame displacement data.
[0101] In summary, in the embodiments of the present application, multiple frames of displacement data are converted into a displacement video, and the size of the displacement video is determined before generating the displacement video, thereby improving the flexibility and real-time performance of subsequent video coding.
[0102] As Figure 4 shown, the embodiments of the present application also provide a decoding method, including:
[0103] Step 401, the second electronic device obtains a displacement bitstream and decodes the displacement bitstream into a displacement video;
[0104] Step 402, the second electronic device extracts pixel values from each frame displacement image of the displacement video in the target order to obtain target data; the target order is the same as the order in which the first electronic device arranges the displacement images;
[0105] Step 403, the second electronic device processes the extracted target data to obtain multiple frames of displacement data.
[0106] In one implementation, as Figure 5 shown, at the decoding end, first, the displacement bitstream is decoded into a displacement video through a video decoder corresponding to the video encoder on the first electronic device side. Then, the displacement video is input into an unpacking module, which extracts pixel values from each frame displacement image of the displacement video in the same order as the arrangement order on the encoding end according to the possible relevant identifiers. Finally, the extracted data is input into a post-processing module to restore the displacement data. The operations of the post-processing module are the reverse implementation of the operations of the pre-processing module of the first electronic device, such as inverse quantization, inverse transform, prediction, etc.
[0107] In at least one embodiment of the present application, step 402 includes:
[0108] For a frame of displacement image, according to the number of vertices of the subdivided reconstructed basic grid corresponding to the displacement image, determine the number of pixel values to be extracted from the displacement image;
[0109] If the number of vertices of the subdivided reconstructed basic grid is less than or equal to the number of pixel values to be extracted from the displacement image, extract pixel values equal to the number of vertices of the subdivided reconstructed basic grid from the displacement image in the target order to obtain target data;
[0110] Alternatively, if the number of vertices of the subdivided reconstructed basic grid is greater than the number of pixel values to be extracted from the displacement image, extract all pixel values from the displacement image in the target order, and perform a filling operation at the end of the extracted data until the number of data is equal to the number of vertices of the subdivided reconstructed basic grid to obtain target data.
[0111] Optionally, there are various ways to determine the filled value, such as being equal to 0, equal to the mean of the extracted pixel values, etc., which are not specifically limited herein.
[0112] In summary, for the solution of converting multiple frames of displacement data into a displacement video and determining the size of the displacement video before generating the displacement video, the embodiments of the present application provide a corresponding decoding method to improve the decoding correctness.
[0113] The solution on the first electronic device side proposed by the present application is mainly applied in the encoding end of the VDMC framework during the process of "two-dimensional arrangement" of displacements, such as Figure 6 shown. Next, a specific description of the entire application process is given, such as Figure 6 shown.
[0114] (1) Mesh simplification
[0115] Mesh simplification is to simplify the currently input mesh to a basic mesh with relatively fewer points and faces, and as much as possible maintain the shape of the original mesh. The key point of mesh simplification lies in the simplification operation and the corresponding error metric. A feasible mesh simplification operation is as Figure 7 shown, merging the vertices at both ends of an edge into one vertex and deleting the connection between these two vertices. Repeat this process in the entire mesh according to a certain rule to reduce the number of faces and vertices of the mesh to the target value.
[0116] During the simplification process, a certain error metric can be selected to optimize the simplification result. For example, the sum of the equation coefficients of all adjacent faces of a vertex can be chosen as the error metric for that vertex, and the error metric for the corresponding edge is the sum of the error metrics of the two vertices on the edge. In other words, the error generated by merging an edge is the sum of the distances from the merged vertex to all the planes adjacent to the original two vertices of the edge.
[0117] After determining the simplification operation and the corresponding error metric, the mesh simplification is carried out iteratively. First, calculate the vertex errors of the initial mesh to obtain the errors of each edge. Then, arrange each edge in ascending order of error, and each time select the edge with the smallest error for merging. At the same time, calculate the position of the merged vertex and update the errors of all the edges related to the merged vertex. That is, update the order of the edge arrangement to ensure that each iteration is based on the global error metric. Through iteration, simplify the faces of the mesh to the number required for lossy coding.
[0118] (2) Mesh parameterization
[0119] Texture coordinates may be regenerated based on the reconstructed base mesh of the current frame. This step requires generating texture coordinates for each attribute map of the input mesh. If there are multiple attribute maps with similar characteristics, the same texture coordinates can be shared. The ways to generate texture coordinates include mesh parameterization, etc. Currently, there are many algorithms used to parameterize meshes, such as the Isocharts algorithm, which uses spectral analysis to achieve stretch-driven 3D mesh parameterization, unfolds, slices, and packs the 3D mesh into a 2D texture domain.
[0120] (3) Subdivision and deformation
[0121] The subdivision and deformation module is an optional module, which can improve the quality of the mesh reconstructed at the decoding end. When the quality of the base mesh can already meet the application requirements, this module can be not used.
[0122] The same concept of the basic idea of the subdivision and deformation module is applied to the input 3D mesh to generate displacement vector information. The input 2D curve (represented by a 2D polyline), called the "original" curve, is first downsampled to generate a basic curve / polyline, called the "simplified" curve. Then, the subdivision scheme is applied to the simplified polyline to generate the "subdivided" curve. Subsequently, the subdivided polyline is deformed to obtain a better approximation of the original curve. That is, calculate the geometric displacement vector for each vertex of the subdivided mesh so that the shape of the subdivided curve is as close as possible to the shape of the original curve. These geometric displacement vectors are the geometric displacement vector information (vertex displacement) output by this module. The same deformation process is also applied to the attribute information corresponding to the vertices to obtain the corresponding attribute displacement vectors.
[0123] (4) Compression of the base mesh
[0124] The basic grid compression module subdivides the deformed module to output the basic grid information. There are mainly two different modes of basic grid compression, namely the intra-frame mode and the inter-frame mode, as Figure 8 shown. In the intra-frame mode, the input of the basic grid compression module is a three-dimensional grid, including geometric coordinates, connection relationships, and attribute information associated with vertices. In the inter-frame mode, the input of the basic grid compression module is the motion vector and its identifier, as well as the possible intra-frame sub-grid. After encoding, the encoded grid needs to be reconstructed to provide it to subsequent modules for processing. The following briefly introduces the encoding of the two modes:
[0125] ① Intra-frame mode
[0126] For the grid encoding in the intra-frame mode, an existing grid encoder can be used to encode the input basic grid, such as Draco, etc. The type of grid encoder is encoded through auxiliary information and transmitted to the decoding end.
[0127] ② Inter-frame mode
[0128] In the inter-frame mode, it is necessary to encode the motion vector and its identifier. The motion vector identifier is an array composed of numbers greater than or equal to 0, and the size of the array is the same as the number of vertices of the basic grid of the reference frame corresponding to the temporal motion vector identifier. The existing entropy encoding algorithm can be directly used to encode this array, such as CABAC (Context Adaptive Binary Arithmetic Coding), etc.
[0129] (5) Reconstruct the basic grid
[0130] After the basic grid is encoded, it is necessary to reconstruct the basic grid. For the grid encoded in the intra-frame mode, the encoded grid can be decoded to obtain the reconstructed grid.
[0131] For the basic grid information encoded in the inter-frame mode, it is necessary to reconstruct the inter-frame sub-grid according to the motion vector and its identifier. If there is an intra-frame sub-grid in the inter-frame mode, it is also necessary to merge the reconstructed inter-frame sub-grid with the intra-frame sub-grid to obtain the reconstructed basic grid.
[0132] (6) Encoding of vertex displacement
[0133] To encode the vertex displacement, first, it is necessary to adjust the vertex displacement order according to the reconstructed basic grid; then, perform wavelet transform and quantization on the displacement to obtain the quantized wavelet transform coefficients; next, arrange the quantized wavelet transform coefficients into a two-dimensional image, and finally form a YUV video, and encode it with a video encoder to obtain the bitstream of the displacement. Next, each module will be introduced specifically:
[0134] ① Adjust the displacement order
[0135] The main function of this module is to adjust the vertex displacement order according to the reconstructed base mesh. The detailed process is as follows: First, the reconstructed base mesh is subdivided, and then the displacement order is adjusted to be the same as the vertex order of the subdivided base mesh. There is an optional step in this process, which is to convert the coordinate system of the vertex displacement from the Cartesian coordinate system to the local coordinate system.
[0136] ② Wavelet transform
[0137] A transform can be applied to the displacement vector to reduce the correlation between its data. An optional transform such as the linear wavelet transform, and its prediction process is defined as shown in Equation (3):
[0138]
[0139] where v is the newly inserted midpoint on the edge (v1, v2), and Signal(v), Signal(v1), and Signal(v2) are the displacement vectors corresponding to vertices v, v1, and v2 respectively. After predicting the displacement vector of vertex v, it is updated, and the update process is defined as shown in Equation (4):
[0140]
[0141] where v * is the set of all vertices adjacent to vertex v. The transformed displacement vector is called the wavelet coefficient.
[0142] ③ Coefficient quantization
[0143] The transformed displacement vector, i.e., the wavelet coefficient, can be quantized. There are various quantization methods, and one method is shown in Equations (5) and (6):
[0144] disp[v].d[k] = floor(disp[v].d[k] * scale[k]) #(5)
[0145]
[0146] where disp[v] represents the value after transformation of the displacement vector of the v-th vertex, d[k] represents the k-th value of the displacement vector, floor represents rounding down. bitDepthPosition represents the bit depth of the geometric position of the current mesh vertex, and qp[k] represents the quantization parameter of the k-th coefficient. As mentioned above, after converting the coordinate system of the displacement vector, the normal component has a more significant effect on the quality than the tangential component. Therefore, a larger quantization parameter can be used for the tangential component.
[0147] Meanwhile, according to the characteristics of wavelet transform, different quantization parameters can also be used for the newly generated vertices and the original vertices after subdivision. That is, for the vertices after subdivision, the quantization parameter is updated as shown in Equation (7):
[0148] scale[k] = scale[k] * lodScale[k] #(7)
[0149] where lodScale[k] represents the coefficient of the quantization parameter at the current subdivision level.
[0150] ④ Two-dimensional arrangement
[0151] After quantization, the quantized wavelet transform coefficients need to be arranged on a two-dimensional image. Before performing the two-dimensional arrangement, the size of the arranged video needs to be determined first. The process of determining the video size is as follows:
[0152] First, it is determined whether the quantized wavelet transform coefficients of the current frame are the selected frame (which can also be called the target frame). There are various ways to determine the selected frame. For example, the quantized wavelet transform coefficients of the first frame are used as the selected frame; or the quantized wavelet transform coefficients of multiple frames are divided into multiple GOPs, and the quantized wavelet transform coefficients of the first frame of each GOP are used as the selected frame; or the quantized wavelet transform coefficients of each I-frame are determined as the selected frame, etc. If the quantized wavelet transform coefficients of the current frame are the selected frame, the size of the arranged image of the selected frame needs to be calculated, and the calculation process is as shown in Equations (8), (9), and (10):
[0153]
[0154] Width = VideoWidthInBlocks * VideoBlockSize #(9)
[0155]
[0156] where Width represents the width of the arranged image; Heigth represents the height of the arranged image; VideoWidthInBlocks represents the number of blocks in each row of the arranged video (this value can be specified by the user); VideoBlockSize represents the size of the block in the arranged video (this value can be specified by the user); NumOfDisplacements LODi represents the number of coefficients included in the i-th LOD level of the quantized wavelet transform coefficients of the selected frame; ceil represents the ceiling operation; NumOfBlocks LODirepresents the number of blocks required to arrange the coefficients in the \(i\)-th LOD level of the wavelet transform coefficients after quantization of the selected frame into a two-dimensional image; \(n\) represents the number of LOD levels, and \(i = 0\sim n\) represents that there are a total of \(n + 1\) levels. After calculating the size of the arranged image of the current selected frame, use it as the size of the video composed of all the arranged images from the current selected frame to the next selected frame. After determining the size of the arranged video, store it in the Buffer (when the next selected frame is encountered, the video size in the Buffer needs to be updated) for the two-dimensional arrangement module. The two-dimensional arrangement module arranges the wavelet transform coefficients after quantization of each frame into a two-dimensional image in a certain order. The embodiments of the present application do not specify a specific arrangement order, as long as the encoding and decoding ends are consistent. For example, the Morton order can be adopted within the block and the raster scan order can be adopted between blocks; during the arrangement process, a feasible arrangement method is: if the sum of the number of blocks required for each LOD level of the wavelet transform coefficients after quantization of the current frame is less than or equal to the number of blocks in one frame of the arranged video (the number of blocks is equal to ), then fill the pixels in the extra blocks with default values (this patent does not specify the specific filling value); if the sum of the number of blocks required for each LOD level of the wavelet transform coefficients after quantization of the current frame is greater than the number of blocks in one frame of the arranged video, then remove the extra displacement data. At the same time, this module outputs a flag to indicate whether to enable the solution proposed in this patent.
[0157] ⑤ Video compression
[0158] After arranging the wavelet coefficients on the two-dimensional image, the video encoder can be directly used to encode them to obtain the displacement bitstream.
[0159] (7) Reconstruction of the deformed grid
[0160] When the displacement vector encoding module performs encoding, it needs to obtain the reconstructed displacement vector values. First, it obtains the decoded and reconstructed video from the video encoder. Then, in the same order as the two-dimensional arrangement order, it extracts pixel values from each frame image of the decoded and reconstructed video. If the flag indicator output by the two-dimensional arrangement module indicates that the solution proposed in this patent is enabled, a feasible extraction method is as follows: for each frame of the image to be extracted, determine the number of pixel values to be extracted according to the number of vertices of each LOD level of the subdivided reconstructed base grid corresponding to it; when the total number of blocks required for the number of vertices of each LOD level of the subdivided reconstructed base grid is less than or equal to the number of blocks of the displacement image to be extracted, extract the same number of pixel values as the number of vertices of the subdivided reconstructed base grid from the image to be extracted in the same order as the arrangement order at the encoding end; when the total number of blocks required for the number of vertices of each LOD level of the subdivided reconstructed base grid is greater than the number of blocks of the image to be extracted, extract all non-filled pixel values in the image to be extracted in the same order as the arrangement order at the encoding end, and perform a filling operation at the end of the extracted data until the number of data is equal to the number of vertices of the subdivided reconstructed base grid. There are various ways to determine the filled values, such as equal to 0, equal to the average of the extracted pixel values, etc. If the flag indicator output by the two-dimensional arrangement module indicates that the solution proposed in this patent is not enabled, directly extract the same number of pixel values as the number of vertices of the subdivided reconstructed base grid from each frame of the decoded and reconstructed video in the same order as the two-dimensional arrangement order. Finally, perform inverse quantization and inverse transformation on the extracted data to obtain the displacement vector consistent with the decoding end. After obtaining the reconstructed geometric displacement vector, the subdivided sub-grid is used to obtain the reconstructed subdivided deformed grid according to the corresponding geometric displacement vector, and it is passed to the attribute map conversion module.
[0161] (8) Texture map conversion
[0162] The texture map conversion module performs texture map conversion according to the input original grid, the input original texture map, and the reconstructed deformed grid. The texture map can also be called the attribute map.
[0163] The steps of attribute map conversion are as follows:
[0164] Calculate the texture coordinates of each pixel on the attribute map to be generated. For example, the texture coordinates corresponding to pixel A(i,j) are P(u,v).
[0165] Determine whether the texture coordinates are within a certain triangular face after parameterization of the subdivided deformed grid.
[0166] If the texture coordinates do not belong to any triangular face, mark the pixel as an empty pixel, and then it can be filled with a filling algorithm.
[0167] If the texture coordinates belong to a triangular face, then
[0168] Mark the pixel as filled.
[0169] Calculate the barycentric coordinates of the texture coordinates in the current triangular face.
[0170] According to the barycentric coordinates and the corresponding triangular face, map the two-dimensional texture coordinates to three-dimensional geometric coordinates, that is, map to the points on the subdivided and deformed mesh corresponding to the texture coordinates, as shown by M(x, y, z) in the figure.
[0171] Find the point on the input original mesh that is closest to the three-dimensional coordinates, as shown by M′(x, y, z) in the figure.
[0172] Calculate the barycentric coordinates of the three-dimensional coordinates according to the triangular face where it is located and map it to two dimensions, calculate its texture coordinates, that is, P′(u′, v′).
[0173] Sample through the texture coordinates on the input original attribute map to obtain the value A′(i′, j′) at the corresponding pixel position.
[0174] Assign this value to the corresponding pixel A(i, j) on the attribute map to be generated.
[0175] (9) Texture map compression
[0176] After obtaining the converted attribute map, for the empty pixels in it, existing filling algorithms (such as the Push-Pull algorithm) can be used to fill these empty pixels. Then, existing video encoders, such as H.264 / AVC, H.265 / HEVC, H.266 / VVC, etc., can be used to encode it to obtain the bitstream of the output attribute map. In addition, operations such as color space conversion and chrominance subsampling can be selectively applied to make the video encoding obtain better rate-distortion performance, such as the color space conversion from RGB 444 to YUV420. For the case where there are multiple attribute maps, auxiliary information is required to identify the texture coordinates corresponding to each attribute map during encoding.
[0177] (10) Auxiliary information
[0178] During the encoding process, there are some alternative schemes in each module, such as the type of mesh encoder, the type of video encoder, the mesh subdivision scheme, the spatial displacement vector transformation scheme, and transformation parameters, etc. The proposed framework allows the use of different schemes. Therefore, the selected scheme needs to be passed to the decoding end to guide correct decoding.
[0179] The auxiliary information includes the type of static mesh encoder, the type of video encoder, the mesh subdivision scheme, the number of mesh iterations, the geometric displacement vector transformation scheme and transformation parameters, the coefficient arrangement scheme, and the color conversion scheme, etc. For the mesh part encoded using the inter-frame coding mode, it also includes the corresponding reference frame list, etc.
[0180] After all modules are encoded, the base mesh part bitstream, texture coordinate part bitstream, displacement vector video bitstream, attribute map video bitstream, and auxiliary information bitstream are mixed to obtain the final encoded bitstream output at the encoding end.
[0181] The solution on the second electronic device side proposed in this application is mainly applied in the "displacement decoding" process at the decoding end of the VDMC framework, as Figure 9 shown. Next, a specific description of the entire application process is given, as Figure 9 shown.
[0182] (1) Auxiliary information decoding
[0183] The auxiliary information includes the type of static mesh encoder, the type of video encoder, the mesh subdivision scheme, the number of mesh iterations, the geometric displacement vector transformation scheme, the coefficient arrangement scheme, the displacement truncation flag, and the color conversion scheme, etc. For the mesh part encoded using the inter-frame coding mode, it also includes the corresponding reference frame list, etc. It can guide the decoding end to perform correct decoding, and this information is represented by the corresponding syntax parameters after decoding.
[0184] (2) Base mesh decoding
[0185] The decoding of the base mesh can be divided into the intra-frame mode and the inter-frame mode, etc. This module is only responsible for decoding the information representing the base mesh. That is, for the intra-frame mode, the output of the decoding of this module is a three-dimensional mesh, including geometric information, connection relationships, and possibly included attribute information and texture coordinate information, etc. For the inter-frame mode, this module is responsible for decoding the motion vector and identifier and possibly existing intra-frame sub-meshes, that is, the output information is the motion vector, identifier, and intra-frame sub-meshes.
[0186] ① Intra-frame mode
[0187] The decoding of the base mesh information in the intra-frame mode, that is, the decoding of the base mesh. This module uses the corresponding mesh decoder to decode the base mesh according to the type of static mesh decoder indicated by the auxiliary information.
[0188] ② Inter-frame mode
[0189] In the inter-frame mode, the motion vector, identifier, and possibly the intra sub-grid need to be decoded. For the intra sub-grid, it can be decoded using the grid decoder indicated by the auxiliary information. The motion vector identifier is obtained through entropy decoding, and after the motion vector is entropy decoded, it is obtained through inverse prediction and inverse quantization.
[0190] (3) Subdivision
[0191] The operation of this subdivision module is the same as the subdivision operation at the encoding end. The number of subdivision iterations in different regions of the current base grid is indicated by the auxiliary information, etc.
[0192] (4) Displacement Vector Video Decoding
[0193] The displacement vector video bitstream is decoded according to the video decoder type indicated in the auxiliary information. Note that this includes geometric displacement vector video and possibly displacement vector videos with different attributes.
[0194] (5) Displacement Vector Decoding
[0195] This module is responsible for restoring the decoded displacement vector image to a displacement vector. First, the pixel values need to be extracted from the displacement vector image. If the flag indicator output from the two-dimensional arrangement module at the encoding end to the decoding end indicates that the solution proposed in this patent is enabled, a feasible extraction method is as follows: for each frame of the image to be extracted, determine the number of pixel values to be extracted according to the number of vertices at each LOD level of the subdivided and reconstructed base grid; when the total number of blocks required for the number of vertices at each LOD level of the subdivided and reconstructed base grid is less than or equal to the number of blocks of the displacement image to be extracted, extract the same number of pixel values as the number of vertices of the subdivided and reconstructed base grid from the image to be extracted in the same order as the encoding end arrangement order; when the total number of blocks required for the number of vertices at each LOD level of the subdivided and reconstructed base grid is greater than the number of blocks of the image to be extracted, extract all non-filled pixel values in the image to be extracted in the same order as the encoding end arrangement order, and perform a filling operation at the end of the extracted data until the data quantity is equal to the number of vertices of the subdivided and reconstructed base grid. There are various ways to determine the filled values, such as being equal to 0, equal to the mean of the extracted pixel values, etc. If the flag indicator output from the two-dimensional arrangement module at the encoding end to the decoding end indicates that the solution proposed in this patent is not enabled, directly extract the same number of pixel values as the number of vertices of the subdivided and reconstructed base grid from each frame of the decoded and reconstructed video in the same order as the two-dimensional arrangement order. Then, perform operations such as inverse quantization and inverse transformation on the extracted data to restore the corresponding displacement vector.
[0196] (6) Reconstruct the Deformed Grid
[0197] After the base mesh and the displacement vector decoding and reconstruction are completed, the deformed mesh is reconstructed based on these two parts. Add the corresponding displacement vector to each vertex of the subdivided mesh, as shown in Equation (11):
[0198] deformedmesh[i].v[k] = subdivmesh[i].v[k] + displacement[k] #(11)
[0199] where subdivmesh[i].v[k] is the geometric coordinate of the k-th vertex after the base mesh subdivision of the current frame (indexed by i), displacement[k] is the spatial displacement vector corresponding to the k-th vertex, and deformedmesh[i].v[k] is the geometric coordinate of the k-th vertex after the current frame subdivision and deformation.
[0200] The attribute displacement vector is applied to the corresponding attribute value in the same way. Taking texture coordinates as an example, as shown in Equation (12).
[0201] deformedmesh[i].vt[k] = subdivmesh[i].vt[k] + attdisplacement[k] #(12)
[0202] where subdivmesh[i].vt[k] represents the k-th texture coordinate after the base mesh subdivision of the current frame, attdisplacement[k] represents the k-th value of the attribute displacement vector, and deformedmesh[i].vt[k] represents the value of the k-th texture coordinate after the current frame subdivision and deformation.
[0203] (7) Attribute map decoding
[0204] The attribute map decoder is responsible for decoding the attribute map bitstream. The attribute map is decoded using the video decoder indicated in the auxiliary information. An optional color space conversion is performed on it to obtain an image format consistent with the input attribute at the encoding end, and the finally decoded output attribute map is obtained. For multiple attribute map bitstreams, after decoding, each type of attribute map is corresponded to the corresponding texture coordinates according to the identifier. Figure 1 After each module is processed, the deformed mesh reconstructed at the decoding end and the corresponding attribute map are finally obtained. Subsequent applications will use the reconstructed deformed mesh and attribute map as inputs for processing.
[0205] After each module is processed, the deformed mesh reconstructed at the decoding end and the corresponding attribute map are finally obtained. Subsequent applications will use the reconstructed deformed mesh and attribute map as inputs for processing.
[0206] In an alternative implementation, when the solutions on the first electronic device side and the methods on the second electronic device side proposed in this application are applied to the VDMC framework, the corresponding syntax structure of the VDMC framework needs to be enhanced. For example, a parameter "asve_custom_displacement_video_size_flag" is added to the syntax structure corresponding to the VDMC framework. When asve_custom_displacement_video_size_flag is equal to 1, it means that the method mentioned in this application is enabled in the VDMC; when it is equal to 0, it means that the method mentioned in this application is not enabled in the VDMC.
[0207] For the displacement information video generation method and decoding method provided in the embodiments of this application, the execution subject can be a displacement information video generation device or a decoding device. In the embodiments of this application, taking the displacement information video generation method device or the decoding device as an example to execute the displacement information video generation method or the decoding method, the displacement information video generation device or the decoding device provided in the embodiments of this application is described.
[0208] As Figure 10 shown, the embodiments of this application also provide a displacement information video generation device, which is applied to a first electronic device and includes:
[0209] A first processing module 1001, configured to process multiple frames of displacement data to obtain multiple frames of first displacement data;
[0210] A first determination module 1002, configured to determine the size of the displacement video corresponding to the multiple frames of displacement data;
[0211] A first conversion module 1003, configured to convert the multiple frames of first displacement data into a displacement video according to the size of the displacement video.
[0212] As an alternative embodiment, the first determination module includes:
[0213] A first determination sub-module, configured to specify the size of the displacement video;
[0214] Or,
[0215] A second determination sub-module, configured to determine the size of the displacement video according to the size of the target frame displacement image.
[0216] As an alternative embodiment, the device further includes:
[0217] A second determination module, configured to determine the size of the target frame displacement image according to the number of data blocks in each row of the displacement video, the size of the data blocks in the displacement video, and the number of target frame displacement data.
[0218] As an alternative embodiment, the second determination module includes:
[0219] A third determination sub-module, configured to determine the width of the target frame displacement image according to the number of data blocks in each row of the displacement video and the size of the data blocks in the displacement video;
[0220] A fourth determination sub-module, configured to determine the height of the target frame displacement image according to the width of the target frame displacement image and the number of target frame displacement data.
[0221] As an optional embodiment, the third determination sub-module includes:
[0222] A first determination unit, configured to determine the width of the target frame displacement image according to a first formula; wherein, the first formula is:
[0223] Width = VideoWidthInBlocks * VideoBlockSize
[0224] Wherein, Width represents the width of the target frame displacement image; VideoWidthInBlocks represents the number of data blocks in each row of the displacement video; VideoBlockSize represents the size of the data blocks in the displacement video.
[0225] As an optional embodiment, the fourth determination sub-module includes:
[0226] A second determination unit, configured to determine the height of the target frame displacement image according to a second formula; wherein, the second formula is:
[0227]
[0228] Wherein, Heigth represents the height of the target frame displacement image; NumOfDisplacements represents the number of target frame displacement data; Width represents the width of the target frame displacement image; ceil represents the ceiling operation.
[0229] As an optional embodiment, the conversion module includes:
[0230] A conversion sub-module, configured to arrange each frame of the first displacement data in a predetermined order into a displacement image according to the size of the displacement video; multiple displacement images constitute the displacement video.
[0231] As an optional embodiment, the device further includes:
[0232] A filling module, configured to fill the extra pixels with a default value if the number of displacement data of the current frame is less than the size of the displacement video;
[0233] Or,
[0234] A removal module, configured to remove redundant displacement data if the number of displacement data of the current frame is greater than the size of the displacement video.
[0235] As an optional embodiment, the apparatus further includes:
[0236] A third determination module, configured to determine whether the displacement data of the current frame is target frame displacement data according to a first manner; wherein the first manner includes at least one of the following:
[0237] Presetting the displacement data of the Mth frame as target frame displacement data;
[0238] Dividing multiple frames of displacement data into multiple groups of pictures (GOPs) in advance, and setting the displacement data of the Nth frame of each GOP as target frame displacement data;
[0239] wherein M and N are respectively integers greater than or equal to 1.
[0240] As an optional embodiment, the apparatus further includes:
[0241] An encoding module, configured to encode the displacement video to obtain a displacement bitstream;
[0242] A sending module, configured to send the displacement bitstream to a second electronic device.
[0243] In the embodiments of the present application, multiple frames of displacement data are converted into a displacement video, and the size of the displacement video is determined before generating the displacement video, so as to improve the flexibility and real-time performance of subsequent video encoding.
[0244] It should be noted that the displacement information video generation apparatus provided in the embodiments of the present application is a device capable of executing the above-mentioned displacement information video generation method. All embodiments of the above-mentioned displacement information video generation method are applicable to the device, and all can achieve the same or similar beneficial effects, which will not be repeated here.
[0245] As Figure 11 shown, the embodiments of the present application further provide a decoding device, which is applied to a second electronic device and includes:
[0246] An obtaining module 1101, configured to obtain a displacement bitstream and decode the displacement bitstream into a displacement video;
[0247] An extraction module 1102, configured to extract pixel values from each frame of displacement image of the displacement video according to a target order to obtain target data; the target order is the same as the order in which the first electronic device arranges the displacement images;
[0248] A second processing module 1103, configured to process the extracted target data to obtain multiple frames of displacement data.
[0249] As an alternative embodiment, the extraction module includes:
[0250] A sixth determination sub-module, configured to determine the number of pixel values to be extracted from a displacement image according to the number of vertices of the subdivided reconstructed basic grid corresponding to the displacement image for a frame of displacement image;
[0251] A first extraction sub-module, configured to, if the number of vertices of the subdivided reconstructed basic grid is less than or equal to the number of pixel values to be extracted from the displacement image, extract pixel values equal to the number of vertices of the subdivided reconstructed basic grid from the displacement image in the target order to obtain target data;
[0252] A second extraction sub-module, configured to, if the number of vertices of the subdivided reconstructed basic grid is greater than the number of pixel values to be extracted from the displacement image, extract all pixel values from the displacement image in the target order, and perform a filling operation at the end of the extracted data until the number of data is equal to the number of vertices of the subdivided reconstructed basic grid, to obtain target data.
[0253] For the solution of converting multiple frames of displacement data into a displacement video and determining the size of the displacement video before generating the displacement video, the embodiments of the present application provide a corresponding decoding method, thereby improving the decoding correctness.
[0254] It should be noted that the decoding device provided by the embodiments of the present application is a device capable of executing the above decoding method. All embodiments of the above decoding method are applicable to this device and can achieve the same or similar beneficial effects, which will not be repeated here.
[0255] The displacement information video generation device or decoding device in the embodiments of the present application may be an electronic device, such as an electronic device with an operating system, or a component in an electronic device, such as an integrated circuit or a chip. The electronic device may be a terminal or other devices other than a terminal. Exemplarily, the terminal may include, but is not limited to, the types of the above-mentioned terminal 11, and other devices may be a server, a Network Attached Storage (NAS), etc., which are not specifically limited in the embodiments of the present application.
[0256] The displacement information video generation device or decoding device provided by the embodiments of the present application can implement Figures 2 to 9 each process implemented by the method embodiments and achieve the same technical effects. To avoid repetition, it will not be elaborated here.
[0257] Such as Figure 12As shown, an embodiment of the present application further provides a communication device 1200, including a processor 1201 and a memory 1202. A program or instruction that can run on the processor 1201 is stored on the memory 1202. For example, when the communication device 1200 is a first electronic device, when the program or instruction is executed by the processor 1201, each step of the above-described embodiment of the displacement information video generation method is implemented, and the same technical effect can be achieved. When the communication device 1200 is a second electronic device, when the program or instruction is executed by the processor 1201, each step of the above-described embodiment of the decoding method is implemented, and the same technical effect can be achieved. To avoid repetition, details are not described here again.
[0258] In the case where the first electronic device is a terminal, or the second electronic device is a terminal, an embodiment of the present application further provides a terminal, including a processor and a communication interface. The communication interface is coupled to the processor, and the processor is configured to run a program or instruction to implement steps in the method embodiment as Figure 10 or Figure 11 shown. This terminal embodiment corresponds to the above-described terminal-side method embodiment. Each implementation process and implementation manner of the above method embodiment can be applied to this terminal embodiment, and the same technical effect can be achieved. Specifically, Figure 13 FIG. is a schematic hardware structure diagram of a terminal for implementing an embodiment of the present application.
[0259] The terminal 1300 includes, but is not limited to, at least some components such as a radio frequency unit 1301, a network module 1302, an audio output unit 1303, an input unit 1304, a sensor 1305, a display unit 1306, a user input unit 1307, an interface unit 1308, a memory 1309, and a processor 1310.
[0260] Those skilled in the art can understand that the terminal 1300 may further include a power supply (such as a battery) for supplying power to each component. The power supply may be logically connected to the processor 1310 through a power management system, so as to implement functions such as management of charging, discharging, and power consumption management through the power management system. Figure 13 The terminal structure shown in does not limit the terminal. The terminal may include more or fewer components than shown, or combine certain components, or have different component arrangements, which are not described here again.
[0261] It should be understood that in the embodiments of the present application, the input unit 1304 may include a Graphics Processing Unit (GPU) 13041 and a microphone 13042. The graphics processor 13041 processes the image data of still pictures or videos obtained by an image capturing device (such as a camera) in a video capture mode or an image capture mode. The display unit 1306 may include a display panel 13061, and the display panel 13061 may be configured in the form of, for example, a liquid crystal display, an organic light emitting diode, etc. The user input unit 1307 includes at least one of a touch panel 13071 and other input devices 13072. The touch panel 13071 is also referred to as a touch screen. The touch panel 13071 may include two parts: a touch detection device and a touch controller. The other input devices 13072 may include, but are not limited to, a physical keyboard, function keys (such as volume control keys, switch keys, etc.), a trackball, a mouse, and a joystick, which will not be elaborated here.
[0262] In the embodiments of the present application, after receiving downlink data from a network-side device, the radio frequency unit 1301 may transmit it to the processor 1310 for processing; in addition, the radio frequency unit 1301 may send uplink data to the network-side device. Generally, the radio frequency unit 1301 includes, but is not limited to, an antenna, an amplifier, a transceiver, a coupler, a low noise amplifier, a duplexer, etc.
[0263] The memory 1309 can be used to store software programs or instructions as well as various data. The memory 1309 may mainly include a first storage area for storing programs or instructions and a second storage area for storing data. Among them, the first storage area may store an operating system, application programs or instructions required for at least one function (such as a sound playback function, an image playback function, etc.). In addition, the memory 1309 may include volatile memory or non-volatile memory. Among them, the non-volatile memory may be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory may be a random access memory (RAM), a static random access memory (SRAM), a dynamic random access memory (DRAM), a synchronous dynamic random access memory (SDRAM), a double data rate synchronous dynamic random access memory (DDR SDRAM), an enhanced synchronous dynamic random access memory (ESDRAM), a synchronous link dynamic random access memory (SLDRAM), and a direct rambus random access memory (DRRAM). The memory 1309 in the embodiments of the present application includes but is not limited to these and any other suitable types of memory.
[0264] The processor 1310 may include one or more processing units; optionally, the processor 1310 integrates an application processor and a modem processor. Among them, the application processor mainly processes operations related to the operating system, user interface, and application programs, etc., and the modem processor mainly processes wireless communication signals, such as a baseband processor. It can be understood that the above modem processor may not be integrated into the processor 1310 either.
[0265] Among them, the radio frequency unit 1301 is used to receive multiple frames of displacement data;
[0266] The processor 1310 is used to process the multiple frames of displacement data to obtain multiple frames of first displacement data; determine the size of the displacement video corresponding to the multiple frames of displacement data; and convert the multiple frames of first displacement data into a displacement video according to the size of the displacement video.
[0267] In the embodiments of the present application, multiple frames of displacement data are converted into a displacement video, and the size of the displacement video is determined before generating the displacement video, thereby improving the flexibility and real-time performance of subsequent video encoding.
[0268] It should be noted that the displacement information video generation device provided in the embodiments of the present application is a device capable of executing the above-mentioned displacement information video generation method. All embodiments of the above-mentioned displacement information video generation method are applicable to the device, and all can achieve the same or similar beneficial effects, which will not be repeated here.
[0269] It can be understood that the implementation processes of the various implementation manners mentioned in this embodiment can refer to the relevant descriptions of the method embodiments, and achieve the same or corresponding technical effects. To avoid repetition, they will not be elaborated here.
[0270] When the first electronic device is a network-side device or the second electronic device is a network-side device, the embodiments of the present application further provide a network-side device, including a processor and a communication interface. The communication interface is coupled to the processor, and the processor is used to run programs or instructions to implement as Figure 10 or the steps of the method embodiments shown in 11. This network-side device embodiment corresponds to the above-mentioned network-side device method embodiments. All implementation processes and implementation manners of the above-mentioned method embodiments are applicable to this network-side device embodiment, and can achieve the same technical effects.
[0271] Specifically, the embodiments of the present application further provide a network-side device. As Figure 14 shown, the network-side device 1400 includes: an antenna 141, a radio frequency device 142, a baseband device 143, a processor 144, and a memory 145. The antenna 141 is connected to the radio frequency device 142. In the uplink direction, the radio frequency device 142 receives information through the antenna 141 and sends the received information to the baseband device 143 for processing. In the downlink direction, the baseband device 143 processes the information to be sent and sends it to the radio frequency device 142. The radio frequency device 142 processes the received information and then sends it out through the antenna 141.
[0272] The methods executed by the network-side device in the above embodiments can be implemented in the baseband device 143, and the baseband device 143 includes a baseband processor.
[0273] The baseband device 143 may include, for example, at least one baseband board, and multiple chips are provided on the baseband board. As Figure 14 shown, one of the chips is, for example, a baseband processor, which is connected to the memory 145 through a bus interface to call the programs in the memory 145 and execute the operations of the network device shown in the above method embodiments.
[0274] The network-side device may further include a network interface 146, such as a Common Public Radio Interface (CPRI).
[0275] Specifically, the network-side device 1400 in the embodiments of the present invention further includes: instructions or programs stored in the memory 145 and executable on the processor 144. The processor 144 calls the instructions or programs in the memory 145 to execute Figure 10 the methods executed by the modules shown in FIGS. 10 or 11, and achieve the same technical effects. To avoid repetition, they will not be elaborated here.
[0276] The embodiments of the present application further provide a readable storage medium, on which programs or instructions are stored. When the programs or instructions are executed by a processor, they implement each process of the above-mentioned displacement information video generation method or decoding method embodiments, and can achieve the same technical effects. To avoid repetition, they will not be elaborated here.
[0277] Wherein, the processor is the processor in the terminal described in the above embodiments. The readable storage medium includes computer-readable storage media, such as computer read-only memory ROM, random access memory RAM, magnetic disks, or optical discs, etc. In some examples, the readable storage medium may be a non-transitory readable storage medium.
[0278] The embodiments of the present application further provide a chip, which includes a processor and a communication interface. The communication interface is coupled to the processor. The processor is used to run programs or instructions to implement each process of the above-mentioned displacement information video generation method or decoding method embodiments, and can achieve the same technical effects. To avoid repetition, they will not be elaborated here.
[0279] It should be understood that the chip mentioned in the embodiments of the present application may also be referred to as a system-on-chip, system chip, chip system, or system-on-chip.
[0280] The embodiments of the present application further provide a computer program / program product, which is stored in a storage medium. The computer program / program product is executed by at least one processor to implement each process of the above-mentioned displacement information video generation method or decoding method embodiments, and can achieve the same technical effects. To avoid repetition, they will not be elaborated here.
[0281] The embodiments of the present application further provide a communication system, including: a terminal and a network-side device. The terminal can be used to execute the steps of the above-mentioned displacement information video generation method, and the network-side device can be used to execute the steps of the above-mentioned decoding method.
[0282] It should be noted that in this text, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, such that a process, method, article or device comprising a series of elements not only includes those elements but also other elements not expressly listed, or further includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising a..." does not exclude the presence of additional identical elements in the process, method, article or device comprising that element. In addition, it should be pointed out that the scope of the methods and devices in the embodiments of the present application is not limited to performing functions in the order shown or discussed, and may also include performing functions in a substantially simultaneous manner or in the reverse order according to the functions involved. For example, the described methods may be performed in an order different from that described, and various steps may be added, omitted or combined. Additionally, the features described with reference to certain examples may be combined in other examples.
[0283] Through the description of the above embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of a computer software product plus a necessary general hardware platform, and of course, can also be implemented by hardware. The computer software product is stored in a storage medium (such as ROM, RAM, magnetic disk, optical disk, etc.) and includes several instructions for causing a terminal or a network-side device to execute the methods described in various embodiments of the present application.
[0284] The embodiments of the present application have been described above in conjunction with the accompanying drawings. However, the present application is not limited to the above specific embodiments. The above specific embodiments are merely illustrative and not restrictive. Those of ordinary skill in the art, under the inspiration of the present application and without departing from the purpose of the present application and the scope protected by the claims, can also make many forms of embodiments, and these embodiments are all within the protection scope of the present application.
Claims
1. A method for generating a displacement information video, characterized in that, Including: The first electronic device processes multiple frames of displacement data to obtain multiple frames of first displacement data; The first electronic device determines the size of the displacement video corresponding to the multiple frames of displacement data; The first electronic device converts the multiple frames of first displacement data into a displacement video according to the size of the displacement video.
2. The method according to claim 1, wherein The first electronic device determining the size of the displacement video corresponding to the multiple frames of displacement data includes: The first electronic device designates the size of the displacement video; Or, The first electronic device determines the size of the displacement video according to the size of the target frame displacement image.
3. The method according to claim 2, wherein The method further includes: The first electronic device determines the size of the target frame displacement image according to the number of data blocks in each row of the displacement video, the size of the data blocks in the displacement video, and the number of target frame displacement data.
4. The method according to claim 3, characterized in that, The first electronic device determining the size of the target frame displacement image according to the number of data blocks in each row of the displacement video, the size of the data blocks in the displacement video, and the number of target frame displacement data includes: Determining the width of the target frame displacement image according to the number of data blocks in each row of the displacement video and the size of the data blocks in the displacement video; Determining the height of the target frame displacement image according to the width of the target frame displacement image and the number of target frame displacement data.
5. The method according to claim 4, wherein Determining the width of the target frame displacement image according to the number of data blocks in each row of the displacement video and the size of the data blocks in the displacement video includes: Determining the width of the target frame displacement image according to the first formula; where the first formula is: Width = VideoWidthInBlocks * VideoBlockSize Where, Width represents the width of the target frame displacement image; VideoWidthInBlocks represents the number of data blocks in each row of the displacement video; VideoBlockSize represents the size of the data blocks in the displacement video.
6. The method according to claim 4, characterized in that Determining the height of the target frame displacement image according to the width of the target frame displacement image and the number of target frame displacement data includes: Determining the height of the target frame displacement image according to the second formula; where the second formula is: Where, Heigth represents the height of the target frame displacement image; NumOfDisplacements represents the number of target frame displacement data; Width represents the width of the target frame displacement image; ceil represents the ceiling operation.
7. The method according to any one of claims 1-6, characterized in that, The first electronic device converting the multiple frames of first displacement data into a displacement video according to the size of the displacement video includes: The first electronic device arranges each frame of first displacement data in a predetermined order into a displacement image according to the size of the displacement video; multiple displacement images constitute a displacement video.
8. The method according to claim 7, wherein The method further includes: If the number of displacement data of the current frame is less than the size of the displacement video, fill the extra pixels with default values; Or, If the number of displacement data of the current frame is greater than the size of the displacement video, remove the extra displacement data.
9. The method according to any one of claims 3-6, characterized in that The method further includes: Determining whether the current frame displacement data is target frame displacement data in the first manner; where the first manner includes at least one of the following: Presetting the Mth frame of displacement data as the target frame displacement data; Pre-divide multiple frames of displacement data into multiple groups of pictures (GOPs), and set the displacement data of the Nth frame in each GOP as the target frame displacement data; where M and N are integers greater than or equal to 1 respectively.
10. The method according to any one of claims 1-9, characterized in that The method further includes: The first electronic device encodes the displacement video to obtain a displacement bitstream; The first electronic device sends the displacement bitstream to the second electronic device.
11. A decoding method, characterized in that, It includes: The second electronic device obtains the displacement bitstream and decodes the displacement bitstream into a displacement video; The second electronic device extracts pixel values from each frame of displacement image in the displacement video according to the target order to obtain target data; the target order is the same as the order in which the first electronic device arranges the displacement images; The second electronic device processes the extracted target data to obtain multiple frames of displacement data.
12. The method according to claim 11, characterized in that, The second electronic device extracts pixel values from each frame of displacement image in the displacement video according to the target order to obtain target data, including: For a frame of displacement image, determine the number of pixel values to be extracted from the displacement image according to the number of subdivided and reconstructed basic grid vertices corresponding to the displacement image; If the number of subdivided and reconstructed basic grid vertices is less than or equal to the number of pixel values to be extracted from the displacement image, extract the same number of pixel values as the number of subdivided and reconstructed basic grid vertices from the displacement image according to the target order to obtain target data; Or, if the number of subdivided and reconstructed basic grid vertices is greater than the number of pixel values to be extracted from the displacement image, extract all pixel values from the displacement image according to the target order, and perform a filling operation at the end of the extracted data until the number of data is equal to the number of subdivided and reconstructed basic grid vertices to obtain target data.
13. A displacement information video generation device, which is applied to a first electronic device, is characterized in that It includes: A first processing module for processing multiple frames of displacement data to obtain multiple frames of first displacement data; A first determination module for determining the size of the displacement video corresponding to the multiple frames of displacement data; A first conversion module for converting the multiple frames of first displacement data into a displacement video according to the size of the displacement video.
14. The device according to claim 13, characterized in that, The first determination module includes: A first determination sub-module for specifying the size of the displacement video; Or, A second determination sub-module for determining the size of the displacement video according to the size of the target frame displacement image.
15. The device according to claim 14, characterized in that, The device further includes: A second determination module for determining the size of the target frame displacement image according to the number of data blocks in each row of the displacement video, the size of the data blocks in the displacement video, and the number of target frame displacement data.
16. The device according to claim 15, characterized in that, The second determination module includes: A third determination sub-module for determining the width of the target frame displacement image according to the number of data blocks in each row of the displacement video and the size of the data blocks in the displacement video; A fourth determination sub-module for determining the height of the target frame displacement image according to the width of the target frame displacement image and the number of target frame displacement data.
17. The device according to claim 16, characterized in that, The third determination sub-module includes: A first determination unit for determining the width of the target frame displacement image according to the first formula; where the first formula is: Width = VideoWidthInBlocks * VideoBlockSize Wherein, Width represents the width of the target frame displacement image; VideoWidthInBlocks represents the number of data blocks in each row of the displacement video; VideoBlockSize represents the size of the data blocks in the displacement video.
18. The device according to claim 16, wherein The fourth determination sub-module includes: A second determination unit, configured to determine the height of the target frame displacement image according to a second formula; wherein, the second formula is: Wherein, Heigth represents the height of the target frame displacement image; NumOfDisplacements represents the number of target frame displacement data; Width represents the width of the target frame displacement image; ceil represents the ceiling operation.
19. The device according to any one of claims 13 - 18, characterized in that The conversion module includes: A conversion sub-module, configured to arrange each frame of the first displacement data into a displacement image in a predetermined order according to the size of the displacement video; multiple frames of displacement images constitute the displacement video.
20. The device according to claim 19, characterized in that, The apparatus further includes: A padding module, configured to, if the number of displacement data of the current frame is less than the size of the displacement video, pad the extra pixels with a default value; Or, A removal module, configured to, if the number of displacement data of the current frame is greater than the size of the displacement video, remove the redundant displacement data.
21. The device according to any one of claims 15-18, characterized in that, The apparatus further includes: A third determination module, configured to determine whether the current frame displacement data is target frame displacement data in a first manner; wherein, the first manner includes at least one of the following: Preset the Mth frame of displacement data as the target frame displacement data in advance; Divide multiple frames of displacement data into multiple image groups GOP in advance, and set the Nth frame of displacement data of each GOP as the target frame displacement data; Wherein, M and N are integers greater than or equal to 1 respectively.
22. The device according to any one of claims 13-21, characterized in that, The apparatus further includes: An encoding module, configured to encode the displacement video to obtain a displacement bitstream; A sending module, configured to send the displacement bitstream to a second electronic device.
23. A decoding device, applied to a second electronic device, characterized in that, Includes: An acquisition module, configured to acquire a displacement bitstream, and decode the displacement bitstream into a displacement video; An extraction module, configured to extract pixel values from each frame of displacement image of the displacement video in a target order to obtain target data; the target order is the same as the order in which the first electronic device arranges the displacement images; A second processing module, configured to process the extracted target data to obtain multiple frames of displacement data.
24. The device according to claim 23, characterized in that, The extraction module includes: A sixth determination sub-module, configured to, for a frame of displacement image, determine the number of pixel values to be extracted from the displacement image according to the number of subdivided reconstructed basic grid vertices corresponding to the displacement image; A first extraction sub-module, configured to, if the number of subdivided reconstructed basic grid vertices is less than or equal to the number of pixel values to be extracted from the displacement image, extract the same number of pixel values as the number of subdivided reconstructed basic grid vertices from the displacement image in the target order to obtain target data; A second extraction sub-module, configured to, if the number of subdivided reconstructed basic grid vertices is greater than the number of pixel values to be extracted from the displacement image, extract all pixel values from the displacement image in the target order, and perform a filling operation at the end of the extracted data until the data quantity is equal to the number of subdivided reconstructed basic grid vertices, to obtain target data.
25. A first electronic device, characterized in that, It includes a processor and a memory, and the memory stores programs or instructions that can run on the processor. When the programs or instructions are executed by the processor, the steps of the displacement information video generation method according to any one of claims 1 to 10 are implemented.
26. A second electronic device, characterized in that, It includes a processor and a memory, and the memory stores programs or instructions that can run on the processor. When the programs or instructions are executed by the processor, the steps of the decoding method according to claim 11 or 12 are implemented.
27. A readable storage medium, characterized in that, Programs or instructions are stored on the readable storage medium. When the programs or instructions are executed by a processor, the steps of the displacement information video generation method according to any one of claims 1 to 10 are implemented, or the steps of the decoding method according to claim 11 or 12 are implemented.