Methods and Devices for Compressing and Decompressing Model Data
By selecting key perspectives from multiple shooting perspectives, reconstructing and rendering sparse models, and compressing them images, the problem of large amount of point cloud data in three-dimensional reconstruction technology is solved, and the resource consumption is reduced and the model is efficiently deployed.
Patent Information
- Application Number
- CN202510033925.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-08
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2045-01-08
AI Technical Summary
In the existing three-dimensional reconstruction technology, the amount of point cloud data is large, making it difficult for the model to be deployed on the end-side device and the storage and transmission resource consumption is high.
By determining N key perspectives in M shooting angles, using projection and back projection at key perspectives to obtain the first sparse model, and then rendering its attribute information to obtain the second sparse model. Then, the second sparse model is projected and image compression is generated to generate compressed data.
While ensuring the scenario representation capability, the consumption of storage and transmission resource of model data is significantly reduced, so that the model can be easily deployed on the end-side device.
Smart Images

Figure CN119444879B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of three-dimensional graphics and computer vision, and in particular to a method and device for compressing and decompressing model data. Background Art
[0002] With the rapid development of science and technology, people have put forward higher requirements for high-quality 3D reconstruction and real-time rendering experience of scenes. 3D reconstruction technology can restore the 3D geometry and appearance information of objects or scenes from 2D data. Point cloud is a commonly used data representation form in 3D reconstruction technology, and the reconstructed 3D model can be composed of point clouds. Point cloud is composed of a large number of points, each of which contains information about an object or scene in 3D space. However, in the application of 3D reconstruction technology, the number of point clouds representing 3D models is too large and the representation parameters are too large, making it difficult to directly deploy the reconstructed model on the end-side device, and also puts forward higher requirements in terms of model storage and transmission.
[0003] For example, 3D Gaussian Splatting (3DGS) reconstruction technology has shown potential for application in multiple fields. 3DGS reconstruction technology can achieve low-cost and high-quality scene reconstruction by collecting a scene video or image and combining it with the implicit discrete Gaussian point cloud representation of the scene. However, during the application of this technology, there will be redundancy in the representation of 3D Gaussian point clouds, making it difficult to directly deploy the scene model reconstructed based on 3DGS reconstruction technology on the end-side device, and it will also consume more resources in terms of model storage and transmission.
[0004] Therefore, how to compress model data in three-dimensional reconstruction technology is a technical problem that needs to be solved urgently.
[0005] The content of the background technology section is only the information known to the inventor personally, and does not mean that the above information has entered the public domain before the application date of this disclosure, nor does it mean that it can become the prior art of the present disclosure. Summary of the invention
[0006] The present specification provides a method and device for compressing and decompressing model data, which can compress model data while ensuring the scene representation capability of the model, thereby reducing the resources consumed when storing and transmitting the model data.
[0007] In a first aspect, this specification provides a method for compressing model data. In this compression method, after obtaining a target model to be compressed, a compression device determines N key viewpoints among M shooting viewpoints, and obtains a first sparse model by projecting and back-projecting the position information of each point cloud in the target model at the N key viewpoints. Then, the attribute information of each point cloud in the first sparse model is rendered to obtain a second sparse model, so that the scene representation ability of the second sparse model approaches that of the target model. The position information and attribute information of each point cloud in the second sparse model are projected at the N key viewpoints to obtain a first set of projected images, and each projected image in the first set of projected images is respectively subjected to image compression to obtain a second set of projected images, and the second set of projected images is used as the compressed data corresponding to the target model.
[0008] In some embodiments, determining N key viewpoints among the M shooting viewpoints includes the following steps: Based on the amount of information respectively included in the shooting images at the M shooting viewpoints, select the viewpoints whose included information meets the target condition from the M shooting viewpoints as the N key viewpoints.
[0009] In some embodiments, the viewpoint difference between two adjacent key viewpoints among the N key viewpoints is greater than a preset threshold.
[0010] In some embodiments, obtaining a first sparse model by projecting and back-projecting the position information of each point cloud in the target model at the N key viewpoints includes the following steps: Projecting the position information of each point cloud in the target model at the N key viewpoints to obtain N projected images; and back-projecting the N projected images at the N key viewpoints to obtain the first sparse model.
[0011] In some embodiments, projecting the position information of each point cloud in the target model at the N key viewpoints to obtain N projected images includes the following steps: For each key viewpoint among the N key viewpoints, based on the camera observation matrix corresponding to the key viewpoint, project the position information of each point cloud in the target model onto the camera plane corresponding to the key viewpoint to obtain a projected image corresponding to the key viewpoint.
[0012] In some embodiments, back-projecting the N projected images at the N key viewpoints to obtain the first sparse model includes the following steps: For each key viewpoint among the N key viewpoints, based on the camera observation matrix corresponding to the key viewpoint, back-project each pixel point in the projected image corresponding to the key viewpoint into three-dimensional space to obtain a point cloud subset corresponding to the key viewpoint; and generating the first sparse model based on the point cloud subsets respectively corresponding to the N key viewpoints.
[0013] In some embodiments, rendering the attribute information of each point cloud in the first sparse model to obtain a second sparse model includes the following steps: using the captured images at the N key viewpoints as supervision, rendering and training the attribute information of each point cloud in the first sparse model to obtain a second sparse model, where, during the rendering and training process, the following training objective is adopted: minimizing the difference between the projected images of the second sparse model at the N key viewpoints and the captured images at the N key viewpoints.
[0014] In some embodiments, the representation dimension of the attribute information of the point clouds in the second sparse model is higher than the representation dimension of the attribute information of the point clouds in the target model.
[0015] In some embodiments, each projected image in the first set of projected images corresponds to an attribute dimension, and the image compression process for each projected image includes the following steps: based on the attribute dimension corresponding to the projected image, determining the compression parameters applicable to compressing the projected image, and performing image compression processing on the projected image through the compression parameters, where the compression parameters determined for different attribute dimensions are different.
[0016] In some embodiments, the step of based on the attribute dimension corresponding to the projected image, determining the compression parameters applicable to compressing the projected image, and performing image compression processing on the projected image through the compression parameters includes the following steps: inputting the projected image and the attribute dimension corresponding to the projected image into a pre-trained adaptive image compression network to perform image compression processing on the projected image through the adaptive image compression network, where the adaptive image compression network is trained to perform image compression on projected images with different attribute dimensions using different compression parameters.
[0017] In a second aspect, this specification also provides a compression device for model data. The compression device includes at least one storage medium and at least one processor. The at least one storage medium stores at least one instruction set for performing data compression. The at least one processor is communicatively connected to the at least one storage medium, where, when the compression device runs, the at least one processor reads the at least one instruction set and executes the method according to any one of the first aspect as instructed by the at least one instruction set.
[0018] Thirdly, this specification also provides a method for decompressing model data. The decompression method includes the following steps: obtaining compressed data corresponding to a target model, where the target model is a model obtained by three-dimensional reconstruction of captured images from M captured perspectives, the target model includes position information and attribute information of a plurality of point clouds, M is an integer greater than 2, and the compressed data is obtained by data compression of the target model by the method described in any one of the first aspects; obtaining a second set of projection images from the compressed data, the second set of projection images includes compressed projection images corresponding to N key perspectives, the N key perspectives are partial perspectives among the M captured perspectives, and N is an integer greater than 1; performing image decompression processing on each projection image in the second set of projection images to obtain a first set of projection images; and back-projecting the first set of projection images at the N key perspectives to obtain decompressed data corresponding to the target model.
[0019] In some embodiments, each projection image in the second set of projection images corresponds to an attribute dimension, and the decompression process corresponding to each projection image includes: determining decompression parameters applicable to decompressing the projection image for the attribute dimension corresponding to the projection image, and performing image decompression processing on the projection image by the decompression parameters, where the decompression parameters determined for different attribute dimensions are different.
[0020] Fourthly, this specification also provides a model data decompression device, including at least one storage medium and at least one processor. The at least one storage medium stores at least one instruction set for performing data decompression; the at least one processor is communicatively connected to the at least one storage medium, where when the decompression device runs, the at least one processor reads the at least one instruction set and executes the method described in any one of the third aspects according to the instructions of the at least one instruction set.
[0021] As can be seen from the above technical solutions, the method for compressing model data provided in this specification can, after obtaining the target model to be compressed, determine N key viewpoints among M shooting viewpoints, and obtain the first sparse model by projecting and back-projecting the position information of each point cloud in the target model at the N key viewpoints. Then, render the attribute information of each point cloud in the first sparse model to obtain the second sparse model, so that the scene representation ability of the second sparse model approaches that of the target model. Furthermore, project the position information and attribute information of each point cloud in the second sparse model at the N key viewpoints to obtain the first set of projection images, and perform image compression on each projection image in the first set of projection images to obtain the second set of projection images, and use the second set of projection images as the compressed data corresponding to the target model. The above method selects some key viewpoints from multiple shooting viewpoints and uses the position projection information at the key viewpoints to reconstruct the first sparse model with a reduced number of point clouds. Then, by rendering the point cloud attribute information in the first sparse model, the scene representation ability of the rendered second sparse model is equivalent to that of the target model, that is, the second sparse model can also achieve efficient scene representation with a limited number of point clouds. Furthermore, the compression method projects the second sparse model at the key viewpoints to convert the second sparse model into structured two-dimensional image data, which can reduce the data volume of the model data. Further, the compression method further reduces the model data volume by performing image compression on the projection images of the second sparse model. The compression method provided in this specification can facilitate the deployment of the model on end-side devices by compressing the model data, and also reduces the resources required for model storage and transmission.
[0022] Other functions of the method, device for compressing and decompressing model data provided in this specification will be partially listed in the following description. The inventive aspects of the method, device for compressing and decompressing model data provided in this specification can be fully explained by practice or using the methods, devices and combinations described in the detailed examples below. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] To more clearly illustrate the technical solutions in the embodiments of this specification, the following will briefly introduce the accompanying drawings required for the description of the embodiments. Obviously, the accompanying drawings in the following description are only some embodiments of this specification. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0024] Figure 1 FIG. shows a schematic diagram of an application scenario for model data processing provided according to an embodiment of this specification;
[0025] Figure 2Shows a schematic diagram of a computing device provided according to an embodiment of the present specification;
[0026] Figure 3 Shows a flowchart of a method for compressing model data provided according to some embodiments of the present specification;
[0027] Figure 4 Shows a schematic diagram of a target model provided according to some embodiments of the present specification;
[0028] Figure 5 Shows a schematic diagram of a first sparse model provided according to some embodiments of the present specification;
[0029] Figure 6 Shows a schematic diagram of a second sparse model provided according to some embodiments of the present specification;
[0030] Figure 7 Shows a schematic diagram of a first set of projection images provided according to some embodiments of the present specification;
[0031] Figure 8 Shows a schematic diagram of a second set of projection images provided according to some embodiments of the present specification; and
[0032] Figure 9 Shows a flowchart of a method for decompressing model data provided according to some embodiments of the present specification. Detailed implementation
[0033] The following description provides specific application scenarios and requirements of the present specification, aiming to enable those skilled in the art to manufacture and use the content in the present specification. For those skilled in the art, various local modifications to the disclosed embodiments are obvious, and without departing from the spirit and scope of the present specification, the general principles defined here can be applied to other embodiments and applications. Therefore, the present specification is not limited to the shown embodiments, but has the broadest scope consistent with the claims.
[0034] The terms used here are only for the purpose of describing specific example embodiments and are not restrictive. For example, unless the context clearly indicates otherwise, the singular forms "a", "an" and "the" used here may also include the plural forms. When used in this specification, the terms "comprising", "including" and / or "containing" mean that the associated integers, steps, operations, elements and / or components exist, but do not exclude the existence of one or more other features, integers, steps, operations, elements, components and / or groups, or the addition of other features, integers, steps, operations, elements, components and / or groups in the system / method.
[0035] In view of the following description, these features of this specification and other features, as well as the operations and functions of the relevant elements of the structure, and the combination and manufacturing economy of the components can be significantly improved. Referring to the accompanying drawings, all of which form a part of this specification. However, it should be clearly understood that the drawings are only for illustrative and descriptive purposes and are not intended to limit the scope of this specification. It should also be understood that the drawings are not drawn to scale.
[0036] The flowcharts used in this specification illustrate the operations implemented by the system according to some embodiments in this specification. It should be clearly understood that the operations of the flowchart may not be implemented in sequence. On the contrary, the operations may be implemented in reverse order or simultaneously. In addition, one or more other operations may be added to the flowchart. One or more operations may be removed from the flowchart.
[0037] For the convenience of description, the terms involved in this specification are first explained as follows.
[0038] Attribute: In the fields of 3D graphics and computer vision, an attribute generally refers to various features and parameters related to objects, scenes, or images in 3D space, also known as 3D attributes. These attributes can be physical, geometric, or visual, and they determine the manifestation form of an object in the 3D world. For example, attributes can include geometric attributes (position, shape, size, orientation, normal, etc.), visual attributes (color, texture, lighting, shadow, etc.), physical attributes (material, density, etc.), image attributes (pixel value, edge, etc.), and spatial relationship attributes (distance, occlusion relationship, etc.).
[0039] Channel: In the fields of 3D graphics and computer vision, a channel can be considered a manifestation form and storage method of an attribute, and the attribute is the information content carried by the channel. For example, in an image, the color attribute can be stored through three channels of RGB (Red, Green, Blue), and each channel represents a color component. Another example is that the normal information of a 3D object can be expressed through a normal channel, while the color information can be expressed through the RGB channels. Channels and attributes together form the basis for describing and processing the 3D world.
[0040] 3D reconstruction technology is a process of converting 2D images or data into 3D models. This technology has a wide range of applications in many fields such as computer vision, graphics, robotics, medical imaging, and archaeology. The goal of 3D reconstruction technology is to recover the 3D geometric shape and appearance information of an object or scene from the given data. Here, the 3DGS reconstruction technology is taken as an example for introduction. And for the convenience of narration, objects and scenes are collectively referred to as scenes.
[0041] The 3DGS reconstruction technology mainly obtains the 3D Gaussian point cloud representation of the scene by collecting multi-view images of the scene and combining with the training of the cutting-edge differentiable neural radiance field. Among them, each Gaussian point cloud data contains the 3D attributes observed from different perspectives in the scene. The 3DGS reconstruction technology then trains and optimizes the 3D Gaussian point cloud through a network based on the differentiable rendering process supervised by multiple view images, so as to achieve high-quality reconstruction and representation of the 3D scene.
[0042] This specification provides a method for processing model data. This method can achieve the compression and reconstruction of any three-dimensional model. Specifically, on the one hand, this specification provides a method for compressing model data. This method can achieve the compression of any three-dimensional model. When this method is applied to the 3DGS reconstruction technology, it can achieve the compression of any 3DGS model. On the other hand, this specification provides a method for decompressing model data. This method can decompress the compressed data obtained by the above compression method to achieve the reconstruction of the three-dimensional model. When this method is applied to the 3DGS reconstruction technology, it can achieve the reconstruction of the 3DGS model.
[0043] Those skilled in the art should understand that the model data compression method and / or the model data decompression method described in this specification are also within the protection scope of this specification when applied to other three-dimensional reconstruction technologies, such as the voxel rendering algorithm based on Neural Radiance Field (NeRF), etc.
[0044] Figure 1 Shows a schematic diagram of an application scenario for processing model data provided according to some embodiments of this specification. As Figure 1 shown, the scene 100 may include a compression device 110 for model data (hereinafter referred to as the compression device 110), a decompression device 120 for model data (hereinafter referred to as the decompression device 120), and a transmission medium 130.
[0045] The compression device 110 can receive the target model to be compressed and use the model data compression method proposed in this specification to compress the target model to obtain the corresponding compressed data. The compression device 110 may store the data or instructions for executing the compression method described in this specification, and execute the data and / or instructions to obtain the compressed data corresponding to the target model.
[0046] In some embodiments, the compression device 110 may also be a generating device for the target model. After the compression device 110 generates the target model, it then compresses the target model to facilitate subsequent transmission and storage of the model data, etc.
[0047] The decompression device 120 can receive the compressed data corresponding to the target model and decompress the compressed data using the decompression method of the model data proposed in this specification to obtain the decompressed data corresponding to the target model. The decompression device 120 can store the data or instructions for performing the decompression method described in this specification and execute the data and / or instructions to obtain the decompressed data corresponding to the target model.
[0048] The compression device 110 and the decompression device 120 can include a wide range of devices. For example, the compression device 110 and the decompression device 120 can include desktop computers, mobile computing devices, notebooks (e.g., laptops) computers, tablet computers, set-top boxes, handheld devices such as smart phones, televisions, cameras, display devices, digital media players, video game consoles, in-vehicle computers, or the like.
[0049] As Figure 1 shown, the compression device 110 and the decompression device 120 can be connected through a transmission medium 130. The transmission medium 130 can facilitate the transmission of information and / or data. The transmission medium 130 can be any data carrier that can transmit the compressed data from the compression device 110 to the decompression device 120. For example, the transmission medium 130 can be a storage medium (e.g., optical disc), a wired or wireless communication medium. The communication medium can be a network. In some embodiments, the transmission medium 130 can be any type of wired or wireless network or a combination thereof. For example, the transmission medium 130 can include a cable network, a wired network, an optical fiber network, a telecommunication network, an intranet, the Internet, a local area network (LAN), a wireless local area network (WLAN), a metropolitan area network (MAN), a wide area network (WAN), a public switched telephone network (PSTN), a Bluetooth network, a ZigBee network, a near field communication (NFC) network, or a similar network.
[0050] One or more components in the compression device 110 and the decompression device 120 may be connected to the transmission medium 130 to transmit data and / or information. The transmission medium 130 may include a router, a switch, a base station, or other devices that facilitate communication from the compression device 110 to the decompression device 120. In other embodiments, the transmission medium 130 may be a storage medium, for example, a mass storage, a removable storage, a volatile read-write memory, a read-only memory (ROM), or the like, or any combination thereof. Exemplary mass storage may include non-transitory storage media such as magnetic disks, optical disks, solid state drives, etc. Removable storage may include flash drives, floppy disks, optical disks, memory cards, zip disks, magnetic tapes, etc. Typical volatile read-write memory may include random access memory (RAM). RAM may include dynamic RAM (DRAM), double data rate synchronous dynamic RAM (DDR SDRAM), static RAM (SRAM), thyristor RAM (T-RAM), and zero-capacitor RAM (Z-RAM), etc. ROM may include masked ROM (MROM), programmable ROM (PROM), virtual programmable ROM (PEROM), electrically programmable ROM (EEPROM), compact disc ROM (CD-ROM), and digital versatile disc ROM, etc. In some embodiments, the transmission medium 130 may be a cloud platform. Merely by way of example, the cloud platform may include forms such as private cloud, public cloud, hybrid cloud, community cloud, distributed cloud, inter-cloud, or forms similar to the above, or any combination of the above forms.
[0051] As Figure 1 shown, the compression device 110 receives the target model to be compressed, and executes the instructions of the model data compression method described in this specification to compress the data of the target model, obtaining the compressed data corresponding to the target model; the compressed data is transmitted to the decompression device 120 through the transmission medium 130; the decompression device 120 executes the instructions of the model data decompression method described in this specification to decompress the compressed data, obtaining the decompressed data corresponding to the target model.
[0052] In some embodiments, the compression device 110 and the decompression device 120 may be different devices. As Figure 1 shown, the target model to be compressed is compressed by a dedicated compression device 110, and decompressed by another dedicated decompression device 120 at the receiving end. In some embodiments, the compression device 110 and the decompression device 120 may be the same device. This means that one device can both complete the compression of the target model and the decompression of the compressed data corresponding to the target model. In this case, compressing the model data can facilitate the storage of the model data.
[0053] Figure 2 shows a schematic diagram of a computing device 200 provided according to some embodiments of the present specification. The compression device 110 and the decompression device 120 may have the structure of the computing device 200 as shown in Figure 2 shown.
[0054] As Figure 2 shown, the computing device 200 includes at least one storage medium 230 and at least one processor 220. In some embodiments, the computing device 200 may further include a communication port 250 and an internal communication bus 210. At the same time, the computing device 200 may further include I / O components 260.
[0055] The internal communication bus 210 can connect different system components, including the storage medium 230 and the processor 220. The I / O components 260 support input / output between the computing device 200 and other components.
[0056] The storage medium 230 may include a data storage device. The data storage device may be a non-transitory storage medium or a transitory storage medium. For example, the data storage device may include one or more of a magnetic disk 232, a read-only storage medium (ROM) 234, or a random access storage medium (RAM) 236. The storage medium 230 further includes at least one set of instructions stored in the data storage device. The set of instructions is computer program code, and the computer program code may include programs, routines, objects, components, data structures, procedures, modules, etc. for executing the compression method or decompression method provided in the present specification.
[0057] The communication port 250 is used for data communication between the compression device 110 or the decompression device 120 and the outside world. For example, the data compression device 110 may be connected to a transmission medium 130 through the communication port 250.
[0058] The processor 220 can execute all the steps included in the compression method or the decompression method. The processor 220 may be in the form of one or more processors. The processor 220 can issue execution instructions. The processor 220 may include one or more hardware processors, such as a microcontroller, a microprocessor, a reduced instruction set computer (RISC), an application specific integrated circuit (ASIC), an application specific instruction set processor (ASIP), a central processing unit (CPU), a graphics processing unit (GPU), a physics processing unit (PPU), a microcontroller unit, a digital signal processor (DSP), a field programmable gate array (FPGA), an advanced RISC machine (ARM), a programmable logic device (PLD), any circuit or processor capable of executing one or more functions, etc., or any combination thereof.
[0059] For purposes of illustration only, only one processor 220 is shown in the computing device 200 in the drawings of this specification. However, it should be noted that the computing device 200 in this specification may also include multiple processors. Therefore, the operations and / or method steps disclosed in this specification may be performed by one processor as described in this specification, or jointly performed by multiple processors. For example, if the processor 220 of the computing device 200 executes step A and step B in this specification, it should be understood that step A and step B may also be jointly or separately executed by two different processors 220 (for example, the first processor executes step A, the second processor executes step B, or the first and second processors jointly execute steps A and B).
[0060] Figure 3 The flowchart of a compression method P300 for model data provided according to some embodiments of this specification is shown. As described above, the compression device 110 may execute the compression method P300. Specifically, the processor 220 in the compression device 110 may read the instruction set stored in its local storage medium, and then execute the compression method P300 for model data according to the provisions of the instruction set. As Figure 3 shown, the method P300 may include:
[0061] P310: Obtain a target model to be compressed. The target model may be a model obtained by performing three-dimensional reconstruction on the captured images from M shooting perspectives, where M is an integer greater than 2. The target model may include the position information and attribute information of multiple point clouds.
[0062] Among them, three-dimensional reconstruction refers to the technology of converting two-dimensional image data into a three-dimensional model.
[0063] The target model is a three-dimensional model constructed by using three-dimensional reconstruction technology for multiple captured images of the target object. The target object may be any object or scene that the user expects to obtain or reconstruct, such as a building, a car, a person, a sea view, etc.
[0064] Figure 4 The schematic diagram of a target model provided according to some embodiments of this specification is shown. As Figure 4 shown therein, the target object is an excavator, and the target model is the three-dimensional model of the excavator.
[0065] The target model may be composed of a bunch of points or point clouds. Therefore, the target model may include the position information and attribute information of multiple point clouds. Among them, the position information may include the coordinates of each point in the point cloud in three-dimensional space. The attribute information may include other information of each point in the point cloud in three-dimensional space except for the position information. For example, the aforementioned color, intensity, normal (direction of the surface), texture, etc. The attribute information may provide more details about the surface characteristics of the target model.
[0066] The target model utilizes multiple captured images of the target object during the reconstruction process. The shooting orientation / angle of the camera when each captured image is taken is called the shooting perspective. By taking pictures from multiple different shooting perspectives, relatively rich captured images are obtained. These captured images can provide visual information from different perspectives of the target object, thus helping to capture different sides and details of the target object. The above shooting perspectives can be arbitrarily selected or arranged according to a certain rule or pattern. For example, the shooting perspectives differ by a fixed angle. In some embodiments, in order to obtain more information about the target object, the shooting perspectives can be as many as possible. Further, in some embodiments, in order not to miss information about the target object, the field of view ranges corresponding to different shooting perspectives can overlap with each other. It can be understood that when the shooting perspectives are relatively rich, the reconstructed target model has a stronger representation ability for the target object, making the target model more vivid and lifelike.
[0067] For example, Figure 4 Eight shooting perspectives, View1~View8, are shown. Among them, M = 8. Eight or eight groups of captured pictures P1~P8 (P8 is not shown) will be obtained under the eight shooting perspectives. A three-dimensional model of an excavator, that is, the target model, can be obtained through three-dimensional reconstruction using the eight captured images.
[0068] As an example, the process of performing three-dimensional reconstruction to obtain the target model described above can be implemented by 3DGS reconstruction technology. Specifically, after collecting the captured images of multiple shooting perspectives, based on the above captured images, a training process of ray rendering and differentiable 3D Gaussian point cloud is performed, and a target model with dense point cloud can be obtained.
[0069] Continuing to refer to Figure 3 , method P300 may further include step P320.
[0070] P320: Determine N key perspectives among the M shooting perspectives. By performing projection and back-projection on the position information of each point cloud in the target model under the N key perspectives, a first sparse model is obtained. The first sparse model includes the position information of some point clouds in the target model, and N is an integer less than M.
[0071] Among them, the key perspective can also be called the supervision perspective, which can refer to the most representative or relatively representative or key perspective among all shooting perspectives. For example, the key perspective may provide a more comprehensive coverage of the scene, and the key perspective may contain more important information about the scene, thus being more capable of representing the object or scene and playing an important role in reconstructing the target model.
[0072] Among multiple shooting perspectives, some shooting perspectives may contain duplicate information, thus there may be a phenomenon of data redundancy. In the case of limited resources, processing data from all shooting perspectives will consume more computing resources. By selecting N key perspectives, data redundancy can be reduced and the use of computing resources can be optimized.
[0073] For example, as Figure 4 shown, four key perspectives, View1, View3, View5, and View7, are determined from eight shooting perspectives. Taking the bucket of the excavator model as the front of the excavator, these four key perspectives can be the forward perspective View1 facing the front of the excavator model, the backward perspective View5 facing the back of the excavator model, the leftward perspective View7 facing the left side of the excavator model, and the rightward perspective View3 facing the right side of the excavator model. Through these four key perspectives, a four-view of the excavator can be obtained, thereby determining the overall structure of the excavator. It should be noted that the actual orientation of the key perspective can be different from Figure 4 the orientation of the perspective shown.
[0074] In some embodiments, determining N key perspectives from M shooting perspectives can be obtained through the following steps: Based on the amount of information contained in the shooting images under M shooting perspectives respectively, select the perspectives whose contained information meets the target condition from M shooting perspectives as N key perspectives.
[0075] The more information contained in the shooting image, the more it can be considered that the shooting image can better represent the scene, so as to ensure that the scene can be efficiently characterized with a limited number of point clouds subsequently. For example, a shooting image with high information content may contain feature points, textures, or structures that are not in other shooting images (i.e., shooting images with lower information content), etc.
[0076] In some embodiments, the amount of information contained in an image can be measured by the information entropy of the image. The higher the amount of information contained in the shooting image, the higher the information entropy of the shooting image. For example, the resolution of the shooting image is 256×256. The shooting image can be composed of pixels in 256 rows and 256 columns. The information entropy of each pixel can be calculated through the information entropy calculation formula, and then the information entropy of all pixels is added up to obtain the information entropy of the entire image.
[0077] In some embodiments, the target condition may include: the information entropy of the shooting image is greater than a preset information entropy threshold. For example, select the perspectives whose information entropy exceeds the information entropy threshold from M shooting perspectives as N key perspectives. Among them, the information entropy threshold can be set according to manual experience or determined through machine learning, which is not limited in this specification. Another example is to select the top N perspectives with the highest information entropy from M shooting perspectives as key perspectives. In this case, N can be a value preset based on the actual situation.
[0078] In some embodiments, the angular difference between two adjacent key perspectives among the N key perspectives is greater than a preset threshold. The angular difference may refer to the spatial angle offset from one key perspective to another. That is to say, the N key perspectives can be relatively evenly distributed. The relatively even distribution of the N key perspectives can enable the scene to be covered as comprehensively as possible within a limited perspective, and the perspectives can complement each other to provide relatively complete scene information. For example, as Figure 4 shown, the four key perspectives View1, View3, View5, and View7 can be relatively evenly distributed in the three-dimensional space.
[0079] After determining the N key perspectives, the processor 220 can obtain a first sparse model by projecting and back-projecting the position information of each point cloud in the target model under the N key perspectives.
[0080] Among them, the first sparse model can represent the basic shape and structure of the scene, but does not have other details of the scene. For example, the first sparse model only includes the shape of the excavator model, but does not include other attribute details such as the color, texture, and material of the excavator.
[0081] Since the first sparse model is generated by projecting and back-projecting the position information of the point cloud data in the target model under the N key perspectives. In the above process of projection and back-projection, the processor 220 can screen and retain the point cloud data related to the N key perspectives, and discard the point cloud data unrelated to the N key perspectives. Thus, compared with the target model, the number of point clouds (data) included in the first sparse model is less. That is to say, the first sparse model only includes part of the point cloud data in the target model.
[0082] Figure 5 shows a schematic diagram of a first sparse model provided according to some embodiments of the present specification. The following combines Figure 5 to describe the generation process of the first sparse model.
[0083] (1) Project the position information of each point cloud in the target model under the N key perspectives to obtain N projection images.
[0084] Through projection, the information in the high-dimensional space can be mapped into the low-dimensional space. For example, through projection, the information in the three-dimensional space can be converted into the information in the two-dimensional space. Through projection, the two-dimensional projection image can reflect the information of the target model in the three-dimensional space.
[0085] In some embodiments, for each of the N key perspectives, based on the camera observation matrix corresponding to the key perspective, the position information of each point cloud in the target model is projected onto the camera plane corresponding to the key perspective to obtain a projection image corresponding to the key perspective.
[0086] The camera observation matrix can represent the position and orientation of the camera during shooting. The camera observation matrix can transform the point cloud in three-dimensional space into the camera coordinate system. Through the camera observation matrix, the point cloud in the camera coordinate system can finally be projected onto the camera plane to form a two-dimensional image. That is to say, the processor 220 can, through the camera observation matrix, transform each point cloud in the target model into the camera coordinate system and then project the points in the camera coordinate system onto the camera plane, thereby obtaining the corresponding projection image. In some embodiments, by performing matrix operations on the camera observation matrix and the three-dimensional coordinates of the point cloud, the projection coordinates of the point cloud on the camera plane can be obtained.
[0087] In some embodiments, the camera observation matrix can be calculated in real time by the processor 220 after the key perspective is determined. In some embodiments, the camera observation matrices for all shooting perspectives can be pre-stored in the memory, and after the processor 220 determines the key perspective, the processor 220 can directly obtain the corresponding camera observation matrix.
[0088] The camera plane can refer to the plane on which the image is formed. For example, the camera plane can refer to the plane where the camera sensor is located.
[0089] The projection image can reflect the position information of each point cloud in the target model. In some embodiments, the projection image can be a depth map to show the distance of each point cloud to the camera plane. In some embodiments, the projection image can also be a grayscale image.
[0090] For example, as Figure 5 shown, after projecting the position information of each point cloud in the target model onto the corresponding camera plane at the key perspective View1, a projection image can be obtained . Similarly, after projecting the position information of each point cloud in the target model onto the corresponding camera plane at the key perspective View3, a projection image can be obtained . After projecting the position information of each point cloud in the target model onto the corresponding camera plane at the key perspective View5, a projection image can be obtained . After projecting the position information of each point cloud in the target model onto the corresponding camera plane at the key perspective View7, a projection image can be obtained .
[0091] In some embodiments, the position information of each point cloud in the target model can also be projected onto other planes. As long as the plane information of the plane (such as normal information, coordinates of points in the plane, etc.) or the observation matrix corresponding to the viewing angle forming the plane is known, the processor 220 can also perform subsequent compression operations based on the projection image projected onto the plane.
[0092] (2) Back-project the N projection images at N key viewing angles to obtain a first sparse model.
[0093] Back-projection can be considered the inverse process of projection. Through back-projection, information in a low-dimensional space can be restored to a high-dimensional space. Through back-projection, information in a two-dimensional space can be transformed into information in a three-dimensional space. Through back-projection, the two-dimensional projection image is reconstructed into a first sparse model in a three-dimensional space.
[0094] In some embodiments, the first sparse model can be obtained in the following manner: for each of the N key viewing angles, based on the camera observation matrix corresponding to the key viewing angle, back-project each pixel point in the projection image corresponding to the key viewing angle into three-dimensional space to obtain a point cloud subset corresponding to the key viewing angle; and generate a first sparse model based on the point cloud subsets corresponding to the N key viewing angles respectively.
[0095] As can be seen from the above, the conversion of three-dimensional data of a point cloud to a projection image can be achieved through a camera observation matrix. Therefore, through the camera observation matrix, an inverse conversion operation can be performed to back-project each pixel point in the projection image into three-dimensional space, thereby obtaining a point cloud subset. In some embodiments, by performing a matrix inverse operation on the camera observation matrix and the two-dimensional coordinates of the projection image, the three-dimensional data (coordinates) of the point cloud in three-dimensional space can be obtained, and thus a point cloud subset can be obtained.
[0096] The point cloud subsets can directly form the first sparse model. The point cloud subsets can also generate a first sparse model through a process of training and optimization. The process of generating the first sparse model can refer to the process of generating the target model and will not be elaborated here. As an example, the first sparse model can be generated through 3DGS reconstruction technology.
[0097] Continue to refer to Figure 3 , step P300 may further include:
[0098] P330: Render the attribute information of each point cloud in the first sparse model to obtain a second sparse model, so that the scene representation ability of the second sparse model tends to the scene representation ability of the target model.
[0099] As described above, since the first sparse model is obtained by projecting and back-projecting the position information of each point cloud in the target model at N key perspectives, only the position information of some point clouds in the target model is included in the first sparse model, and the attribute information of the point clouds is not included. That is to say, the scene representation ability of the first sparse model obtained through P320 is weaker than that of the target model. Therefore, in P330, the processor 220 can render the attribute information of each point cloud in the first sparse model to obtain a second sparse model, so that the scene representation ability of the second sparse model approaches that of the target model.
[0100] The attribute information may include, but is not limited to, the color information, transparency information, curvature information, density information, material information, etc. described above. As mentioned above, the attribute information can provide more details about the surface characteristics of the model. Rendering the first sparse model with the attribute information can improve the representation ability of the point clouds in the model, and thus obtain a more realistic and detailed second sparse model.
[0101] The scene representation ability may refer to the ability to represent the real scene in three-dimensional space, or it may refer to the degree of realism of restoring the scene. For example, when the user observes the target model and the second sparse model, visually, the user believes that the scene presented by the second sparse model is the same or approximately the same as the scene presented by the target model. In this case, it can be considered that the scene representation ability of the second sparse model approaches that of the target model. Another example is that the difference between the amount of information contained in the target model and the amount of information contained in the second sparse model is less than a preset value. In this case, it can be considered that the scene representation ability of the second sparse model approaches that of the target model.
[0102] The number of point clouds in the second sparse model is less than the number of point clouds in the target model. When the scene representation ability of the second sparse model approaches that of the target model, it means that the second sparse model achieves the same or similar scene representation ability with fewer point clouds. In this way, the second sparse model reduces the model data volume compared with the target model.
[0103] Figure 6 A schematic diagram of a second sparse model according to some embodiments of the present specification is shown. Figure 6 The second sparse model obtained after rendering and training by the first sparse model is shown. The second sparse model has completely or almost presented all the characteristics of the excavator. The excavator presented by the second sparse model is completely or almost the same as the excavator presented by the target model visually.
[0104] In some embodiments, the second sparse model can be obtained through the following steps: using the captured images under N key perspectives as supervision, rendering and training the attribute information of each point cloud in the first sparse model to obtain the second sparse model. During the rendering and training process, minimizing the difference between the projected images of the second sparse model under N key perspectives and the captured images under N key perspectives is used as the training objective.
[0105] Since the captured images under key perspectives include rich and crucial scene or object information. By using these captured images as supervision to perform rendering and training on the first sparse model, the obtained second sparse model can be made closer to the real scene. Rendering and training the attribute information of each point cloud in the first sparse model can be understood as "assigning" the attributes of the pixels in the captured images to the corresponding point clouds. For example, the rendering and training include processes such as "coloring", "light estimation", and "material rendering" for each point cloud in the first sparse model to obtain a second sparse model close to the real scene.
[0106] As an example, the process of performing rendering and training on the first sparse model to obtain the second sparse model can be achieved through 3DGS reconstruction technology. The 3DGS reconstruction technology uses the point clouds in the first sparse model, the captured images (i.e., the ground-truth images under key perspectives), and the rendered images (i.e., the projected images of the rendered sparse model under key perspectives), and trains the model by minimizing the difference between the rendered images and the ground-truth images. By calculating the loss function and then backpropagating to update the point cloud data in the rendered sparse model, the scene representation ability of the rendered sparse model gradually approaches that of the target model.
[0107] To enhance the representation ability of the point cloud data, in some embodiments, the representation dimension of the attribute information of the point clouds in the second sparse model is higher than that of the attribute information of the point clouds in the target model. The representation dimension can refer to the number of vector dimensions or parameters required to describe or represent the attribute information. For example, a certain point cloud in the target model uses a ten-dimensional vector to represent color, and the corresponding point cloud in the second sparse model uses a twenty-dimensional vector to represent color. Therefore, the color information of this point cloud in the second sparse model trained with the captured images as supervision information is richer, thereby enhancing the representation ability of the point cloud data, and further enhancing the representation ability of the limited point cloud data for the scene. Or rather, since the representation dimension left for the second sparse model is relatively high, when using the captured images to render the second sparse model, the processor 220 can extract more and richer information from the captured images, so as to obtain more information about the scene, making the rendered second sparse model closer to the real scene.
[0108] In some embodiments, the number of point clouds included in the second sparse model is different from the number of point clouds included in the first sparse model. For example, the number of point clouds included in the second sparse model is greater than the number of point clouds included in the first sparse model. During the rendering training, the scene representation ability of the model is enhanced by increasing a small number of point clouds. For example, in order to make the scene representation ability of the second sparse model equivalent to that of the target model, a small amount of point cloud data can be added to the second sparse model in very necessary cases during the training process. It can be understood that the number of point clouds added during the rendering process is much smaller than the number of point clouds discarded in P320, that is, the number of point clouds added during the rendering process can be ignored.
[0109] In some embodiments, the number of point clouds included in the second sparse model is the same as the number of point clouds included in the first sparse model. That is to say, during the rendering process, the processor 220 enables the second sparse model to efficiently represent the scene with a smaller number of point clouds by only increasing the representation dimension of the attribute information without adding point cloud data.
[0110] Continue to refer to Figure 3 , method P300 further includes:
[0111] P340: Project the position information and attribute information of each point cloud in the second sparse model at N key viewpoints to obtain a first set of projection images.
[0112] Specifically, the processor 220 can project the position information and attribute information of each point cloud in the second sparse model through the camera observation matrix corresponding to the key viewpoint, project it onto the camera plane, and then obtain a first set of projection images. The projection method of the attribute information is similar to that of the position information. The specific projection method and process can refer to the projection of the target model described above.
[0113] The attribute information of the point clouds in the second sparse model can involve multiple attribute dimensions. The attribute dimension can be the aforementioned attribute or a more specific division of the aforementioned attribute. For example, if the color attribute of the point clouds in the second sparse model is represented by RGB (Red, Green, Blue), then the attribute dimensions involved in the attribute information of the point clouds in the second sparse model include at least three dimensions of R, G, and B.
[0114] When performing projection processing on the second sparse model, the processor may project the information of each attribute dimension in the second sparse model separately under the N key perspectives, so as to obtain a first set of projection images. For example, taking the attribute information of the point cloud in the second sparse model involving five attribute dimensions as an example, assuming the five attribute dimensions are attribute dimension A, attribute dimension B, attribute dimension C, attribute dimension D, and attribute dimension E respectively. The processor 220 may project the attribute information of each point cloud in the second sparse model on attribute dimension A separately under the N key perspectives to obtain N projection images. These N projection images correspond to attribute dimension A. Similarly, the processor 220 may project the attribute information of each point cloud in the second sparse model on attribute dimension B separately under the N key perspectives to obtain N projection images. These N projection images correspond to attribute dimension B. And so on, the projection processes of the processor 220 for attribute dimension C, attribute dimension D, and attribute dimension E are similar.
[0115] Figure 7 shows a schematic diagram of a first set of projection images provided according to some embodiments of the present specification. As Figure 7 shown, the first set of projection images includes projection images under four key perspectives. Specifically, the first set of projection images includes the projected under View1, the projected under View3, the projected under View5, and the projected under View7. The first projection image under each key perspective also includes projection images of different attribute dimensions. Different attribute dimensions are represented by different colors. For example, the projected under View1 includes five projection images corresponding to different attribute dimensions (attribute dimensions A, B, C, D, E). Similarly, the projected under View3, the projected under View5, and the projected under View7 also respectively include the projection images corresponding to the above five attribute dimensions.
[0116] By converting the three-dimensional model structure into a two-dimensional image through the above steps, the discrete point cloud data can be structurally characterized, the data volume of the model data can be reduced, and the subsequent operations for compressing the two-dimensional image can be made simpler, and more image compression algorithms can be used.
[0117] In some embodiments, the first set of projection images may directly serve as the compressed data corresponding to the target model.
[0118] In some embodiments, the projection images in the first projection image set may be further compressed, thereby further improving storage efficiency. Figure 3 , method P300 may further include:
[0119] P350: Perform image compression on each projection image in the first projection image set to obtain a second projection image set, and use the second projection image set as compressed data corresponding to the target model.
[0120] By further compressing the first projection image to obtain the second projection image, the model data can be further compressed, thereby further improving the storage efficiency while ensuring that the scene can be efficiently represented by the model data.
[0121] As described above, each projection image in the first projection image set may correspond to an attribute dimension. In some embodiments, the image compression process of each projection image may include: determining compression parameters applicable to compressing the projection image based on the attribute dimension corresponding to the projection image, and performing image compression processing on the projection image using the compression parameters.
[0122] The compression parameters determined for different attribute dimensions are different. For example, the compression parameter corresponding to attribute dimension A is a, the compression parameter corresponding to attribute dimension B is b, the compression parameter corresponding to attribute dimension C is c, the compression parameter corresponding to attribute dimension D is d, and the compression parameter corresponding to attribute dimension E is e.
[0123] In some embodiments, the same attribute dimension under different perspectives may correspond to the same compression parameter. Figure 7 The compression parameters corresponding to View1, View3, View5 and View7 in the attribute dimension A are the same. Similarly, the compression parameters corresponding to View1, View3, View5 and View7 in the attribute dimensions B~D are also the same.
[0124] In some embodiments, the compression parameter may be a quantization matrix. The quantization matrix may reduce the amount of image data by quantizing discrete cosine transform coefficients.
[0125] In the attribute information of the three-dimensional model, the physical meanings of the information representation of different attribute dimensions are different. Therefore, the physical features that need to be compressed in different attribute dimensions are also different during data compression. Therefore, in the above scheme, for the attribute dimension corresponding to the projection image, the compression parameters suitable for the attribute dimension are adaptively adopted, so that the projection images of different attribute dimensions can be compressed to the maximum extent while ensuring that no useful information is lost.
[0126] In some embodiments, the determining of the compression parameters and the image compression processing of the projection image through the compression parameters can be achieved through the following steps: inputting the projection image and the corresponding attribute dimension of the projection image into a pre-trained adaptive image compression network to perform image compression processing on the projection image through the adaptive image compression network.
[0127] Among them, the adaptive image compression network is trained to perform image compression on projection images with different attribute dimensions using different compression parameters. The adaptive image compression network can determine the corresponding compression parameters based on the input attribute dimension, and then perform compression processing on the projection image based on the compression parameters. For example, the adaptive image compression network can be designed as a differentiable compression network with adaptive channels. This compression network can combine the algorithm network architecture of Diff-JPEG, perform quantization matrix adaptive optimization according to the data characteristics of different feature channels (attribute dimensions), and use different quantization matrices to compress different feature channels, so as to achieve the maximum compression and sparse representation of the image, and then effectively improve the compression efficiency.
[0128] In some embodiments, the adaptive image compression network can be pre-trained and deployed on the compression device 110. The training data of the adaptive image compression network can include the sample projection images of the sample model under different attribute dimensions and the corresponding sample compression parameters. The training objective of the adaptive image compression network can include constraining the difference between the compressed image output by it and the pre-labeled compressed image within a second preset difference threshold. Among them, the second preset difference threshold can be determined based on experience or based on the network training process.
[0129] Figure 8 A schematic diagram of the second set of projection images provided according to some embodiments of the present specification is shown. Figure 8 The process of obtaining the second set of projection images corresponding to each attribute dimension based on the adaptive image compression network and the second set of projection images are shown. As Figure 8 shown, the second set of projection images includes the projection images obtained after compressing the projection image , the projection images obtained after compressing the projection image , the projection images obtained after compressing the projection image , the projection images obtained after compressing the projection image .
[0130] In the above solution, an adaptive image compression network is obtained through pre-training, enabling the adaptive image compression network to have the ability to compress projection images using different compression parameters based on different attribute dimensions. In this way, when compressing each projection image in the first set of projection images, it is only necessary to input the projection image into the pre-trained adaptive image compression network, and then it is possible to adaptively select appropriate compression parameters according to the attribute dimensions of the projection image for compression processing, which can improve the processing efficiency of data compression.
[0131] Thus, the processor 220 has completed the compression of the model data.
[0132] During the compression process of the model data, the processor 220 first selects some key viewpoints from multiple shooting viewpoints, and uses the projection information under the key viewpoints to reconstruct a first sparse model with a reduced number of point clouds. Then, by rendering the point cloud attribute information in the first sparse model, the scene representation ability of the rendered second sparse model is made equivalent to that of the target model, that is, the second sparse model can also achieve efficient scene representation using a limited number of point clouds. Furthermore, the compression method projects the second sparse model under the key viewpoints to convert the second sparse model into structured two-dimensional image data, which can reduce the data volume of the model data. Further, the compression method further reduces the model data volume by compressing the projection images of the second sparse model. The compression method provided in this specification can facilitate the deployment of the scene model on the edge device by compressing the model data, and also reduces the resources required for model storage and transmission.
[0133] After the data of the target model is compressed, when the target model needs to be used, the decompression device 120 can perform decompression processing on the compressed data of the target model to obtain the decompressed data corresponding to the target model. Next, the decompression method of the model data provided in the embodiments of this specification will be described in detail.
[0134] Figure 9 The flowchart of a decompression method P900 of model data provided according to some embodiments of this specification is shown. As described above, the decompression device 120 can execute the decompression method P900. Specifically, the processor 220 in the decompression device 120 can read the instruction set stored in its local storage medium, and then execute the decompression method P900 of this specification according to the provisions of the instruction set. As Figure 9 shown, the method P900 may include:
[0135] P910: Obtain the compressed data corresponding to the target model, where the target model is a model obtained by three-dimensional reconstruction of the captured images from M captured perspectives, the target model includes the position information and attribute information of multiple point clouds, M is an integer greater than 2, and the compressed data is obtained by compressing the target model through a compression method.
[0136] P920: Obtain a second set of projection images from the compressed data, where the second set of projection images includes the compressed projection images corresponding to N key perspectives, and the N key perspectives are partial perspectives among the M captured perspectives, and N is an integer greater than 1.
[0137] For example, the second set of projection images can be the set of images obtained in the above step P350.
[0138] P930: Perform image decompression processing on each projection image in the second set of projection images to obtain a first set of projection images.
[0139] Among them, the first set of projection images are the projection images corresponding to the N key perspectives respectively. For example, the first set of projection images can be the set of images obtained in the above step P340.
[0140] As mentioned above, each projection image in the second set of projection images corresponds to an attribute dimension. Therefore, in some embodiments, the decompression processing corresponding to each projection image can be implemented through the following steps: For the attribute dimension corresponding to the projection image, determine the decompression parameters applicable to decompress the projection image, and perform image decompression processing on the projection image through the decompression parameters.
[0141] Among them, similar to the compression parameters, the decompression parameters determined for different attribute dimensions are different. In some embodiments, in order to ensure that the compressed data can be completely restored, the decompression parameters should maintain a one-to-one correspondence with the compression parameters. For example, when the compression parameter is a quantization matrix, the decompression parameter is an inverse quantization matrix, that is, the inverse matrix of this quantization matrix.
[0142] In some embodiments, the decompression of the projection image can be implemented through a network. For example, an adaptive image decompression network is pre-trained, and the adaptive image decompression network is trained to determine different decompression parameters for different attribute dimensions, and then decompress the projection image based on the decompression parameters. For example, input the projection image and the attribute dimension corresponding to the projection image into the adaptive image decompression network, and the adaptive image decompression network can determine the required inverse quantization matrix based on this attribute dimension, and then use this inverse quantization matrix to decompress the projection image.
[0143] P940: Back-project the first set of projection images from N key perspectives to obtain the decompressed data corresponding to the target model.
[0144] Through back-projection, information in a low-dimensional space can be restored to a high-dimensional space. Through back-projection, information in a two-dimensional space can be transformed into information in a three-dimensional space. The decompressed data obtained by back-projection can be a three-dimensional model. A three-dimensional model (decompressed data) can be generated from the images in the first set of projection images. The process of generating the three-dimensional model (decompressed data) can refer to the process of generating the target model, which will not be elaborated here. As an example, a three-dimensional model (decompressed data) can be generated through 3DGS reconstruction technology.
[0145] The above decompression process can enable the decompressed data to be completely restored or approximately restored to the state of the second sparse model without considering other computational errors. Since the scene representation ability of the second sparse model is comparable to that of the target model, the second sparse model can be used as the decompressed data corresponding to the target model.
[0146] In some embodiments, the processor 220 can also perform other processing on the decompressed data. For example, enhancing the point cloud in the decompressed data, etc., which are not limited in this specification.
[0147] So far, the processor 220 has completed the decompression of the model data.
[0148] During the decompression process of the model data, through decompressing the projection images and performing back-projection operations, decompressed data that is restored or substantially restored to the second sparse model can be obtained, thereby enabling a realistic scene to be restored, thus ensuring the user's immersion and experience.
[0149] This specification further provides a non-transitory storage medium storing at least one set of executable instructions for compressing model data. When the executable instructions are executed by a processor, the executable instructions direct the processor to implement the steps of method P300.
[0150] This specification further provides a non-transitory storage medium storing at least one set of executable instructions for decompressing model data. When the executable instructions are executed by a processor, the executable instructions direct the processor to implement the steps of method P900.
[0151] In some possible embodiments, various aspects of this specification can also be implemented in the form of a program product, which includes program code. For example, when the program product runs on the compression device 110, the program code is used to cause the compression device 110 to execute the steps of the compression method P300 described in this specification. Another example is that when the program product runs on the decompression device 120, the program code is used to cause the decompression device 120 to execute the steps of the decompression method P900 described in this specification.
[0152] The program product for implementing the above method can be stored on a portable compact disc read-only memory (CD-ROM) and run on the compression device 110 or the decompression device 120. However, the program product of this specification is not limited to this. In this specification, the readable storage medium can be any tangible medium that contains or stores a program, and this program can be used by or in combination with an instruction execution system. The program product can adopt any combination of one or more readable media. The readable medium can be a readable signal medium or a readable storage medium. The readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the readable storage medium include: an electrical connection with one or more wires, a portable disc, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. The computer-readable storage medium can include a data signal propagated in a baseband or as part of a carrier wave, in which the readable program code is carried. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The readable storage medium can also be any readable medium other than the readable storage medium, and this readable medium can send, propagate, or transmit a program for use by or in combination with an instruction execution system, apparatus, or device. The program code contained on the readable storage medium can be transmitted using any appropriate medium, including but not limited to wireless, wired, optical cable, RF, etc., or any suitable combination of the above. The program code for performing the operations of this specification can be written in any combination of one or more programming languages, including object-oriented programming languages - such as Java, C++, etc., and also including conventional procedural programming languages - such as the "C" language or similar programming languages.
[0153] The program code for implementing the compression method of model data can be executed entirely on the compression device 110, partially on the compression device 110, executed as an independent software package, partially on the compression device 110 and partially on a remote computing device, or entirely on a remote computing device. The program code for implementing the decompression method of model data can be executed entirely on the decompression device 120, partially on the decompression device 120, executed as an independent software package, partially on the decompression device 120 and partially on a remote computing device, or entirely on a remote computing device. In the case involving a remote computing device, the remote computing device can be connected to the compression device 110 or the decompression device 120 through a transmission medium 130, or can be connected to an external computing device.
[0154] The specific embodiments of this specification have been described above. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be performed in a different order than in the embodiments and still achieve the desired result. Additionally, the processes depicted in the drawings do not necessarily require a specific order or a sequential order to achieve the desired result. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0155] In summary, after reading this detailed disclosure, those skilled in the art will understand that the foregoing detailed disclosure may be presented only by way of example and may not be restrictive. Although not explicitly stated here, those skilled in the art can understand that this specification is intended to encompass various reasonable changes, improvements, and modifications to the embodiments. These changes, improvements, and modifications are intended to be proposed by this specification and are within the spirit and scope of the exemplary embodiments of this specification.
[0156] Furthermore, certain terms in this specification have been used to describe the embodiments of this specification. For example, "one embodiment", "an embodiment", and / or "some embodiments" mean that the specific features, structures, or characteristics described in connection with that embodiment can be included in at least one embodiment of this specification. Therefore, it should be emphasized and understood that two or more references to "an embodiment" or "one embodiment" or "alternative embodiments" in various parts of this specification do not necessarily all refer to the same embodiment. Additionally, the specific features, structures, or characteristics can be appropriately combined in one or more embodiments of this specification.
[0157] It should be understood that in the foregoing description of the embodiments of this specification, for the purpose of helping to understand a feature and for the purpose of simplifying this specification, this specification combines various features in a single embodiment, drawing, or its description. However, this does not mean that the combination of these features is necessary. When reading this specification, those skilled in the art may very well mark out some of the devices as separate embodiments for understanding. That is to say, the embodiments in this specification can also be understood as the integration of multiple sub-embodiments. And the content of each sub-embodiment is also valid when it has fewer features than all the features of a single foregoing disclosed embodiment.
[0158] Each patent, patent application, published patent application, and other materials cited herein, such as articles, books, specifications, publications, documents, items, etc., except for those that are inconsistent with or conflict with this document, or those that have a restrictive effect on the broadest scope of the claims, may be incorporated herein by reference and used for all purposes now or hereafter related to this document. In addition, in the event of any inconsistency or conflict between the description, definition, and / or use of relevant terms in any material and the description, definition, and / or use of relevant terms in this document, the terms in this document shall prevail.
[0159] Finally, it should be understood that the embodiments of the application disclosed herein are illustrative of the principles of the embodiments of this specification. Other modified embodiments are also within the scope of this specification. Therefore, the embodiments disclosed in this specification are merely examples and not limitations. Those skilled in the art can adopt alternative configurations according to the embodiments in this specification to implement the application in this specification. Therefore, the embodiments of this specification are not limited to the embodiments precisely described in the application.
Claims
1. A method for compressing model data, comprising: Obtaining a target model to be compressed, wherein the target model is a model obtained by performing three-dimensional reconstruction on images captured at M shooting angles, the target model includes position information and attribute information of multiple point clouds, and M is an integer greater than 2; Determine N key viewing angles from the M shooting viewing angles, and obtain a first sparse model by projecting and back-projecting the position information of each point cloud in the target model under the N key viewing angles, wherein the first sparse model includes the position information of part of the point cloud in the target model, and N is an integer less than M; Rendering the attribute information of each point cloud in the first sparse model to obtain a second sparse model, so that the scene representation capability of the second sparse model approaches the scene representation capability of the target model; Projecting the position information and attribute information of each point cloud in the second sparse model at the N key viewing angles to obtain a first projection image set; as well as Perform image compression on each projection image in the first projection image set to obtain a second projection image set, and use the second projection image set as compressed data corresponding to the target model.
2. The method according to claim 1, wherein: The determining of N key viewing angles from the M shooting viewing angles includes: Based on the information amounts respectively contained in the captured images under the M shooting perspectives, perspectives whose information amounts satisfy the target condition are selected from the M shooting perspectives as the N key perspectives.
3. The method according to claim 2, wherein: A viewing angle difference between two adjacent key viewing angles among the N key viewing angles is greater than a preset threshold.
4. The method according to claim 1, wherein: The first sparse model is obtained by projecting and back-projecting the position information of each point cloud in the target model at N key viewing angles, including: Projecting the position information of each point cloud in the target model at the N key viewing angles to obtain N projection images; and The first sparse model is obtained by back-projecting the N projection images at the N key viewing angles.
5. The method according to claim 4, wherein: The projecting the position information of each point cloud in the target model under the N key viewing angles to obtain N projection images includes: For each of the N key perspectives, based on the camera observation matrix corresponding to the key perspective, the position information of each point cloud in the target model is projected onto the camera plane corresponding to the key perspective to obtain a projection image corresponding to the key perspective.
6. The method according to claim 4, wherein: The back-projecting the N projection images at the N key viewing angles to obtain the first sparse model includes: For each key perspective of the N key perspectives, based on a camera observation matrix corresponding to the key perspective, back-project each pixel point in the projection image corresponding to the key perspective into a three-dimensional space to obtain a point cloud subset corresponding to the key perspective; and The first sparse model is generated based on the point cloud subsets corresponding to the N key perspectives respectively.
7. The method according to claim 1, wherein: The rendering of the attribute information of each point cloud in the first sparse model to obtain a second sparse model includes: Using the captured images at the N key viewing angles as supervision, the attribute information of each point cloud in the first sparse model is rendered and trained to obtain a second sparse model, wherein the following training objective is adopted during the rendering training process: minimizing the difference between the projection image of the second sparse model at the N key viewing angles and the captured images at the N key viewing angles.
8. The method according to claim 7, wherein: The representation dimension of the attribute information of the point cloud in the second sparse model is higher than the representation dimension of the attribute information of the point cloud in the target model.
9. The method according to claim 1, wherein: Each projection image in the first projection image set corresponds to an attribute dimension, and the image compression process of each projection image includes: Based on the attribute dimension corresponding to the projection image, a compression parameter suitable for compressing the projection image is determined, and image compression processing is performed on the projection image using the compression parameter, wherein different compression parameters are determined for different attribute dimensions.
10. The method according to claim 9, wherein: The step of determining a compression parameter applicable to compressing the projection image based on the attribute dimension corresponding to the projection image, and performing image compression processing on the projection image by using the compression parameter, includes: The projection image and the attribute dimension corresponding to the projection image are input into a pre-trained adaptive image compression network so as to perform image compression processing on the projection image through the adaptive image compression network, wherein the adaptive image compression network is trained to use different compression parameters for image compression for projection images with different attribute dimensions.
11. A model data compression device, comprising: at least one storage medium storing at least one instruction set for data compression; as well as at least one processor, in communication with the at least one storage medium, Wherein, when the compression device is running, the at least one processor reads the at least one instruction set and executes the method according to any one of claims 1 to 10 according to the instructions of the at least one instruction set.
12. A method for decompressing model data, comprising: Obtaining compressed data corresponding to a target model, wherein the target model is a model obtained by three-dimensionally reconstructing images captured at M shooting angles, the target model includes location information and attribute information of a plurality of point clouds, M is an integer greater than 2, and the compressed data is obtained by compressing the target model by the method according to any one of claims 1 to 10; Obtaining a second projection image set from the compressed data, the second projection image set including compressed projection images corresponding to N key viewing angles, where the N key viewing angles are partial viewing angles of the M shooting viewing angles, and N is an integer greater than 1; performing image decompression processing on each projection image in the second projection image set to obtain a first projection image set; and The first projection image set is back-projected at the N key viewing angles to obtain decompressed data corresponding to the target model.
13. The method according to claim 12, wherein: Each projection image in the second projection image set corresponds to an attribute dimension, and the decompression processing process corresponding to each projection image includes: Decompression parameters applicable to decompressing the projection image are determined for the attribute dimensions corresponding to the projection image, and image decompression processing is performed on the projection image using the decompression parameters, wherein different decompression parameters are determined for different attribute dimensions.
14. A decompression device for model data, comprising: at least one storage medium storing at least one instruction set for decompressing data; as well as at least one processor, in communication with the at least one storage medium, Wherein, when the decompression device is running, the at least one processor reads the at least one instruction set and executes the method according to any one of claims 12-13 according to the instructions of the at least one instruction set.
Citation Information
Patent Citations
Face modeling method and device, electronic device, storage medium, and product
CN109376698A
Multi-view video compression method and device based on light field, equipment and medium
CN111757125A