Method, apparatus and computer program product for compressing two-dimensional image
Through the transformer network, the image importance score is predicted and the image subset is selected for compression is solved, and efficient image data processing and high-quality image reconstruction are realized.
Patent Information
- Application Number
- CN202410110870.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-01-25
- Publication Date
- 2025-07-25
AI Technical Summary
The prior art often struggles to efficiently process large-scale data sets when compressing two-dimensional images at the expense of image quality, and traditional methods do not work well when dealing with challenging objects or scenes.
Using an image compressor network based on the transformer network, the image importance score of the image is predicted by training the model and selecting a subset of high-importance images for compression, avoiding manual labeling and calibration processes.
Achieve high compression ratio while maintaining image reconstruction quality, reducing manual intervention and computing resources, and is suitable for efficient data processing of challenging objects or scenarios.
Smart Images

Figure CN120378609A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of data compression, and more particularly, to a method, apparatus, and computer program product for compressing two-dimensional images. Background Art
[0002] In many application scenarios of computer vision, a large number of two-dimensional (2D) images are required as input. However, storing and processing such a large amount of data can be costly and inefficient. To save storage resources and improve processing speed, the prior art usually compresses the input images at the expense of the quality of the output images.
[0003] Recently, some neural networks (e.g., transformer networks) have been applied to three-dimensional (3D) vision tasks, such as point cloud classification, shape generation, or scene completion. The transformer network is a powerful sequence modeling architecture that relies on self-attention mechanisms to capture long-term dependencies. Summary of the Invention
[0004] Embodiments of the present invention provide a method, apparatus, and computer program product for compressing 2D images.
[0005] According to a first aspect of an embodiment of the present invention, there is provided a method for compressing images, the method comprising: determining, by a trained image compressor network, multiple importance scores of multiple images based on pixel values of the multiple images; selecting an image subset from the multiple images according to the multiple importance scores of the multiple images; and compressing the multiple images by retaining the selected image subset and discarding the remaining images.
[0006] According to a second aspect of an embodiment of the present invention, there is provided an electronic device, comprising: at least one processor; and a memory coupled to the at least one processor and having instructions stored thereon that, when executed by the at least one processor, cause the electronic device to perform operations, the operations comprising: determining, by a trained image compressor network, multiple importance scores of multiple images based on pixel values of the multiple images; selecting an image subset from the multiple images according to the multiple importance scores of the multiple images; and compressing the multiple images by retaining the selected image subset and discarding the remaining images.
[0007] According to a third aspect of an embodiment of the present invention, there is provided a computer program product tangibly stored on a non-transitory computer-readable medium and including machine-executable instructions that, when executed, cause a machine to perform operations, the operations including: determining, based on pixel values of a plurality of images, a plurality of importance scores of the plurality of images through a trained image compressor network; selecting, based on the plurality of importance scores of the plurality of images, an image subset from the plurality of images; and compressing the plurality of images by retaining the selected image subset and discarding the remaining images.
[0008] It should be understood that the content described in the summary of the invention is not intended to limit the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0009] In conjunction with the accompanying drawings and with reference to the following detailed description, the above and other features, advantages, and aspects of the embodiments of the present disclosure will become more apparent. In the drawings, the same or similar reference numerals always denote the same or similar elements, and in the drawings:
[0010] Figure 1 A schematic diagram showing an overall design for implementing 3D view synthesis using an image compressor network according to some embodiments of the present disclosure;
[0011] Figure 2 A flowchart showing compressing an image set according to some embodiments of the present disclosure;
[0012] Figure 3 A flowchart showing training an image compressor network according to some embodiments of the present disclosure;
[0013] Figure 4 A flowchart showing determining importance scores of a 2D image set according to some embodiments of the present disclosure; and
[0014] Figure 5 A schematic block diagram of an exemplary device suitable for implementing the embodiments of the present disclosure. DETAILED DESCRIPTION
[0015] Embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although some embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Instead, these embodiments are provided to more thoroughly and completely understand the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are for exemplary purposes only and are not used to limit the protection scope of the present disclosure.
[0016] In the description of the embodiments of the present disclosure, the terms "comprising", "having" and their like terms should be understood as open inclusion, that is, "including but not limited to". The term "based on" should be understood as "at least partially based on". The terms "embodiment", "one embodiment" or "the embodiment" should be understood as "at least one embodiment".
[0017] Image compression is applied to many scenarios, such as for the reconstruction of 2D to 3D images. Such reconstruction usually requires a large number of 2D images as input. However, storing and processing such a large amount of data can be costly and inefficient. Methods such as voxels, point clouds, meshes, or implicit functions are typically used to learn 3D representations from 2D images. But when processing more challenging scenarios with these methods and generalizing to unseen objects, the resulting reconstructed images usually suffer from problems such as low resolution, aliasing artifacts, or lack of fine details.
[0018] To efficiently and highly accurately complete the 2D to 3D image reconstruction task, it is necessary to reduce the storage space and computing resources required for 2D to 3D image reconstruction while maintaining the quality and accuracy of the output image. This problem can be formulated as: Given a set of 2D images, where each image in the set is an image of the same object or scene from different viewpoints, it is necessary to find a subset of images and its corresponding set of camera poses such that: the size of the subset is much smaller than the set of 2D images, and the quality of the 3D image reconstructed from the subset and its corresponding set of camera poses is comparable to the quality of the 3D image reconstructed from the set of 2D images and its corresponding set of camera poses, and the accuracy of the resulting novel view synthesis is also comparable to the accuracy obtained from the set of 2D images and its corresponding set of camera poses.
[0019] This problem is challenging for the following reasons: First, the number of possible subsets of images is exponential with respect to the number of input images, which makes it difficult to computationally find the optimal subset. Second, the quality and accuracy of 3D reconstruction and novel view synthesis depend on the content and geometric structure of the images, so simple heuristic methods (such as selecting the most diverse or representative images) may also not work well. Finally, the 3D reconstruction model is a black-box function, and it is difficult to analyze the contribution of each image to 3D reconstruction and how to measure its importance.
[0020] To this end, the present disclosure proposes a new framework that utilizes an image compressor network based on a transformer network to perform content-aware lossless compression of 2D images. In an embodiment of the present disclosure, by using a trained image compressor network to determine the importance score of an image and retaining or discarding the image according to the importance score of the image, compression of an image set is achieved. Using this solution, a high compression ratio is achieved while maintaining a high image reconstruction quality. In addition, since the process of manual labeling or calibration is avoided, large-scale data sets can be processed with less manual intervention and computing resources.
[0021] Figure 1 FIG. shows a schematic diagram of an environment 100 for synthesizing a 3D view 110 using an image compressor network 104 according to some embodiments of the present disclosure. As Figure 1 shown, the environment 100 includes a computing device 112, and the computing device 112 includes an image compressor network 104, a memory 106, and a 3D reconstruction model 108. The image compressor network 104 may be a trained compressor network obtained after training, and it is composed of an encoder and a decoder. The image compressor network 104 according to the embodiments of the present disclosure can be used to implement various tasks, and the present disclosure does not limit its structure and the specific tasks implemented.
[0022] The computing device 112 may include, but is not limited to, a personal computer, a server computer, a handheld or laptop device, a mobile device (such as a mobile phone, a personal digital assistant (PDA), a media player, etc.), a multiprocessor system, a consumer electronic product, a wearable electronic device, a smart home device, a small computer, a large computer, an edge computing device, a distributed computing environment including any one of the above systems or devices, etc.
[0023] In some embodiments, as Figure 1 shown, a 2D image set 102 is input into the image compressor network 104 in the computing device 112. Generally, the 2D image set 102 includes a plurality of two-dimensional images of the same object or scene captured from different perspectives. The image compressor network 104 predicts the importance score and position of each 2D image according to the pixel values of the 2D image set 102, then selects a subset of images with high importance scores, and discards the remaining images.
[0024] Next, the process by which the image compressor network 104 predicts the importance score and position of the 2D image set 102 will be described in detail.
[0025] The input set of 2D images 102 is first processed by the encoder of the image compressor network 104, which is a vision transformer that divides each image in the 2D image set 102 into multiple pixel blocks and converts these pixel blocks into a sequence of tokens. The sequence of tokens is then fed into a stack of the following layers: the layers apply self-attention and feed-forward operations to encode the global and local features of each image. Finally, the encoded sequence of tokens is output, where each image corresponds to an encoded sequence of tokens.
[0026] The encoded sequence of tokens is input into a decoder, which is also a transformer that uses masked self-attention and cross-attention mechanisms to decode the features from the encoder and generate importance scores and positions for each 2D image. In the way described above, the use of geometric techniques (stereo matching, structure from motion, or multi-view stereo) to estimate the position and importance scores of images is avoided, enabling learning to be performed given only the images, and thus can be generalized to more challenging and invisible objects or scenes.
[0027] In some embodiments, the subset of images with high importance scores selected by the image compressor network 104, along with the corresponding importance scores and positions, are stored in the memory 106 as the data required by the 3D reconstruction model 108 for reading from the memory 106 when 3D reconstruction is needed.
[0028] In some embodiments, the subset of images with high importance scores selected by the image compressor network 104, along with the corresponding importance scores and positions, are directly transmitted into the 3D reconstruction model 108 as the data required by the 3D reconstruction model 108.
[0029] The 3D reconstruction model 108 performs 3D reconstruction on the subset of 2D images using the training data obtained through the image compressor network 104. By continuously training and evaluating the 3D reconstruction model 108, a higher-quality 3D view 110 of the object or scene is finally output.
[0030] Examples of the 3D reconstruction model 108 include but are not limited to neural radiance field (NeRF) reconstruction models. The NeRF reconstruction model is a new 3D scene representation method that can learn a scene only with images and poses, which models the 3D scene as a continuous function of 3D coordinates and viewing directions. Given a set of posed images of a scene, the NeRF model learns a multi-layer perceptron (MLP) network that takes 3D coordinates and viewing directions as inputs and outputs the color and density at that point. By sampling and accumulating the color and density along each camera ray, new views of the scene can be synthesized with high fidelity and consistency.
[0031] By combining Figure 1In the described manner, compared with the traditional methods that sacrifice the output image quality for the sake of compression, the method of the present disclosure achieves a high compression ratio and maintains a high image reconstruction quality. In addition, compared with the traditional methods that require manual marking, calibration, or optimization, the method of the present disclosure can efficiently process large-scale data sets with less manual intervention and computing resources. In addition, since the trained image compressor network 104 only needs images to make predictions, the method of the present disclosure is more applicable to challenging and invisible objects or scenes.
[0032] The following combines Figure 2 Describe a method for compressing a 2D image set 102 according to an embodiment of the present disclosure. Figure 2 A flowchart of a method 200 for compressing a 2D image set 102 according to some embodiments of the present disclosure is shown. The method 200 can be executed at Figure 1 The computing device 112 and any suitable computing device in. In addition, the numbers in the flowchart do not indicate the order in which these steps are executed. Some or all of these steps can be executed in parallel, or the execution order can be interchanged, and the present disclosure does not limit this.
[0033] At block 202, based on the pixel values of multiple images, the trained image compressor network can give an importance score for each 2D image. In some embodiments, in the encoder of the trained image compressor network 104, the pixels of each image are divided into multiple pixel blocks, and these pixel blocks are converted into a sequence of tokens, and then these sequences of tokens are encoded, and then input into the decoder in the trained image compressor network 104, and the decoder decodes the sequence of tokens from the encoder, and then generates an importance score for each 2D image.
[0034] In block 204, based on the importance score of each image, the trained image compressor network selects a subset of images with high importance scores. In some embodiments, an importance score threshold can be set, and images with importance scores above this threshold are selected to form a subset of images with high importance scores.
[0035] In block 206, the trained image compressor network retains the selected subset of images and discards the remaining images. In some embodiments, the retained subset of images will be used as input data for the 3D reconstruction model 108 to synthesize the 3D view 110. In the above manner, the trained image compressor network 104 realizes the compression of multiple 2D images, saves storage resources and computing resources, and at the same time maintains the quality and accuracy of image reconstruction.
[0036] The following combines Figure 3Describe a method 300 for training an image compressor network 104 according to an embodiment of the present disclosure. Figure 3 A flowchart for training an image compressor network 104 according to some embodiments of the present disclosure is shown. Method 300 can be executed at Figure 1 the computing device 112 in
[0037] and any suitable computing device. In addition, the numbers in the flowchart do not represent the order in which these steps are executed. Some or all of these steps can be executed in parallel, or the execution order can be interchanged, and the present disclosure does not limit this. N} and its corresponding pose set P = {P1, P2,..., P N} are input into the 3D reconstruction model 108. In this embodiment, a NeRF reconstruction model is used. Where N represents the number of images in the 2D image set 102. Generally, the pose set P can be read using the COLMAP software. A pose is a six-dimensional vector including the position (x, y, z) and the viewing direction d of the camera that captured the image.
[0038] At block 304, the 3D reconstruction model 108 uses the image set I and the pose set P to learn a multi-layer perceptron (MLP) network F θ . The MLP network F θ is capable of mapping the position and viewing direction of the camera to color c and density σ in order to synthesize a new view of the object or scene using the color and density.
[0039] In some embodiments, given the pose set P, the 3D reconstruction model 108 samples N r points along each camera line of sight r that passes through the center of the camera and the pixels of the image plane. The positions and viewing directions of these N r points are fed into F θ to obtain c i and σ i , i = 1, 2,..., N r . By using alpha compositing to blend the colors along the line of sight r which is expressed as
[0040]
[0041] where, T i is expressed as
[0042]
[0043] represents the transmittance accumulated along the line of sight r up to point i, and alpha i is expressed as
[0044]
[0045] is the α value at point i, Δ i is the distance between point i and point i + 1. Using the difference between the synthesized color and the real color, the loss of view synthesis can be determined. The 3D reconstruction model 108 is optimized by minimizing this loss through the gradient descent method.
[0046] At block 306, an importance score for each 2D image is determined. The following is described in conjunction with Figure 4 a method 400 for determining the importance scores of the 2D image set 102 according to an embodiment of the present disclosure. Figure 4 FIG. shows a flowchart of determining the importance scores of the 2D image set 102 according to some embodiments of the present disclosure. The method 400 can be executed at Figure 1 the computing device 112 in and any suitable computing device.
[0047] In the present disclosure, the method 400 is also referred to as a SHAP-like method. Since SHAP is a framework for explaining machine learning models by assigning feature importance scores based on Shapley values, which is a game theory concept for quantifying the contribution of each player to a cooperative game, in the solution of the present disclosure, each image is regarded as a player and the 3D reconstruction model 108 is regarded as the game. To evaluate the contribution of each image to the 3D reconstruction model 108, the negative L2 loss between the new view synthesized by the 3D reconstruction model 108 and the real new view is used to define the return of the game, and the Shapley value of each image is defined as
[0048]
[0049] where f(S) is the return function returned when training and evaluating the 3D reconstruction model 108 using a subset S of the images and their corresponding poses, f(S ∪ I k ) is the return function returned when training and evaluating the 3D reconstruction model 108 using the union of the subset S of the images and the image I k and their corresponding poses. The Shapley value φ k represents the average marginal contribution to the return when the image I k is added to an arbitrary subset of the images.
[0050] However, since calculating the exact Shapley value requires evaluating the return function for all possible subsets of images, it is computationally time-consuming and expensive. To avoid calculating all possible subsets of images, a sampling-based approximation method is used in the present disclosure. The method includes randomly sampling subsets from the 2D image set 102 (step 402), then calculating the return difference when adding image I k to the randomly sampled subset of images S m and when only having the subset of images S m (step 404), then determining whether the above sampling and calculation process has been performed M times (step 406). If it has not been performed M times, return to step 402. If it has been performed M times, average the calculation results (step 408) to obtain an estimated value of φ k , which estimated value is expressed as
[0051]
[0052] Using this estimated value as the importance score of image I k , each importance score is normalized by summing all importance scores, so that a probability distribution of the 2D image set can be obtained. By using Figure 4 the method 400 shown, it is possible to select the most informative images describing the scene content and geometry according to its execution result.
[0053] Return Figure 3 , at block 308, a new data set consisting of an image tuple, a position tuple, and an importance tuple is created. In this new data set, the data in the image tuple and the position tuple directly come from the input image set and pose set, and only the importance scores in the importance tuple are indirectly obtained through the method 400.
[0054] In some embodiments, the purpose of including the importance scores in this new data set is to train the image compressor network 104 so that it no longer experiences the Figure 4 process shown at inference time for a new image set, but can directly infer the importance scores of the images.
[0055] In some embodiments, the purpose of including the camera position and images in this new data set is to train the image compressor network 104 so that it can estimate the camera position based only on the pixel values of the images at inference time for a new image set, for use as input data for the 3D reconstruction model 108.
[0056] At block 310, the newly created dataset is used to train the image compressor network 104. Different from the NeRF reconstruction model which requires given camera positions and images, the image compressor network 104 only needs the pixel values of the images as input to be able to predict the camera positions corresponding to the images. To train the image compressor network 104, a total loss function is defined to evaluate the accuracy of its predictions. This total loss function is the position loss function and the importance loss function weighted sum, expressed as
[0057]
[0058] where λ1 and λ2 are the weight factors corresponding to and respectively, and their sum is less than or equal to 1. is expressed as
[0059]
[0060] where is the predicted value of the camera position predicted by the image compressor network 104 for image I k , and P k is the ground truth value of the camera position for image I k from the new dataset. is expressed as
[0061]
[0062] where is the predicted importance score predicted by the image compressor network 104 for image I k , and φ k is the importance score for image I k from the new dataset. The image compressor network 104 is optimized by minimizing the total loss function using the gradient descent method.
[0063] In some embodiments, when using the trained image compressor network 104 to synthesize the 3D view 110, only the 2D image set 102 can be input without having to include the positions of the cameras corresponding to the 2D image set 102, because the trained image compressor network 104 can directly predict the camera positions. In addition, the trained image compressor network 104 can also directly predict the importance scores for each 2D image without having to perform Figure 4The process shown. In this way, the inference speed is improved, and at the same time, learning is achieved only from the pixel values of the image without relying on precise camera poses or geometric assumptions, and thus can be generalized to objects or scenes with different shapes, textures, lighting conditions and viewpoints, as well as objects or scenes that cannot be seen, etc.
[0064] Figure 5 FIG. shows a schematic block diagram of an example device 500 that can be used to implement embodiments of the present disclosure. As shown, the device 500 includes a processor 501 that can perform various appropriate actions and processes according to computer program instructions stored in a read-only memory (ROM) 502 or computer program instructions loaded from a storage unit 508 into a random access memory (RAM) 503. In the RAM 503, various programs and data required for the operation of the device 500 can also be stored. The processor 501, the ROM 502, and the RAM 503 are connected to each other via a bus 504. An input / output (I / O) interface 505 is also connected to the bus 504.
[0065] A plurality of components in the device 500 are connected to the I / O interface 505, including: an input unit 506, such as a keyboard, a mouse, etc.; an output unit 507, such as various types of displays, speakers, etc.; a storage unit 508, such as a magnetic disk, an optical disk, etc.; and a communication unit 509, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 509 allows the device 500 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0066] The processor 501 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the processor 501 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The processor 501 executes the various methods and processes described above, such as method 200. For example, in some embodiments, method 200 can be implemented as a computer software program that is tangibly contained in a machine-readable medium, such as the storage unit 508. In some embodiments, part or all of the computer program can be loaded and / or installed onto the device 500 via the ROM 502 and / or the communication unit 509. When the computer program is loaded into the RAM 503 and executed by the processor 501, one or more steps of method 200 described above can be executed. Alternatively, in other embodiments, the processor 501 can be configured to execute method 200 in any other suitable way (e.g., by means of firmware).
[0067] The functions described above in this document can be performed, at least in part, by one or more hardware logic components. For example, without limitation, exemplary types of hardware logic components that can be used include: Field Programmable Gate Arrays (FPGAs), Application Specific Integrated Circuits (ASICs), Application Specific Standard Products (ASSPs), Systems on Chip (SOCs), Complex Programmable Logic Devices (CPLDs), and the like.
[0068] The program code for implementing the methods of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing devices, such that when the program codes are executed by the processor or controller, the functions / operations specified in the flowcharts and / or block diagrams are implemented. The program code can be executed entirely on the machine, partially on the machine, as a stand-alone software package partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0069] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer disk, a hard disk, a Random Access Memory (RAM), a Read-Only Memory (ROM), an Erasable Programmable Read-Only Memory (EPROM or Flash Memory), an optical fiber, a portable Compact Disc Read-Only Memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing. Additionally, although the operations are depicted in a particular order, this should be understood to require that the operations be performed in the particular order shown or in sequential order, or that all illustrated operations be performed to achieve the desired result. In certain environments, multitasking and parallel processing may be advantageous. Similarly, although several specific implementation details are included in the above discussion, these should not be construed as limitations on the scope of the present disclosure. Certain features described in the context of separate embodiments can also be implemented in combination in a single implementation. Conversely, various features described in the context of a single implementation can also be implemented separately or in any suitable sub-combination in multiple implementations.
[0070] Although the subject matter has been described in language specific to structural features and / or methodological acts, it is to be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. On the contrary, the specific features and acts described above are merely example forms of implementing the claims.
Claims
1. A method for compressing images, the method comprising: Determining, by a trained image compressor network, multiple importance scores for the multiple images based on pixel values of the multiple images; Selecting an image subset from the multiple images according to the multiple importance scores of the multiple images; And Compressing the multiple images by retaining the selected image subset and discarding the remaining images.
2. The method according to claim 1, wherein the retained image subset is used for reconstruction of a three-dimensional (3D) scene.
3. The method according to claim 1, wherein determining the multiple importance scores for the multiple images comprises causing an encoder of the trained image compressor network to perform the following steps: Dividing each of the multiple images into multiple pixel blocks; Converting the multiple pixel blocks into a token sequence; Feeding the token sequence into a stack of the following layers of the encoder: the layers apply self-attention and feed-forward operations to encode global features and local features of each image; and Outputting the encoded token sequence, wherein each image corresponds to an encoded token sequence.
4. The method according to claim 3, wherein determining the multiple importance scores for the multiple images further comprises causing a decoder of the trained image compressor network to perform the following steps: Taking the encoded token sequence as an input; Using masked self-attention and cross-attention mechanisms to decode features from the encoder; And Generating a set containing the importance scores for each image.
5. The method according to claim 1, wherein the method further comprises: Performing 3D reconstruction on a set of two-dimensional (2D) images of a 3D scene and a pose set corresponding to the 2D image set using a 3D reconstruction model and synthesizing a new 3D view.
6. The method according to claim 5, wherein the method further comprises: Determining an importance score for each image based on the contribution of each image in the 2D image set to the 3D reconstruction.
7. The method according to claim 6, wherein determining the importance score for each image comprises: Randomly sampling a subset from the 2D image set; Determining the difference in return when adding an image in the 2D image set to the subset and when not adding the image to the subset; Repeating the determination of the difference in return and the sampling a number of times; And Averaging the obtained results to obtain an estimated value of the importance score for the image.
8. The method according to claim 1, wherein the method further comprises: Creating a new data set consisting of image tuples, position tuples, and importance tuples.
9. The method according to claim 8, wherein the method further comprises: Using the created new data set to train the image compressor network.
10. The method according to claim 9, wherein training the image compressor network comprises: Minimizing a total loss function using the gradient descent method; Wherein the total loss function is a weighted sum of a position loss function and an importance loss function; Wherein the position loss function is determined based on the number of images, the ground truth of the camera position, and the estimated value for the camera position obtained by the image compressor network; and Wherein the importance loss function is determined based on the number of images, the estimated value of the importance score according to claim 7, and the estimated value of the importance score for the image obtained by the image compressor network.
11. An electronic device, the electronic device comprising: at least one processor; and a memory, coupled to the at least one processor and having instructions stored thereon, the instructions, when executed by the at least one processor, cause the electronic device to perform actions, the actions including: determining, by a trained image compressor network, multiple importance scores of the multiple images based on pixel values of the multiple images; selecting, from the multiple images, an image subset based on the multiple importance scores of the multiple images; and compressing the multiple images by retaining the selected image subset and discarding the remaining images.
12. The device according to claim 11, wherein the retained image subset is used for reconstruction of a three-dimensional (3D) scene.
13. The device according to claim 11, wherein determining the multiple importance scores of the multiple images includes causing an encoder of the trained image compressor network to perform the following steps: dividing each of the multiple images into multiple pixel blocks; converting the multiple pixel blocks into a token sequence; feeding the token sequence into a stack of the following layers of the encoder: the layers apply self-attention and feed-forward operations to encode global features and local features of each image; and outputting an encoded token sequence, wherein each image corresponds to an encoded token sequence.
14. The device according to claim 13, wherein determining the multiple importance scores of the multiple images further includes causing a decoder of the trained image compressor network to perform the following steps: taking the encoded token sequence as an input; using masked self-attention and cross-attention mechanisms to decode features from the encoder; and generating a set including the importance scores of each image.
15. The device according to claim 11, wherein the action further comprises: Performing 3D reconstruction on a set of 2D images of a 3D scene and a set of poses corresponding to the 2D image set using a 3D reconstruction model and synthesizing a new 3D view.
16. The apparatus according to claim 15, wherein the action further comprises: Determining an importance score of each image based on the contribution of each image in the 2D image set to the 3D reconstruction.
17. The device according to claim 16, wherein determining the importance score of each image includes: randomly sampling a subset from the 2D image set; determining a difference in rewards when adding an image in the 2D image set to the subset and when not adding the image to the subset; repeating the determination of the difference in rewards and the sampling a number of times; and averaging the obtained results to obtain an estimated value of the importance score of the image.
18. The device according to claim 11, wherein the action further comprises: Creating a new data set consisting of image tuples, position tuples, and importance tuples.
19. The apparatus according to claim 18, wherein the action further comprises: Using the created new data set to train the image compressor network.
20. A computer program product, the computer program product being tangibly stored on a non-transitory computer-readable medium and including machine-executable instructions, the machine-executable instructions, when executed, cause the machine to perform actions, the actions including: determining, by a trained image compressor network, multiple importance scores of the multiple images based on pixel values of the multiple images; Select an image subset from the multiple images according to the multiple importance scores of the multiple images; and Compress the multiple images by retaining the selected image subset and discarding the remaining images.