Image processing method, apparatus, device, and product
By generating and displaying aggregated images based on quantified attribute information from multiple videos, and using three-dimensional Gaussian parameters to characterize the scene, the problem of high computational resources for 3D reconstruction models in existing technologies is solved, achieving efficient, realistic, and three-dimensional new perspective image generation, and optimizing the user's immersive live streaming experience.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- BEIJING ZITIAO NETWORK TECH CO LTD
- Filing Date
- 2025-06-20
- Publication Date
- 2026-04-23
AI Technical Summary
Existing technologies consume a lot of computational resources and time when generating 3D reconstruction models, resulting in low efficiency and poor quality in generating images from new perspectives, which cannot meet users' needs for an immersive live streaming experience.
By generating aggregated images based on the quantized attribute information of multiple videos, using three-dimensional Gaussian parameters to characterize the scene, dynamically adjusting the display screen to meet user interaction needs, and generating and displaying the screen position or viewpoint that the user wants to view.
It improves the flexibility and personalization of image generation, resulting in more realistic and three-dimensional images, and significantly enhances the efficiency and quality of rendering new views.
Smart Images

Figure CN2025102526_23042026_PF_FP_ABST
Abstract
Description
Image processing methods, apparatus, equipment and products
[0001] Cross-reference to related applications
[0002] This application claims priority to Chinese Patent Application No. 202411441916.0, filed on October 15, 2024, entitled "Image Processing Method, Apparatus, Device and Product", the entire contents of which are incorporated herein by reference. Technical Field
[0003] This disclosure relates to the field of image processing, and more specifically, to image processing methods, apparatus, devices and products. Background Technology
[0004] With the development of high-definition video and audio technologies, people have increasingly diverse entertainment options. For example, live streaming technology, as an emerging interactive media technology, is gradually gaining popularity due to its diversified and personalized characteristics. Users have high demands for live streaming content, quality, and interactivity, requiring live streaming platforms to continuously innovate and optimize to meet market demands. For instance, users may not be satisfied with a single live streaming method, or they may desire a more immersive live streaming experience. Summary of the Invention
[0005] Embodiments of this disclosure provide an image processing method, apparatus, device, and product.
[0006] In a first aspect of this disclosure, an image processing method is provided. The method includes determining target display adjustment information corresponding to a screen to be displayed based on interactive operations on a display screen. The method further includes determining an aggregated image corresponding to the target display adjustment information, wherein the aggregated image is an image obtained from multiple videos storing multiple quantized attribute information. Each of the multiple videos corresponds to multiple adjacent acquisition viewpoints. The multiple quantized attribute information is stored in the aggregated image through aggregation. The multiple quantized attribute information includes quantized three-dimensional Gaussian parameters, wherein the three-dimensional Gaussian parameters are determined based on multiple images contained in the multiple videos, and the three-dimensional Gaussian parameters are used to characterize the scene contained in the video. The video content of the multiple videos corresponds to the target display adjustment information. The method further includes generating a screen to be displayed and displaying the screen to be displayed based on the aggregated image.
[0007] In a second aspect of this disclosure, an image processing method is provided. The method includes receiving target display adjustment information corresponding to a screen to be displayed from a user terminal. The method further includes determining an aggregated image corresponding to the target display adjustment information based on the target display adjustment information. The aggregated image is an image obtained from multiple videos and storing multiple quantized attribute information. The multiple acquisition viewpoints corresponding to the multiple videos are adjacent. The multiple quantized attribute information includes quantized three-dimensional Gaussian parameters, wherein the three-dimensional Gaussian parameters are determined based on multiple images contained in the multiple videos. The three-dimensional Gaussian parameters are used to characterize the scenes contained in the multiple videos, and the video content of the multiple videos corresponds to the target display adjustment information. The method further includes sending the aggregated image to the user terminal, whereby the user terminal generates and displays the screen to be displayed.
[0008] In a third aspect of this disclosure, an image processing apparatus is provided. The apparatus includes a display adjustment information determining unit configured to determine target display adjustment information corresponding to a display screen based on an interactive operation on the display screen. The apparatus also includes an aggregated image determining unit configured to determine an aggregated image corresponding to the target display adjustment information, wherein the aggregated image is an image obtained from multiple videos storing multiple quantized attribute information, each video corresponding to multiple adjacent acquisition viewpoints, and the multiple quantized attribute information is stored in the aggregated image through aggregation. The multiple quantized attribute information includes quantized three-dimensional Gaussian parameters, wherein the three-dimensional Gaussian parameters are determined based on multiple images contained in the multiple videos, and the three-dimensional Gaussian parameters are used to characterize the scene contained in the video, and the video content of the multiple videos corresponds to the target display adjustment information. The apparatus also includes a display screen generation unit configured to generate and display the display screen based on the aggregated image.
[0009] In a fourth aspect of this disclosure, an image processing apparatus is provided. The apparatus includes a display adjustment information receiving unit configured to receive target display adjustment information corresponding to a screen to be displayed from a user terminal. The apparatus also includes an aggregated image determining unit configured to determine an aggregated image corresponding to the target display adjustment information based on the target display adjustment information. The aggregated image is an image obtained from multiple videos and storing multiple quantized attribute information. Multiple acquisition viewpoints corresponding to the multiple videos are adjacent. The multiple quantized attribute information includes quantized three-dimensional Gaussian parameters, which are determined based on multiple images contained in the multiple videos. The three-dimensional Gaussian parameters characterize the scenes contained in the multiple videos, and the video content of the multiple videos corresponds to the target display adjustment information. The apparatus also includes an aggregated image sending unit configured to send the aggregated image to a user terminal, whereby the user terminal generates and displays the screen to be displayed.
[0010] In a fifth aspect of this disclosure, an electronic device is provided. The electronic device includes one or more processors; and a storage device for storing one or more programs, which, when executed by the one or more processors, cause the one or more processors to implement the methods provided according to a first or second aspect of this disclosure.
[0011] In a sixth aspect of this disclosure, a computer-readable storage medium is provided. The computer-readable storage medium stores computer-executable instructions, which are executed by a processor to implement the method provided according to a first or second aspect of this disclosure.
[0012] In a seventh aspect of this disclosure, a computer program product is provided. The computer program product is tangibly stored on a non-transitory computer-readable medium and includes machine-executable instructions that, when executed, cause a machine to implement the method provided according to a first or second aspect of this disclosure.
[0013] It should be understood that the description in the Summary of the Invention section is not intended to limit the key or essential features of the embodiments of this disclosure, nor is it intended to restrict the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description.
[0014] The summary section is provided to present the chosen concepts in a simplified form, which will be further described in the detailed description below. The summary section is not intended to identify key or principal features of the claimed subject matter, nor is it intended to limit the scope of the claimed subject matter. Attached Figure Description
[0015] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent from the accompanying drawings and the following detailed description. In the drawings, the same or similar reference numerals denote the same or similar elements, wherein:
[0016] Figure 1 illustrates a schematic diagram of an example environment 100 in which various embodiments of the present disclosure may be implemented;
[0017] Figure 2 illustrates a schematic diagram of the workflow of various execution entities during image processing according to some embodiments of the present disclosure;
[0018] Figure 3 shows a flowchart of an image processing method according to some embodiments of the present disclosure;
[0019] Figure 4 illustrates a schematic diagram of acquiring multi-view video according to some embodiments of the present disclosure;
[0020] Figure 5 illustrates a schematic diagram of an exemplary process for generating an image according to some embodiments of the present disclosure;
[0021] Figure 6 illustrates a schematic diagram of determining attribute information of an image according to some embodiments of the present disclosure;
[0022] Figure 7 illustrates a schematic diagram of attribute information contained in an image according to some embodiments of the present disclosure;
[0023] Figure 8 shows a schematic diagram of a new perspective image rendered according to some embodiments of the present disclosure;
[0024] Figure 9 shows a block diagram of an image processing apparatus according to some embodiments of the present disclosure;
[0025] Figure 10 shows a block diagram of an image processing apparatus according to other embodiments of the present disclosure; and
[0026] Figure 11 shows a block diagram of a device that can implement several embodiments of the present disclosure. Detailed Implementation
[0027] It is understood that all user-related data involved in this technical solution should be obtained and used only after authorization from the user. This means that if it is necessary to use a user's personal information in this technical solution, the user's explicit consent and authorization are required before obtaining this data; otherwise, no related data collection and use will be carried out. It should also be understood that when implementing this technical solution, relevant laws and regulations should be strictly followed in the process of data collection, use, and storage, and necessary technical measures should be taken to protect user data security and ensure the secure use of data.
[0028] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.
[0029] In the description of embodiments of this disclosure, the term "comprising" and similar terms should be understood as open-ended inclusion, i.e., "including but not limited to". The term "based on" should be understood as "at least partially based on". The term "one embodiment" or "the embodiment" should be understood as "at least one embodiment". The terms "first", "second", etc., may refer to different or the same objects. Other explicit and implicit definitions may also be included below.
[0030] As mentioned above, users want to freely change the viewing angle of the displayed image and move around in space to experience the video content more comprehensively. In some related technologies, Neural Radiance Fields (NeRF) can be used to generate high-quality 3D reconstruction models. For example, deep learning techniques can be used to extract the geometric shape and texture information of objects from images from multiple perspectives, and then this information can be used to generate a continuous 3D radiation field, thus presenting a highly realistic 3D model at any angle and distance. Based on this 3D model, images from specific perspectives can be generated. However, NeRF requires significant computational resources and time to generate 3D reconstruction models, resulting in low efficiency and poor quality, which in turn leads to low efficiency and poor image quality when generating images from new perspectives.
[0031] Therefore, embodiments of this disclosure provide an image processing method. The method includes determining target display adjustment information corresponding to the screen to be displayed based on interactive operations on the displayed screen. The method further includes determining an aggregated image corresponding to the target display adjustment information, wherein the aggregated image is an image obtained from multiple videos storing various quantized attribute information. Each of the multiple videos corresponds to multiple adjacent acquisition viewpoints. The various quantized attribute information is stored in the aggregated image through aggregation. The various quantized attribute information includes quantized three-dimensional Gaussian parameters, wherein the three-dimensional Gaussian parameters are determined based on multiple images contained in the multiple videos. The three-dimensional Gaussian parameters are used to characterize the scene contained in the video, and the video content of the multiple videos corresponds to the target display adjustment information. The method further includes generating and displaying the screen to be displayed based on the aggregated image.
[0032] In this way, the desired viewing position or perspective can be dynamically generated and displayed based on user interactions, making the presentation and switching of the displayed image more flexible and personalized, further optimizing the user's viewing experience. Moreover, the aggregated image used to generate the image stores various attribute information. Based on these attributes, the scene the user wants to view can be synthesized and presented more accurately, making the generated image more realistic and three-dimensional. Furthermore, the 3D Gaussian parameters in the attribute information can clearly express scene information, and the generation of 3D Gaussian parameters is also faster. Therefore, rendering new views in complex scenes based on 3D Gaussian parameters is more efficient and produces higher image quality.
[0033] Figure 1 illustrates a schematic diagram of an example environment 100 in which various embodiments of the present disclosure may be implemented. As shown in Figure 1, environment 100 includes a user terminal 102, a server 106 communicating with the user terminal 102, and a data acquisition device 104 communicating with the server 106. The user terminal 102 can communicate with the server 106 via a wireless network or a wired network. The data acquisition device 104 can communicate with the server 106 via a wireless network or a wired network. The networks between the user terminal 102 and the server 106, and between the data acquisition device 104 and the server 106, may be the same or different. In one example, in a live streaming scenario, the user terminal 102 may be a terminal device carried by a user watching the live stream, such as, but not limited to, a mobile phone, tablet computer, smart wearable device, laptop computer, desktop computer, etc. In some embodiments, the user watching the live stream can view the live stream content or interact with it (e.g., like, comment, share, etc.) through the user terminal 102. The data acquisition device 104 may be a device with image acquisition capabilities used by the user watching the live stream. The device can be a camera, a video camera, or a terminal device with shooting capabilities, such as, but not limited to, mobile phones, tablet computers, smart wearable devices, laptop computers, and desktop computers. Server 106 can be a monolithic server, a distributed server, a cloud server, or a server cluster deployed in the cloud. In some embodiments, server 106 may include, but is not limited to, blade servers, cloud servers, or a server group composed of multiple servers.
[0034] In some embodiments of this disclosure, the acquisition device 104 can be used to capture images or videos of a scene or objects contained within the scene from multiple perspectives. For example, multiple acquisition devices 104, such as multiple cameras, can be used to capture multiple images of the scene, such as a live broadcast scene, from different positions and perspectives. In one example, n acquisition devices 104 (e.g., optical cameras) can be arranged in a circle around a preset area centered on the scene or object (e.g., a preset polygon or circle centered on the scene or object), and the n acquisition devices 104 can perform surround shooting of the scene. The n acquisition devices 104 can capture multiple videos at the same frame rate from different positions using different shooting perspectives.
[0035] In some embodiments of this disclosure, multiple videos from multiple perspectives can be transmitted to server 106, where server 106 determines a corresponding aggregated image based on the multiple videos with adjacent acquisition perspectives. It is understood that the multiple videos with adjacent acquisition perspectives can partially or jointly cover the scene within the same viewpoint range. The image determined from the multiple videos can store various attribute information, which can be used to characterize the scene within that viewpoint range. These various attribute information may include quantized three-dimensional Gaussian parameters, quantized depth information, quantized color information, etc. The three-dimensional Gaussian parameters can explicitly characterize the corresponding scene. For example, the three-dimensional Gaussian parameters generated from video 1 and video 2 can be used to characterize the scene within that viewpoint range (viewpoint 1 corresponding to video 1 + viewpoint 2 corresponding to video 2).
[0036] In some embodiments, to provide users with multi-view perspectives and a more immersive and diverse live streaming experience, corresponding views can be provided based on the user's selection. For example, a user watching a live stream can input the desired viewing location (such as a specific position in the image) or angle (such as tilt or rotation) through the user terminal 102. In one example, the user can select the desired viewing location by dragging the live stream image. In other examples, the user can also select the desired viewing location through voice or text input. In some embodiments of this disclosure, after determining the desired viewing location or angle, a corresponding aggregated image can be determined from the server 106. For example, the corresponding aggregated image can be determined based on spatial calculations. Based on the quantified attribute information stored in the aggregated image, the image at the corresponding location or angle can be rendered. The user terminal 102 can display this image on a screen, providing the user with a new perspective.
[0037] This approach allows for the dynamic generation and display of the desired viewing position or perspective based on user interactions, making the presentation and switching of the displayed image more flexible and personalized, further optimizing the user's viewing experience. Moreover, the images used to generate the image contain various attribute information, more accurately synthesizing and presenting the scene the user wants to see, resulting in a more realistic and three-dimensional image. Furthermore, 3D Gaussian parameters can clearly express scene information, and their generation is also faster. Based on this, rendering new views in complex scenes using 3D Gaussian parameters is highly efficient and produces high-quality images.
[0038] Figure 2 illustrates a schematic diagram of the workflow of various execution entities in an image processing process according to some embodiments of the present disclosure. As shown in Figure 2, the image generation process 200 can be executed by the user terminal 2002, the server 2004, and the acquisition device 2006. It should be understood that the execution entities involved in the image generation process 200 shown in Figure 2 are merely examples of the present disclosure and should not be construed as limiting the present disclosure. As shown in Figure 2, the image generation process 200 may include steps 201 to 212.
[0039] In step 201, the acquisition device 2006 can acquire multiple videos from multiple perspectives for a scene (e.g., a live broadcast scene) or an object. For example, multiple acquisition devices 2006 can be arranged around the scene, with different shooting angles and positions. It can be understood that multiple acquisition devices 2006 can acquire multiple videos corresponding to multiple perspectives at the same frame rate within the same time period. The videos can include multiple consecutive frames. In one example, the multiple videos can include video A corresponding to perspective A, video B corresponding to perspective B, and video C corresponding to perspective C. The perspectives corresponding to video A and video B can be adjacent perspectives. Adjacent perspectives refer to situations where the shooting ranges of two or more acquisition devices 2006 are adjacent or partially overlap.
[0040] In step 202, the acquisition device 2006 can send the acquired multiple videos to the server 2004. It can be understood that the multiple videos can be sent to the server 2004 separately, or the multiple videos can be spliced together and then sent to the server 2004.
[0041] In step 203, after receiving multiple videos, server 2004 can determine the 3D Gaussian (3D GS) parameters of the scene based on multiple videos from adjacent viewpoints. A 3D Gaussian sphere is a centered, ellipsoidal object in 3D space; it uses a set of Gaussian spheres to represent scene information. The 3D Gaussian sphere can include various parameters, such as the position of the center point, the size and rotation of the Gaussian spheres, opacity, depth information, or sphere harmonic coefficients. For example, the 3D Gaussian parameters can be determined based on video A and video B, and based on video B and video C. The two 3D Gaussian parameters can contain different parameters; different 3D Gaussian parameters are used to characterize the 3D representation of the scene within a certain viewpoint range.
[0042] In some embodiments of this disclosure, to obtain a better quality rendered image, aggregated images can be used to store multiple attribute information. To save transmission resources and improve transmission efficiency, these attribute information can be quantized; for example, three-dimensional Gaussian parameters can be quantized to determine the quantized three-dimensional Gaussian parameters. In some embodiments of this disclosure, an encoder can be used to encode multiple aggregated images to determine the corresponding bitstream. The encoding process may include compression. The bitstream may include multiple frames of aggregated images, and the aggregated images may store multiple quantized attribute information, such as quantized three-dimensional Gaussian parameters.
[0043] In step 204, the user terminal 2002 can determine the target display adjustment information of the screen to be viewed based on the user's interactive operations on the displayed screen. For example, when the user focuses on a specific object in the scene or is interested in a certain position in the screen, the user terminal 2002 can determine the target display adjustment information by recognizing the user's interactive operations (such as clicking, dragging, swiping, etc.).
[0044] In step 205, server 2004 can determine the bitstream corresponding to the target display adjustment information from a stored pool of bitstreams based on the target display adjustment information. It can be understood that the viewpoint corresponding to the target position or angle is included within the corresponding viewpoint area in the bitstream.
[0045] In step 206, the server 2004 can send the corresponding bitstream to the user terminal 2002. The various quantized attribute information stored in the bitstream can be transmitted to the user terminal 2002 in a compressed manner.
[0046] In step 207, the user terminal 2002 can render the corresponding image based on the received bitstream. For example, it can render the corresponding image based on the three-dimensional Gaussian parameters stored in the image. For example, the pixel color corresponding to each pixel in the image can be determined according to the three-dimensional Gaussian parameters, thereby generating a complete two-dimensional image.
[0047] This approach allows for the dynamic generation and display of the desired viewing position or perspective based on user interactions, making the presentation and switching of the displayed image more flexible and personalized, further optimizing the user's viewing experience. Furthermore, 3D Gaussian parameters can effectively express scene information, and their generation is faster. Based on this, rendering new views in complex scenes using 3D Gaussian parameters is highly efficient and produces high-quality images.
[0048] Figure 3 illustrates a flowchart of an image processing method 300 according to some embodiments of the present disclosure. Method 300 can be executed by the user terminal 102 or server 106 shown in Figure 1. As shown in Figure 3, at block 302, method 300 may include determining target display adjustment information corresponding to the screen to be displayed based on interactive operations on the display screen. The display screen can be a screen displaying text content, image content, video content, or any combination thereof. In some examples, the display screen can be a screen displaying a certain scene. This screen can be a screen corresponding to a certain angle of the scene. The scene can be a live broadcast scene, such as a game live broadcast scene, a concert live broadcast scene, a football match scene, etc. It is understood that different users may have different viewing needs. For example, when watching a concert live broadcast scene, a user may want to switch from the view in front of the stage to the view from the side of the stage; or, when watching a football match, a user may want to switch the view to the perspective behind the goal to better capture the moment of a goal. Based on this, display adjustment information for the screen to be displayed can be determined according to the user's viewing needs. The display adjustment information can be related information about the screen the user wants to view, such as including but not limited to the angle or position corresponding to the screen. Here, angle can be a specific viewpoint value, used to describe the direction of the scene the user wants to view. Position can be a two-dimensional coordinate point on the screen, used to indicate the area or object of interest to the user. In this way, by adjusting the angle or position of the screen, users can view the scene from different angles and obtain richer visual information.
[0049] In some embodiments of this disclosure, users can interact with the displayed screen using a touchscreen, mouse, keyboard, gesture recognition, or other interactive methods. These interactions may include clicking, dragging, swiping, zooming, rotating, etc., to indicate the angle or position of the scene the user wants to view. In other embodiments, image recognition technology can be used to determine the user's gaze based on an image containing the user's face. The angle or position of the screen the user wants to view is then determined based on the user's gaze.
[0050] In box 304, method 300 may include determining an aggregated image corresponding to the target display adjustment information based on the target display adjustment information. The aggregated image is an image obtained from multiple videos that stores various quantized attribute information. Each of the multiple videos corresponds to multiple adjacent acquisition viewpoints. The various quantized attribute information is stored in the aggregated image through aggregation. The various quantized attribute information includes quantized three-dimensional Gaussian parameters, which are determined based on multiple images contained in the multiple videos. The three-dimensional Gaussian parameters are used to characterize the scene contained in the multiple videos, and the video content of the multiple videos corresponds to the target display adjustment information. The multiple videos with adjacent acquisition viewpoints can partially or jointly cover the scene within the same viewpoint range. The image determined from the multiple videos may store various attribute information, which can be used to characterize the scene within that viewpoint range. The various attribute information may include quantized three-dimensional Gaussian parameters, quantized depth information, quantized color information, etc.
[0051] In some embodiments of this disclosure, the quantized three-dimensional Gaussian parameters refer to the parameters determined after quantization of the three-dimensional Gaussian parameters. Quantization can be a process of approximating the continuous values (or a large number of possible discrete values) of a signal into a finite number (or a small number) of discrete values. The three-dimensional Gaussian parameters can explicitly characterize the corresponding scene. It can be understood that the three-dimensional Gaussian parameters are used to represent relevant information about a centered ellipsoidal object, and multiple ellipsoids (Gaussian spheres) can be used to represent scene information. The three-dimensional Gaussian parameters can be related parameters of a three-dimensional Gaussian function. The three-dimensional Gaussian function can include various parameters, such as the position of the center point, the size and rotation of the Gaussian sphere, opacity, depth information, or sphere harmonic coefficients. The depth information of a pixel refers to the distance from the pixel to the three-dimensional Gaussian function that intersects with it.
[0052] In some embodiments of this disclosure, the three-dimensional Gaussian parameters can be determined based on images contained in two videos from adjacent viewpoints. For example, the three-dimensional Gaussian parameters corresponding to the first frame image 1 in video 1 and the first frame image 2 in video 2 can be determined based on the first frame image 1 in video 1 and the first frame image 2 in video 2. The acquisition viewpoints of video 1 and video 2 are adjacent. For example, feature points or feature regions can be extracted from the first frame image 1 and the first frame image 2, and a correspondence can be established between these feature points or feature regions. Based on this correspondence, a transformation matrix from one view (first frame image 1) to another view (first frame image 2) can be determined. Based on this, three-dimensional point cloud data can be reconstructed according to the transformation matrix and multiple matching pixels. On the three-dimensional point cloud, the shape or distribution of objects contained in the image is described by fitting Gaussian parameters. The parameters included in this Gaussian function can be three-dimensional Gaussian parameters.
[0053] In some embodiments of this disclosure, to further improve the quality of the rendered image, other image attributes can be considered during the rendering process. For example, the RGB colors of pixels in the image, and the masks corresponding to multiple objects in the image. To save transmission resources during the transmission of these attributes, multiple attribute information can be aggregated to obtain aggregated attribute information. That is, the aggregated image stores multiple attribute information, and each attribute information can provide supporting data for rendering the image from a new perspective.
[0054] In box 306, method 300 includes generating a viewable image based on the image and displaying the viewable image. For example, the user terminal can render the viewable image based on various attribute information stored in the image. For example, a three-dimensional Gaussian parameter (ellipse) can be projected onto a two-dimensional image space (ellipse), and the corresponding image can be rendered based on the two-dimensional Gaussian parameter. After the corresponding image is rendered, it can be displayed on the screen of the user terminal (e.g., user terminal 102 shown in Figure 1). In this way, during a live concert broadcast, users can switch to a close-up shot of their favorite singer or band; or in online education, users can adjust the screen angle to obtain a better learning perspective.
[0055] This approach allows for the dynamic generation and display of the desired viewing position or perspective based on user interactions, making the presentation and switching of the displayed image more flexible and personalized, further optimizing the user's viewing experience. Moreover, the images used to generate the image contain various attribute information, more accurately synthesizing and presenting the scene the user wants to see, resulting in a more realistic and three-dimensional image. Furthermore, 3D Gaussian parameters can clearly express scene information, and their generation is also faster. Based on this, rendering new views in complex scenes using 3D Gaussian parameters is highly efficient and produces high-quality images.
[0056] Figure 4 illustrates a schematic diagram of acquiring multi-view video according to some embodiments of the present disclosure. As shown in Figure 4, multiple cameras (e.g., n cameras), such as camera 1, camera 2, camera 3, camera 4, camera n, and camera n+1, can be arranged around object 402. Different cameras are located at different shooting positions, with different shooting ranges and shooting angles, resulting in different captured images. In this way, information about object 402 can be captured from different angles. It can be understood that different cameras can acquire different videos at the same frame rate. The video can include multiple consecutive m frames of images. That is, different cameras can capture one frame of image at the same point in time, thereby ensuring that multiple videos are synchronized in time. For example, camera 1 can capture image 404 at time A, camera 2 can capture image 406 at time A, camera 3 can capture image 408 at time A, camera n-1 can capture image 410 at time A, and camera n can capture image 412 at time A. In some embodiments of this disclosure, the videos captured by camera 1 and camera 2 can be videos from adjacent viewpoints, and the videos captured by camera 2 and camera 3 can also be videos from adjacent viewpoints. Alternatively, the videos captured by camera 1, camera 2, and camera 3 can also be videos from adjacent viewpoints.
[0057] In this way, different images or videos can be captured from different positions and perspectives, thus providing rich information for subsequent image processing and bitstream determination.
[0058] Figure 5 illustrates a schematic diagram of an exemplary process for generating images according to some embodiments of the present disclosure. As shown in Figure 5, a multi-view acquisition module can be used to acquire multiple videos from multiple perspectives of an object or scene. The multi-view acquisition module may include n+1 cameras, which can acquire n+1 videos (e.g., video 1, video 2, video 3... video n, video n+1, etc.) from n+1 perspectives. The multiple videos from multiple perspectives can be stored or uploaded to a content delivery network. The content delivery network can be integrated into a server (e.g., server 2004 shown in Figure 2) or it can be a separate distributed server.
[0059] As shown in Figure 5, the Gaussian generation module can infer corresponding Gaussian data from multiple videos from adjacent viewpoints. This Gaussian data can be data used to characterize the three-dimensional information of a scene. In other words, Gaussian data is essentially the Gaussian distribution of a scene or object in three-dimensional space. For example, Gaussian data 1 within the viewpoint region can be inferred from video 1 and video 2. Gaussian data 2 within the viewpoint region can be inferred from video 2 and video 3. Gaussian data n within the viewpoint region can be inferred from video n and video n+1. In some embodiments of this disclosure, the Gaussian compression module can encode the Gaussian data to obtain corresponding bitstreams. The encoding process can include quantization and compression. Quantization refers to quantizing the parameters contained in the Gaussian data or quantizing other attribute information in the image (such as RGB color, resolution, etc.). For example, by encoding Gaussian data 1, bitstream 1 is obtained; by encoding Gaussian data 2, bitstream 2 is obtained; and by encoding Gaussian data n, bitstream n is obtained. The bitstream can contain multiple frames of images. Gaussian data is stored in the images, meaning that Gaussian data is stored in the pixels of the images. For example, pixels in an image can be assigned quantized Gaussian data and other quantized attributes. In this way, the three-dimensional information of the scene can be stored on the pixels of a two-dimensional image. By transmitting this two-dimensional image, sufficient data support can be provided for rendering while saving transmission resources and improving transmission efficiency.
[0060] It is understood that the Gaussian generation module and the Gaussian compression module can be located in a server (e.g., server 2004 shown in Figure 2), and the determination of the bitstream is completed by the server communicating with the user terminal and the acquisition device. In some embodiments of this disclosure, after encoding multiple bitstreams, the multiple bitstreams can be transmitted to a content delivery network for subsequent user terminals to retrieve the corresponding bitstreams. For example, users can select the desired viewing position or perspective according to their needs. Based on this position or perspective, the corresponding bitstream can be retrieved from the content delivery network. In some embodiments of this disclosure, after determining the corresponding bitstream, the bitstream can be decoded to determine the image contained in the bitstream and the Gaussian data contained in the image. Based on the Gaussian data, an image under a new perspective or position can be generated. After rendering the corresponding image, the image can be displayed to the audience.
[0061] Figure 6 illustrates a schematic diagram of determining image attribute information according to some embodiments of the present disclosure. As shown in Figure 6, the image may include various attributes, which may be basic image attributes or three-dimensional Gaussian parameters determined based on multiple videos from adjacent viewpoints. The basic image attributes may include, but are not limited to, resolution, size, brightness, color channels, etc. It is understood that, to better represent scene information, a mask corresponding to the image can be used as an attribute. The mask can be determined based on the segmentation result obtained by segmenting the image (e.g., semantic segmentation, instance segmentation, etc.). The segmentation result may include the category corresponding to each pixel in the image. The region formed by pixels belonging to the same object can be the mask of that object. In some examples, attribute 1 can be RGB color, attribute 2 can be the opacity of the three-dimensional Gaussian sphere, attribute 3 can be the size of the three-dimensional Gaussian sphere, attribute 4 can be the mask, and attribute 5 can be depth information.
[0062] In some embodiments of this disclosure, to save transmission resources and improve transmission efficiency, different quantization methods can be used to quantize different attributes, resulting in quantized attributes (e.g., quantized attribute 1, quantized attribute 2, quantized attribute 3, quantized attribute 4, quantized attribute 5). Quantization refers to the process of converting continuous or discrete attribute values with a large range into discrete attribute values with a smaller range. Quantization methods can include various types, such as linear quantization, nonlinear quantization, adaptive quantization, etc. Linear quantization can linearly map one data interval to another. This mapping method can preserve the relative size and relationships in the original data while changing the absolute size or range of the data. For example, each value in the original data interval [a, b] can be mapped to a new data interval [c, d] according to a preset linear relationship.
[0063] In one example, linear quantization can be used for attribute 1 (RGB color) to reduce the number of colors and decrease data complexity. For example, floating-point RGB color values can be quantized to a specified numerical range, such as [0, 255] or [0, 127]. The quantized numerical range can be determined based on the number of bytes storing attribute 1. For example, with 8 bits of storage, RGB color values can be quantized to the [0, 255] range. In some embodiments of this disclosure, attribute 2 (opacity) can be quantized from a continuous range to a discrete range, for example, opacity can be quantized to discrete values within the range of 0-100. In some embodiments of this disclosure, the quantization method for attribute 5 (depth information) differs from that of other attributes. When quantizing depth information, dilation can be applied to the depth information to preserve edge precision. For example, the three values obtained after quantizing a depth value can be stored in the Y, U, and V channels respectively.
[0064] It should be noted that during the quantization process, the quantization step size and quantization parameters can be determined by the user based on actual quantization needs or transmission efficiency. The quantization step size refers to the step size of attribute changes during quantization. Quantization parameters refer to the coarseness of the quantization; the larger the quantization parameters, the coarser the quantization and the higher the data compression rate.
[0065] As shown in Figure 6, after quantizing multiple attributes, corresponding quantized attributes can be determined, such as quantized attribute 1, quantized attribute 2, quantized attribute 3, quantized attribute 4, and quantized attribute 5. To integrate information and improve transmission efficiency, multiple quantized attributes can be aggregated and concatenated. After obtaining the aggregated and concatenated quantized attribute information, to further save transmission resources, the aggregated and concatenated quantized attribute information can be compressed to obtain the corresponding bitstream. For example, a video compressor can be used to compress the aggregated and concatenated quantized attribute information to obtain the corresponding bitstream. The bitstream can contain multiple aggregated images, with various quantized attribute information added to each aggregated image. It can be understood that the information contained in this bitstream can, to a certain extent, characterize the information of the scene.
[0066] In some embodiments of this disclosure, the compressed bitstream can be stored in a content delivery network (CDN). It is understood that the CDN can be a standalone server or integrated into a server used to determine the bitstream. The CDN can store multiple bitstreams, with different bitstreams corresponding to different viewpoint regions. As shown in Figure 5, data distribution is performed based on interaction information transmitted from the user terminal. The user terminal can retrieve the corresponding bitstream to its local storage based on the interaction information. For example, the user can interact with the user terminal, and the interaction methods can include eye tracking, gesture recognition, text input, etc. Based on the interaction information, the user terminal can determine the position of the screen it wants to view. After determining the position of the screen the user wants to view, the corresponding bitstream can be determined using spatial calculation methods. For example, the desired viewing angle can be determined based on the screen position, thereby determining the bitstream for the corresponding viewing angle.
[0067] As shown in Figure 6, after retrieving the corresponding bitstream, the user terminal or server can decode the bitstream to determine the various quantization attributes contained within it. It can be understood that these multiple quantization attributes are aggregated and spliced information. To determine one of the multiple quantization attributes, they can be separated after decoding to identify quantization attribute 1, quantization attribute 2, quantization attribute 3, quantization attribute 4, and quantization attribute 5. To restore information as close as possible to the original video or image, each quantization attribute can be dequantized to determine its corresponding attribute. In some embodiments of this disclosure, the dequantization method can correspond to the quantization method. For example, the dequantization steps match the quantization steps, and the dequantization parameters also match the quantization parameters (e.g., the dequantization parameters can be the inverse mapping of the quantization parameters in linear quantization). That is, when dequantizing different quantization attributes, the correct dequantization method can be determined based on the quantization method used for that attribute, and the dequantization parameters can be adjusted accordingly. Based on this, the original attribute value can be restored as accurately as possible.
[0068] By selecting appropriate quantization and dequantization methods in this way, effective compression and transmission can be achieved while ensuring the quality of the bitstream, thereby further improving the efficiency of bitstream transmission.
[0069] Figure 7 illustrates a schematic diagram of attribute information contained in an image according to some embodiments of the present disclosure. As shown in Figure 7, the aggregated image 702 (e.g., the image after aggregation and stitching in Figure 6) can store five types of attribute information, such as RGB attributes, opacity, depth, rotation, and mask. For example, the aggregated image 702 can be divided into multiple regions, and pixels in each region can store one attribute. For example, pixels in region 1 can store RGB attributes, and pixels in region 2 can store opacity. It should be noted that pixels in each region describe the same type of attribute information. In some embodiments of the present disclosure, the aggregated image 702 stores multiple types of attribute information, and these multiple types of attribute information are transmitted through aggregation and stitching. It is understood that other types of attribute information can also be added to the aggregated image 702. The attribute information shown in this disclosure is for illustrative purposes only and is not intended to limit the scope of protection of this disclosure.
[0070] In some embodiments of this disclosure, the corresponding image can be rendered based on the three-dimensional Gaussian parameters stored in the image. For example, the three-dimensional Gaussian functions intersecting each pixel in the image can be determined based on the three-dimensional Gaussian parameters. These Gaussian functions are then sorted based on their depth information (e.g., the distance from the center point of the Gaussian function to the camera's viewpoint). For example, the Gaussian functions can be sorted in ascending order. For each Gaussian function in the sorted list, its contribution to the pixel color can be determined. For example, its color can be adjusted based on the opacity of the three-dimensional Gaussian function and its relative position to the pixel. In some embodiments of this disclosure, after determining the contribution of each Gaussian function to the pixel color, a blending technique (e.g., alpha blending) can be used to accumulate the contributions of all Gaussian functions to the pixel color, generating the final pixel color. In this process, potential occlusion relationships between each Gaussian function can be considered, thereby generating a more accurate pixel color. Based on this, the above process is repeated for all pixels in the image, ultimately generating a complete two-dimensional image where each pixel contains the synthesized color value.
[0071] Figure 8 illustrates a schematic diagram of a newly rendered viewpoint according to some embodiments of the present disclosure. As shown in Figure 8, the user can drag and drop on the original live stream screen 802 displayed on the user terminal to select the desired viewing location based on actual needs. For example, the user can move the screen by touching it to select the desired viewing location. For example, the endpoint of the movement can be determined based on the movement trajectory, and the desired viewing location or viewing angle can be determined based on the endpoint. In one example, after determining the desired viewing location, the corresponding bitstream can be retrieved from the content delivery network, and through a series of decoding, separation, and dequantization operations, various attribute information for rendering the screen can be determined. Based on the various attribute information, the user terminal can render a new viewpoint 804, which is the desired viewing location. In this way, the live stream screen moves in real time according to the user's operation, allowing the user to select and view different locations within the screen.
[0072] Figure 9 shows a block diagram of an image processing apparatus 900 according to some embodiments of the present disclosure. As shown in Figure 9, the apparatus 900 includes a display adjustment information determination unit 902, configured to determine target display adjustment information corresponding to a display screen based on an interactive operation on the display screen. The apparatus 900 also includes an aggregated image determination unit 904, configured to determine an aggregated image corresponding to the target display adjustment information based on the target display adjustment information. The aggregated image is an image obtained from multiple videos that stores multiple quantized attribute information. The multiple videos each correspond to multiple adjacent acquisition viewpoints. The multiple quantized attribute information is stored in the aggregated image in an aggregated manner. The multiple quantized attribute information includes quantized three-dimensional Gaussian parameters, wherein the three-dimensional Gaussian parameters are determined based on multiple images contained in the multiple videos. The three-dimensional Gaussian parameters are used to characterize the scene contained in the multiple videos, and the video content of the multiple videos corresponds to the target display adjustment information. The apparatus 900 also includes a display screen generation unit 906, configured to generate a display screen to be displayed and display the display screen based on the aggregated image.
[0073] It is understood that by utilizing the apparatus 900 of this disclosure, at least one of the many advantages that can be achieved by the methods or processes described above can be realized.
[0074] Figure 10 shows a block diagram of an image processing apparatus 1000 according to some other embodiments of the present disclosure. As shown in Figure 10, the apparatus 1000 includes a display adjustment information receiving unit 1002, configured to receive target display adjustment information corresponding to a screen to be displayed from a user terminal. The apparatus 1000 also includes an aggregated image determining unit 1004, configured to determine an aggregated image corresponding to the target display adjustment information based on the target display adjustment information. The aggregated image is an image obtained from multiple videos that stores multiple quantized attribute information. The multiple acquisition perspectives corresponding to the multiple videos are adjacent. The multiple quantized attribute information is stored in the aggregated image in an aggregated manner. The multiple quantized attribute information includes quantized three-dimensional Gaussian parameters, wherein the three-dimensional Gaussian parameters are determined based on multiple images contained in the multiple videos. The three-dimensional Gaussian parameters are used to characterize the scene contained in the multiple videos. The video content of the multiple videos corresponds to the target display adjustment information. The apparatus 1000 also includes an aggregated image sending unit 1006, configured to send the aggregated image to the user terminal, so that the user terminal generates and displays the screen to be displayed.
[0075] It is understood that by utilizing the apparatus 900 of this disclosure, at least one of the many advantages that can be achieved by the methods or processes described above can be realized.
[0076] Figure 11 shows a schematic block diagram of an example device 1100 that can be used to implement embodiments of the present disclosure. As shown, device 1100 includes a computing unit 1101, which can perform various appropriate actions and processes according to computer program instructions stored in read-only memory (ROM) 1102 or loaded from storage unit 1108 into random access memory (RAM) 1103. Various programs and data required for the operation of device 1100 may also be stored in RAM 1103. The computing unit 1101, ROM 1102, and RAM 1103 are interconnected via bus 1104. Input / output (I / O) interface 1105 is also connected to bus 1104.
[0077] Multiple components in device 1100 are connected to I / O interface 1105, including: input unit 1106, such as keyboard, mouse, etc.; output unit 1107, such as various types of monitors, speakers, etc.; storage unit 1108, such as disk, optical disk, etc.; and communication unit 1109, such as network card, modem, wireless transceiver, etc. Communication unit 1109 allows device 1100 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.
[0078] The computing unit 1101 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 1101 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 1101 performs the various methods and processes described above, such as method 300. For example, in some embodiments, method 300 may be implemented as a computer software program tangibly contained in a machine-readable medium, such as storage unit 1108. In some embodiments, part or all of the computer program may be loaded and / or installed on device 1100 via ROM 1102 and / or communication unit 1109. When the computer program is loaded into RAM 1103 and executed by the computing unit 1101, one or more steps of method 300 described above may be performed. Alternatively, in other embodiments, the computing unit 1101 may be configured to execute method 300 by any other suitable means (e.g., by means of firmware).
[0079] The functions described above in this document can be performed at least in part by one or more hardware logic components. For example, exemplary types of hardware logic components that can be used, without limitation, include: field programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload programmable logic devices (CPLDs), and so on.
[0080] The program code used to implement the methods of this disclosure may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.
[0081] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. Machine-readable media can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing. Furthermore, although operations are depicted in a specific order, this should be understood as requiring that such operations be performed in the specific order shown or in sequential order, or requiring that all illustrated operations be performed to achieve the desired result. In certain environments, multitasking and parallel processing may be advantageous. Similarly, while several specific implementation details are included in the foregoing discussion, these should not be construed as limiting the scope of this disclosure. Certain features described in the context of individual embodiments may also be implemented in combination in a single implementation. Conversely, various features described in the context of a single implementation may also be implemented individually or in any suitable sub-combination in multiple implementations.
[0082] The following are some example implementations of this disclosure.
[0083] Example 1. An image processing method, comprising:
[0084] Based on the interactive operations on the display screen, determine the target display adjustment information corresponding to the screen to be displayed;
[0085] Based on the target display adjustment information, a corresponding aggregated image is determined. This aggregated image is an image obtained from multiple videos, storing various quantized attribute information. Each video corresponds to multiple adjacent acquisition viewpoints. The various quantized attribute information is stored in the aggregated image through aggregation. This quantized attribute information includes quantized three-dimensional Gaussian parameters, which are determined based on multiple images contained in the videos. These three-dimensional Gaussian parameters characterize the scenes contained in the videos, and the video content of the multiple videos corresponds to the target display adjustment information.
[0086] Based on the aggregated image, generate and display the screen to be displayed.
[0087] Example 2. Following the method in Example 1, the target display adjustment information includes the target angle and target position, and the target adjustment information corresponding to the screen to be displayed includes:
[0088] In response to the original displayed image being moved, the movement trajectory of the original displayed image is obtained; and
[0089] The endpoint of the movement trajectory of the original displayed screen is taken as the target position of the screen to be displayed.
[0090] Example 3. Following the method in Example 1 or 2, generating the screen to be displayed includes:
[0091] By decoding the aggregated image, the various quantized attribute information stored in the aggregated image is determined;
[0092] By separating the aggregated multiple quantitative attribute information, each quantitative attribute information in the multiple quantitative attribute information is determined;
[0093] Based on multiple inverse quantization methods, the attribute information corresponding to each quantized attribute information is determined; and
[0094] Based on the attribute information corresponding to each quantified attribute, a display screen is generated.
[0095] Example 4. Using the method from any of Examples 1-3, where multiple quantization attribute information includes quantized depth information, quantized color information, and quantized 3D Gaussian parameters, and the attribute information corresponding to each quantization attribute information includes:
[0096] Based on the first preset inverse quantization method, determine the depth information corresponding to the quantized depth information;
[0097] Based on the second preset inverse quantization method, determine the color information corresponding to the quantized color information; and
[0098] Based on the third preset inverse quantization method, the three-dimensional Gaussian parameters corresponding to the quantized three-dimensional Gaussian parameters are determined.
[0099] Example 5. Based on any one of Examples 1-4, the method for generating the screen to be displayed includes:
[0100] The two-dimensional Gaussian parameters corresponding to the three-dimensional Gaussian parameters are determined by projecting the three-dimensional Gaussian parameters onto the image plane;
[0101] The aggregated image is uniformly divided into a preset number of regions, and the depth information of the two-dimensional Gaussian parameters in each region is determined.
[0102] Based on depth information, the order of the two-dimensional Gaussian parameters in each region is determined;
[0103] Based on the arrangement order and three-dimensional Gaussian parameters, determine the pixel values corresponding to the pixels in the aggregated image; and
[0104] The image to be displayed is generated based on the pixel values.
[0105] Example 6. The method is based on any one of Examples 1-5, where the three-dimensional Gaussian parameters include opacity parameters and rotation parameters.
[0106] Example 7. An image processing method, comprising:
[0107] Receive target display adjustment information corresponding to the screen to be displayed from the user terminal;
[0108] Based on the target display adjustment information, an aggregated image corresponding to the target display adjustment information is determined. The aggregated image is an image obtained from multiple videos that stores various quantized attribute information. The multiple acquisition viewpoints corresponding to the multiple videos are adjacent. The various quantized attribute information is stored in the aggregated image through aggregation. The various quantized attribute information includes quantized three-dimensional Gaussian parameters, which are determined based on multiple images contained in the multiple videos. The three-dimensional Gaussian parameters are used to characterize the scenes contained in the multiple videos, and the video content of the multiple videos corresponds to the target display adjustment information; and
[0109] The system sends an aggregated image to the user terminal, which then generates and displays the image to be displayed.
[0110] Example 8. Based on the method in Example 7, it also includes:
[0111] Acquire multiple videos captured from multiple perspectives of a target scene, each video comprising multiple consecutive frames; and
[0112] Based on multiple images from multiple videos at the same time in adjacent viewpoints, the three-dimensional Gaussian parameters corresponding to the multiple images are determined. The three-dimensional Gaussian parameters are used to characterize the target scene.
[0113] Example 9. Following the method of Example 7 or 8, it also includes:
[0114] Obtain depth and color information from multiple images;
[0115] By quantizing depth information, color information, and 3D Gaussian parameters, the quantized depth information, quantized color information, and quantized 3D Gaussian parameters are determined; and
[0116] By aggregating and quantizing depth information, color information, and quantized 3D Gaussian parameters, an aggregated image storing multiple aggregated quantized attribute information is determined.
[0117] Example 10. Using the method from any of Examples 7-9, determining the quantized depth information, quantized color information, and the quantized 3D Gaussian parameters includes:
[0118] Based on the first preset quantization method, determine the quantized depth information corresponding to the depth information;
[0119] Based on the second preset quantization method, the quantized color information corresponding to the color information is determined; and
[0120] Based on the third preset quantization method, the quantized three-dimensional Gaussian parameters corresponding to the three-dimensional Gaussian parameters are determined.
[0121] Example 11. The method according to any one of Examples 7-10 also includes:
[0122] The aggregated image is compressed using a video encoder to determine the compressed aggregated image.
[0123] Example 12. According to the method of any one of Examples 7-11, wherein the target display adjustment information includes the target location, and determining the aggregated image corresponding to the target display adjustment information includes:
[0124] Based on the pre-built correspondence between image positions and aggregated images, an aggregated image corresponding to the target position is selected from multiple aggregated images, and the viewpoint corresponding to the target position is included in the viewpoint area corresponding to the aggregated image.
[0125] Example 13. An image processing apparatus, comprising:
[0126] The display adjustment information determination unit is configured to determine the target display adjustment information corresponding to the screen to be displayed based on the interactive operation on the display screen;
[0127] The aggregated image determination unit is configured to determine an aggregated image corresponding to the target display adjustment information based on the target display adjustment information. The aggregated image is an image obtained from multiple videos that stores various quantized attribute information. The multiple acquisition viewpoints corresponding to the multiple videos are adjacent. The various quantized attribute information is stored in the aggregated image through aggregation. The various quantized attribute information includes quantized three-dimensional Gaussian parameters, which are determined based on multiple images contained in the multiple videos. The three-dimensional Gaussian parameters are used to characterize the scenes contained in the multiple videos, and the video content of the multiple videos corresponds to the target display adjustment information; and
[0128] The display screen generation unit is configured to generate and display the screen to be displayed based on the aggregated image.
[0129] Example 14. The apparatus according to Example 13, wherein the target display adjustment information includes the target angle and the target position, and the display adjustment information determining unit is further configured to:
[0130] In response to the original displayed image being moved, the movement trajectory of the original displayed image is obtained; and
[0131] The endpoint of the movement trajectory of the original displayed screen is taken as the target position of the screen to be displayed.
[0132] Example 15. The apparatus according to Example 13 or 14, wherein the display screen generation unit is further configured to:
[0133] By decoding the aggregated image, the various quantized attribute information stored in the aggregated image is determined;
[0134] By separating the aggregated multiple quantitative attribute information, each quantitative attribute information in the multiple quantitative attribute information is determined;
[0135] Based on multiple inverse quantization methods, the attribute information corresponding to each quantized attribute information is determined; and
[0136] Based on the attribute information corresponding to each quantified attribute, a display screen is generated.
[0137] Example 16. An apparatus according to any one of Examples 13-15, wherein multiple quantization attribute information includes quantized depth information, quantized color information, and quantized three-dimensional Gaussian parameters, and the attribute information corresponding to each quantization attribute information includes:
[0138] Based on the first preset inverse quantization method, determine the depth information corresponding to the quantized depth information;
[0139] Based on the second preset inverse quantization method, determine the color information corresponding to the quantized color information; and
[0140] Based on the third preset inverse quantization method, the three-dimensional Gaussian parameters corresponding to the quantized three-dimensional Gaussian parameters are determined.
[0141] Example 17. An apparatus according to any one of Examples 13-16, wherein the display screen generation unit is further configured as follows:
[0142] The two-dimensional Gaussian parameters corresponding to the three-dimensional Gaussian parameters are determined by projecting the three-dimensional Gaussian parameters onto the image plane;
[0143] The aggregated image is uniformly divided into a preset number of regions, and the depth information of the two-dimensional Gaussian parameters in each region is determined.
[0144] Based on depth information, the order of the two-dimensional Gaussian parameters in each region is determined;
[0145] Based on the arrangement order and three-dimensional Gaussian parameters, determine the pixel values corresponding to the pixels in the aggregated image; and
[0146] The image to be displayed is generated based on the pixel values.
[0147] Example 18. An apparatus according to any one of Examples 13-17, wherein the three-dimensional Gaussian parameters include an opacity parameter and a rotation parameter.
[0148] Example 19. An image processing apparatus, comprising:
[0149] The display adjustment information determination unit is configured to receive target display adjustment information corresponding to the screen to be displayed from the user terminal.
[0150] The aggregated image determination unit is configured to determine an aggregated image corresponding to the target display adjustment information based on the target display adjustment information. The aggregated image is an image obtained from multiple videos that stores various quantized attribute information. The multiple acquisition viewpoints corresponding to the multiple videos are adjacent. The various quantized attribute information is stored in the aggregated image through aggregation. The various quantized attribute information includes quantized three-dimensional Gaussian parameters, which are determined based on multiple images contained in the multiple videos. The three-dimensional Gaussian parameters are used to characterize the scenes contained in the multiple videos, and the video content of the multiple videos corresponds to the target display adjustment information; and
[0151] The display screen generation unit is configured to send an aggregated image to the user terminal, and the user terminal generates and displays the screen to be displayed.
[0152] Example 20. The apparatus according to Example 19, wherein the apparatus is further configured as follows:
[0153] Acquire multiple videos captured from multiple perspectives of a target scene, each video comprising multiple consecutive frames; and
[0154] Based on multiple images from multiple videos at the same time in adjacent viewpoints, the three-dimensional Gaussian parameters corresponding to the multiple images are determined. The three-dimensional Gaussian parameters are used to characterize the target scene.
[0155] Example 21. The apparatus according to Example 19 or 20, wherein the apparatus is further configured as follows:
[0156] Obtain depth and color information from multiple images;
[0157] By quantizing depth information, color information, and 3D Gaussian parameters, the quantized depth information, quantized color information, and quantized 3D Gaussian parameters are determined; and
[0158] By aggregating and quantizing depth information, color information, and quantized 3D Gaussian parameters, an aggregated image storing multiple aggregated quantized attribute information is determined.
[0159] Example 22. An apparatus according to any one of Examples 19-21, wherein the apparatus is further configured as follows:
[0160] Based on the first preset quantization method, determine the quantized depth information corresponding to the depth information;
[0161] Based on the second preset quantization method, the quantized color information corresponding to the color information is determined; and
[0162] Based on the third preset quantization method, the quantized three-dimensional Gaussian parameters corresponding to the three-dimensional Gaussian parameters are determined.
[0163] Example 23. An apparatus according to any one of Examples 19-22, wherein the apparatus is further configured as follows:
[0164] The aggregated image is compressed using a video encoder to determine the compressed aggregated image.
[0165] Example 24. An apparatus according to any one of Examples 19-23, wherein the target display adjustment information includes the target position, and the aggregated image determination unit is further configured to:
[0166] Based on the pre-built correspondence between image positions and aggregated images, an aggregated image corresponding to the target position is selected from multiple aggregated images, and the viewpoint corresponding to the target position is included in the viewpoint area corresponding to the aggregated image.
[0167] Example 25. An electronic device comprising:
[0168] At least one processor; and
[0169] A memory, coupled to at least one processor and having instructions stored thereon, which, when executed by the at least one processor, cause the device to perform the following actions:
[0170] Based on the interactive operations on the display screen, determine the target display adjustment information corresponding to the screen to be displayed;
[0171] Based on the target display adjustment information, a corresponding aggregated image is determined. This aggregated image is an image obtained from multiple videos, storing various quantized attribute information. Each video corresponds to multiple adjacent acquisition viewpoints. The various quantized attribute information is stored in the aggregated image through aggregation. This quantized attribute information includes quantized three-dimensional Gaussian parameters, which are determined based on multiple images contained in the videos. These three-dimensional Gaussian parameters characterize the scenes contained in the videos, and the video content of the multiple videos corresponds to the target display adjustment information.
[0172] Based on the aggregated image, generate and display the screen to be displayed.
[0173] Example 26. The electronic device according to Example 25, wherein the target display adjustment information includes the target angle and the target position, and the target display adjustment information corresponding to the screen to be displayed includes:
[0174] In response to the original displayed image being moved, the movement trajectory of the original displayed image is obtained; and
[0175] The endpoint of the movement trajectory of the original displayed screen is taken as the target position of the screen to be displayed.
[0176] Example 27. Based on the electronic device of Example 25 or 26, generating the screen to be displayed includes:
[0177] By decoding the aggregated image, the various quantized attribute information stored in the aggregated image is determined;
[0178] By separating the aggregated multiple quantitative attribute information, each quantitative attribute information in the multiple quantitative attribute information is determined;
[0179] Based on multiple inverse quantization methods, the attribute information corresponding to each quantized attribute information is determined; and
[0180] Based on the attribute information corresponding to each quantified attribute, a display screen is generated.
[0181] Example 28. An electronic device based on any one of Examples 25-27, wherein multiple quantization attribute information includes quantized depth information, quantized color information, and quantized three-dimensional Gaussian parameters, and the attribute information corresponding to each quantization attribute information includes:
[0182] Based on the first preset inverse quantization method, determine the depth information corresponding to the quantized depth information;
[0183] Based on the second preset inverse quantization method, determine the color information corresponding to the quantized color information; and
[0184] Based on the third preset inverse quantization method, the three-dimensional Gaussian parameters corresponding to the quantized three-dimensional Gaussian parameters are determined.
[0185] Example 29. An electronic device based on any one of Examples 25-28, wherein generating the screen to be displayed includes:
[0186] The two-dimensional Gaussian parameters corresponding to the three-dimensional Gaussian parameters are determined by projecting the three-dimensional Gaussian parameters onto the image plane;
[0187] The aggregated image is uniformly divided into a preset number of regions, and the depth information of the two-dimensional Gaussian parameters in each region is determined.
[0188] Based on depth information, the order of the two-dimensional Gaussian parameters in each region is determined;
[0189] Based on the arrangement order and three-dimensional Gaussian parameters, determine the pixel values corresponding to the pixels in the aggregated image; and
[0190] The image to be displayed is generated based on the pixel values.
[0191] Example 30. An electronic device based on any of Examples 25-29, wherein the three-dimensional Gaussian parameters include an opacity parameter and a rotation parameter.
[0192] Example 31. An electronic device comprising:
[0193] At least one processor; and
[0194] A memory, coupled to at least one processor and having instructions stored thereon, which, when executed by the at least one processor, cause the device to perform the following actions:
[0195] Receive target display adjustment information corresponding to the screen to be displayed from the user terminal;
[0196] Based on the target display adjustment information, an aggregated image corresponding to the target display adjustment information is determined. The aggregated image is an image obtained from multiple videos that stores various quantized attribute information. The multiple acquisition viewpoints corresponding to the multiple videos are adjacent. The various quantized attribute information is stored in the aggregated image through aggregation. The various quantized attribute information includes quantized three-dimensional Gaussian parameters, which are determined based on multiple images contained in the multiple videos. The three-dimensional Gaussian parameters are used to characterize the scenes contained in the multiple videos, and the video content of the multiple videos corresponds to the target display adjustment information; and
[0197] The system sends an aggregated image to the user terminal, which then generates and displays the image to be displayed.
[0198] Example 32. According to the electronic device of Example 31, the action also includes:
[0199] Acquire multiple videos captured from multiple perspectives of a target scene, each video comprising multiple consecutive frames; and
[0200] Based on multiple images from multiple videos at the same time in adjacent viewpoints, the three-dimensional Gaussian parameters corresponding to the multiple images are determined. The three-dimensional Gaussian parameters are used to characterize the target scene.
[0201] Example 33. For an electronic device according to Example 31 or 32, the action also includes:
[0202] Obtain depth and color information from multiple images;
[0203] By quantizing depth information, color information, and 3D Gaussian parameters, the quantized depth information, quantized color information, and quantized 3D Gaussian parameters are determined; and
[0204] By aggregating and quantizing depth information, color information, and quantized 3D Gaussian parameters, an aggregated image storing multiple aggregated quantized attribute information is determined.
[0205] Example 34. An electronic device according to any one of Examples 31-33, wherein determining the quantized depth information, quantized color information, and quantized three-dimensional Gaussian parameters includes:
[0206] Based on the first preset quantization method, determine the quantized depth information corresponding to the depth information;
[0207] Based on the second preset quantization method, the quantized color information corresponding to the color information is determined; and
[0208] Based on the third preset quantization method, the quantized three-dimensional Gaussian parameters corresponding to the three-dimensional Gaussian parameters are determined.
[0209] Example 35. For any of the electronic devices in Examples 31-34, the actions also include:
[0210] The aggregated image is compressed using a video encoder to determine the compressed aggregated image.
[0211] Example 36. An electronic device according to any one of Examples 31-35, wherein the target display adjustment information includes the target location, and determining the aggregated image corresponding to the target display adjustment information includes:
[0212] Based on the pre-built correspondence between image positions and aggregated images, an aggregated image corresponding to the target position is selected from multiple aggregated images, and the viewpoint corresponding to the target position is included in the viewpoint area corresponding to the aggregated image.
[0213] Example 37. A computer program product, tangibly stored on a non-transitory computer-readable medium and comprising machine-executable instructions, which, when executed, cause a machine to perform the following actions:
[0214] Based on the interactive operations on the display screen, determine the target display adjustment information corresponding to the screen to be displayed;
[0215] Based on the target display adjustment information, a corresponding aggregated image is determined. This aggregated image is an image obtained from multiple videos, storing various quantized attribute information. Each video corresponds to multiple adjacent acquisition viewpoints. The various quantized attribute information is stored in the aggregated image through aggregation. This quantized attribute information includes quantized three-dimensional Gaussian parameters, which are determined based on multiple images contained in the videos. These three-dimensional Gaussian parameters characterize the scenes contained in the videos, and the video content of the multiple videos corresponds to the target display adjustment information.
[0216] Based on the aggregated image, generate and display the screen to be displayed.
[0217] Example 38. According to the computer program product of Example 37, the target display adjustment information includes the target angle and the target position, and the target display adjustment information corresponding to the screen to be displayed includes:
[0218] In response to the original displayed image being moved, the movement trajectory of the original displayed image is obtained; and
[0219] The endpoint of the movement trajectory of the original displayed screen is taken as the target position of the screen to be displayed.
[0220] Example 39. A computer program product based on Example 37 or 38, wherein generating a screen to be displayed includes:
[0221] By decoding the aggregated image, the various quantized attribute information stored in the aggregated image is determined;
[0222] By separating the aggregated multiple quantitative attribute information, each quantitative attribute information in the multiple quantitative attribute information is determined;
[0223] Based on multiple inverse quantization methods, the attribute information corresponding to each quantized attribute information is determined; and
[0224] Based on the attribute information corresponding to each quantified attribute, a display screen is generated.
[0225] Example 40. A computer program product based on any one of Examples 37-39, wherein multiple quantized attribute information includes quantized depth information, quantized color information, and quantized 3D Gaussian parameters, and the attribute information corresponding to each quantized attribute information includes:
[0226] Based on the first preset inverse quantization method, determine the depth information corresponding to the quantized depth information;
[0227] Based on the second preset inverse quantization method, determine the color information corresponding to the quantized color information; and
[0228] Based on the third preset inverse quantization method, the three-dimensional Gaussian parameters corresponding to the quantized three-dimensional Gaussian parameters are determined.
[0229] Example 41. A computer program product based on any one of Examples 37-40, wherein generating the screen to be displayed includes:
[0230] The two-dimensional Gaussian parameters corresponding to the three-dimensional Gaussian parameters are determined by projecting the three-dimensional Gaussian parameters onto the image plane;
[0231] The aggregated image is uniformly divided into a preset number of regions, and the depth information of the two-dimensional Gaussian parameters in each region is determined.
[0232] Based on depth information, the order of the two-dimensional Gaussian parameters in each region is determined;
[0233] Based on the arrangement order and three-dimensional Gaussian parameters, determine the pixel values corresponding to the pixels in the aggregated image; and
[0234] The image to be displayed is generated based on the pixel values.
[0235] Example 42. A computer program product based on any one of Examples 37-41, wherein the three-dimensional Gaussian parameters include an opacity parameter and a rotation parameter.
[0236] Example 43. A computer program product, tangibly stored on a non-transitory computer-readable medium and comprising machine-executable instructions, which, when executed, cause a machine to perform the following actions:
[0237] Receive target display adjustment information corresponding to the screen to be displayed from the user terminal;
[0238] Based on the target display adjustment information, an aggregated image corresponding to the target display adjustment information is determined. The aggregated image is an image obtained from multiple videos that stores various quantized attribute information. The multiple acquisition viewpoints corresponding to the multiple videos are adjacent. The various quantized attribute information is stored in the aggregated image through aggregation. The various quantized attribute information includes quantized three-dimensional Gaussian parameters, which are determined based on multiple images contained in the multiple videos. The three-dimensional Gaussian parameters are used to characterize the scenes contained in the multiple videos, and the video content of the multiple videos corresponds to the target display adjustment information; and
[0239] The system sends an aggregated image to the user terminal, which then generates and displays the image to be displayed.
[0240] Example 44. According to the computer program product of Example 43, the actions also include:
[0241] Acquire multiple videos captured from multiple perspectives of a target scene, each video comprising multiple consecutive frames; and
[0242] Based on multiple images from multiple videos at the same time in adjacent viewpoints, the three-dimensional Gaussian parameters corresponding to the multiple images are determined. The three-dimensional Gaussian parameters are used to characterize the target scene.
[0243] Example 45. According to the computer program product of Example 43 or 44, the actions also include:
[0244] Obtain depth and color information from multiple images;
[0245] By quantizing depth information, color information, and 3D Gaussian parameters, the quantized depth information, quantized color information, and quantized 3D Gaussian parameters are determined; and
[0246] By aggregating and quantizing depth information, color information, and quantized 3D Gaussian parameters, an aggregated image storing multiple aggregated quantized attribute information is determined.
[0247] Example 46. A computer program product based on any one of Examples 43-45, wherein determining the quantized depth information, the quantized color information, and the quantized 3D Gaussian parameters includes:
[0248] Based on the first preset quantization method, determine the quantized depth information corresponding to the depth information;
[0249] Based on the second preset quantization method, the quantized color information corresponding to the color information is determined; and
[0250] Based on the third preset quantization method, the quantized three-dimensional Gaussian parameters corresponding to the three-dimensional Gaussian parameters are determined.
[0251] Example 47. For a computer program product based on any one of Examples 43-46, the actions also include:
[0252] The aggregated image is compressed using a video encoder to determine the compressed aggregated image.
[0253] Example 48. A computer program product according to any one of Examples 43-47, wherein the target display adjustment information includes a target location, and determining the aggregated image corresponding to the target display adjustment information includes:
[0254] Based on the pre-built correspondence between image positions and aggregated images, an aggregated image corresponding to the target position is selected from multiple aggregated images, and the viewpoint corresponding to the target position is included in the viewpoint area corresponding to the aggregated image.
[0255] Although the subject matter has been described using language specific to structural features and / or methodological logic, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or actions described above. Rather, the specific features and actions described above are merely illustrative examples of implementing the claims.
Claims
1. An image processing method, comprising: Based on the interactive operations on the display screen, determine the target display adjustment information corresponding to the screen to be displayed; Based on the target display adjustment information, an aggregated image corresponding to the target display adjustment information is determined. The aggregated image is an image obtained from multiple videos, storing various quantized attribute information. Each of the multiple videos corresponds to multiple adjacent acquisition viewpoints. The various quantized attribute information is stored in the aggregated image through aggregation. The various quantized attribute information includes quantized three-dimensional Gaussian parameters, which are determined based on multiple images contained in the multiple videos. The three-dimensional Gaussian parameters are used to characterize the scenes contained in the multiple videos, and the video content of the multiple videos corresponds to the target display adjustment information. Based on the aggregated image, the screen to be displayed is generated and then displayed.
2. The method according to claim 1, wherein the target display adjustment information includes the target angle and the target position, and determining the target adjustment information corresponding to the screen to be displayed includes: In response to the original display screen being moved, the movement trajectory of the original display screen is obtained; as well as The endpoint of the movement trajectory of the original displayed screen is taken as the target position corresponding to the screen to be displayed.
3. The method according to claim 1, wherein generating the screen to be displayed comprises: By decoding the aggregated image, the aggregated multiple quantized attribute information stored in the aggregated image is determined; By separating the aggregated multiple quantized attribute information, each quantized attribute information among the multiple quantized attribute information is determined; Based on multiple inverse quantization methods, the attribute information corresponding to each quantization attribute information is determined; as well as The screen to be displayed is generated based on the attribute information corresponding to each quantized attribute information.
4. The method according to claim 3, wherein the plurality of quantized attribute information includes quantized depth information, quantized color information, and the quantized three-dimensional Gaussian parameters, and determining the attribute information corresponding to each quantized attribute information includes: Based on the first preset inverse quantization method, the depth information corresponding to the quantization depth information is determined; Based on the second preset inverse quantization method, the color information corresponding to the quantized color information is determined; as well as Based on the third preset inverse quantization method, the three-dimensional Gaussian parameters corresponding to the quantized three-dimensional Gaussian parameters are determined.
5. The method according to claim 4, wherein generating the screen to be displayed comprises: The two-dimensional Gaussian parameters corresponding to the three-dimensional Gaussian parameters are determined by projecting the three-dimensional Gaussian parameters onto the image plane; The aggregated image is uniformly divided into a preset number of regions, and the depth information of the two-dimensional Gaussian parameters in each region is determined. Based on the depth information, the arrangement order of the two-dimensional Gaussian parameters in each region is determined; Based on the arrangement order and the three-dimensional Gaussian parameters, the pixel values corresponding to the pixels in the aggregated image are determined; as well as The image to be displayed is generated based on the pixel values.
6. The method according to claim 4, wherein the three-dimensional Gaussian parameters include an opacity parameter and a rotation parameter.
7. An image processing method, comprising: Receive target display adjustment information corresponding to the screen to be displayed from the user terminal; Based on the target display adjustment information, an aggregated image corresponding to the target display adjustment information is determined. The aggregated image is an image obtained from multiple videos that stores multiple quantized attribute information. The multiple acquisition perspectives corresponding to the multiple videos are adjacent. The multiple quantized attribute information is stored in the aggregated image in an aggregated manner. The multiple quantized attribute information includes quantized three-dimensional Gaussian parameters, wherein the three-dimensional Gaussian parameters are determined based on multiple images contained in the multiple videos. The three-dimensional Gaussian parameters are used to characterize the scene contained in the multiple videos. The video content of the multiple videos corresponds to the target display adjustment information. as well as The aggregated image is sent to the user terminal, which then generates and displays the screen to be displayed.
8. The method according to claim 7, further comprising: Acquire multiple videos captured from multiple perspectives of a target scene, each of the multiple videos comprising multiple consecutive frames of images; as well as Based on multiple images at the same time in multiple videos from adjacent viewpoints, the three-dimensional Gaussian parameters corresponding to the multiple images are determined, and the three-dimensional Gaussian parameters are used to characterize the target scene.
9. The method according to claim 8, further comprising: Obtain the depth and color information corresponding to the multiple images; The quantized depth information, quantized color information, and quantized three-dimensional Gaussian parameters are determined by quantizing the depth information, the color information, and the three-dimensional Gaussian parameters. as well as By aggregating the quantized depth information, quantized color information, and the quantized three-dimensional Gaussian parameters, an aggregated image storing multiple aggregated quantized attribute information is determined.
10. The method according to claim 9, wherein determining the quantized depth information, the quantized color information, and the quantized three-dimensional Gaussian parameters comprises: Based on the first preset quantization method, the quantized depth information corresponding to the depth information is determined; Based on the second preset quantization method, the quantized color information corresponding to the color information is determined; as well as Based on the third preset quantization method, the quantized three-dimensional Gaussian parameters corresponding to the three-dimensional Gaussian parameters are determined.
11. The method of claim 9, further comprising: The aggregated image is compressed using a video encoder to determine the compressed aggregated image.
12. The method of claim 7, wherein the target display adjustment information includes a target location, and determining the aggregated image corresponding to the target display adjustment information includes: Based on a pre-built correspondence between image positions and aggregated images, an aggregated image corresponding to the target position is selected from multiple aggregated images, and the viewpoint corresponding to the target position is included in the viewpoint area corresponding to the aggregated image.
13. An image processing apparatus, comprising: The display adjustment information determination unit is configured to determine the target display adjustment information corresponding to the screen to be displayed based on the interactive operation on the display screen; An aggregated image determination unit is configured to determine an aggregated image corresponding to the target display adjustment information based on the target display adjustment information. The aggregated image is an image obtained from multiple videos that stores multiple quantized attribute information. The multiple acquisition viewpoints corresponding to the multiple videos are adjacent. The multiple quantized attribute information is stored in the aggregated image in an aggregated manner. The multiple quantized attribute information includes quantized three-dimensional Gaussian parameters, wherein the three-dimensional Gaussian parameters are determined based on multiple images contained in the multiple videos. The three-dimensional Gaussian parameters are used to characterize the scene contained in the multiple videos. The video content of the multiple videos corresponds to the target display adjustment information. as well as The display screen generation unit is configured to generate the screen to be displayed and display the screen to be displayed based on the aggregated image.
14. An image processing apparatus, comprising: The display adjustment information receiving unit is configured to receive target display adjustment information corresponding to the screen to be displayed from the user terminal. An aggregated image determination unit is configured to determine an aggregated image corresponding to the target display adjustment information based on the target display adjustment information. The aggregated image is an image obtained from multiple videos that stores multiple quantized attribute information. The multiple acquisition viewpoints corresponding to the multiple videos are adjacent. The multiple quantized attribute information is stored in the aggregated image in an aggregated manner. The multiple quantized attribute information includes quantized three-dimensional Gaussian parameters, wherein the three-dimensional Gaussian parameters are determined based on multiple images contained in the multiple videos. The three-dimensional Gaussian parameters are used to characterize the scene contained in the multiple videos. The video content of the multiple videos corresponds to the target display adjustment information. as well as The aggregated image sending unit is configured to send the aggregated image to the user terminal, and the user terminal generates and displays the screen to be displayed.
15. An electronic device comprising: At least one processor; as well as A memory coupled to the at least one processor and having instructions stored thereon, which, when executed by the at least one processor, cause the device to perform the method according to any one of claims 1-12.
16. A computer program product tangibly stored on a non-transitory computer-readable medium and comprising machine-executable instructions that, when executed, cause a machine to perform the method according to any one of claims 1 to 12.
Citation Information
Patent Citations
Point cloud data generation method and device based on 3D Gaussian splashing, equipment and medium
CN118102044A
View angle synthesis method of 3D Gaussian splashing technology based on space coupling
CN118674905A
System and method for 3D modeling
US20240265630A1