Image rendering method for an object

By jointly training a hybrid neural radiation field model and a super-resolution model, the problem of low rendering efficiency of the neural radiation field model is solved, achieving efficient image rendering and quality improvement.

CN115690291BActive Publication Date: 2026-07-14ALIBABA (CHINA) CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
ALIBABA (CHINA) CO LTD
Filing Date
2022-11-15
Publication Date
2026-07-14

AI Technical Summary

Technical Problem

Existing neural radiation field models are prone to detail blurring and fog-like noise when rendering images, resulting in low rendering efficiency.

Method used

A joint training method combining a hybrid neural radiation field model and a super-resolution model is adopted. By rendering low-resolution images into high-resolution images, the hybrid neural radiation field model stores image information in a voxel grid, and the super-resolution model is used to amplify the resolution of the rendered image.

Benefits of technology

It improves image rendering efficiency, solves the problem of low rendering efficiency, and enhances image quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115690291B_ABST
    Figure CN115690291B_ABST
Patent Text Reader

Abstract

The application discloses an image rendering method of an object. The method comprises the following steps: starting image monitoring on an object to be monitored, and acquiring an original image monitored under a current view angle; calling a hybrid neural radiance field model and a super-resolution model generated after joint training; rendering the original image by using the hybrid neural radiance field model to obtain a first rendered image monitored under the current view angle, wherein the resolution of the first rendered image is less than a first resolution threshold; and performing resolution amplification processing on the first rendered image by using the super-resolution model to obtain a second rendered image monitored under the current view angle, wherein the resolution of the second rendered image is higher than a second resolution threshold. The application solves the technical problem of low image rendering efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of image processing, and more specifically, to an image rendering method for an object. Background Technology

[0002] As algorithms and hardware related to 3D scenes mature, concepts such as virtual reality, augmented reality, and the metaverse have gained popularity. Among the 3D scene-related algorithms developed in recent years, neural radiation fields are relatively convenient to use, requiring only multiple photos taken from different angles with a regular mobile phone for the same object. They also offer high-quality rendering effects. Therefore, neural radiation fields have become one of the fastest-growing 3D algorithms.

[0003] Currently, neural radiation field models have many problems in use. For example, the rendered images are prone to problems such as blurred details and fog-like noise, resulting in low rendering efficiency.

[0004] There is currently no effective solution to the above problems. Summary of the Invention

[0005] This application provides an image rendering method for an object, which at least solves the technical problem of low image rendering efficiency.

[0006] According to one aspect of the embodiments of this application, an image rendering method for an object is provided. The method may include: initiating image monitoring of the object to be monitored, acquiring an original image monitored from the current viewpoint, wherein the resolution of the original image is lower than a first resolution threshold; invoking a hybrid neural radiation field model and a super-resolution model generated after joint training, wherein the samples used during joint training are a first image sample set captured from different viewpoints, the resolution of the image samples in the first image sample set is higher than a second resolution threshold, the second resolution threshold is greater than the first resolution threshold, and the hybrid neural radiation field model is used to store at least the target image information of the first image sample set in a voxel grid, the frequency of the target image information being lower than a frequency threshold; rendering the original image using the hybrid neural radiation field model to obtain a first rendered image monitored from the current viewpoint, wherein the resolution of the first rendered image is lower than the first resolution threshold; and amplifying the resolution of the first rendered image using the super-resolution model to obtain a second rendered image monitored from the current viewpoint, wherein the resolution of the second rendered image is higher than the second resolution threshold.

[0007] According to another aspect of the embodiments of this application, another image rendering method for an object is also provided. The method may include: acquiring a first image sample set of an object to be monitored captured from different viewpoints, wherein the resolution of the image samples in the first image sample set is higher than a resolution threshold; jointly training an initial hybrid neural radiation field model and an initial super-resolution model using the first image sample set to obtain a hybrid neural radiation field model and a super-resolution model, wherein the hybrid neural radiation field model is used to store at least the target image information of the first image sample set in a voxel grid, the frequency of the target image information is lower than a frequency threshold, and is used to render the original image of the object to be monitored monitored from the current viewpoint to obtain a first rendered image monitored from the current viewpoint, wherein the resolution of the original image is lower than another resolution threshold, and the other resolution threshold is lower than a resolution threshold; the super-resolution model is used to enlarge the resolution of the first rendered image to obtain a second rendered image monitored from the current viewpoint.

[0008] According to another aspect of the embodiments of this application, another image rendering method for an object is also provided. The method may include: initiating image monitoring of the object to be monitored; displaying the original image monitored from the current viewpoint on the display screen of a virtual reality (VR) device or an augmented reality (AR) device, wherein the resolution of the original image is lower than a first resolution threshold; invoking a hybrid neural radiation field model and a super-resolution model generated after joint training, wherein the samples used during joint training are a first image sample set captured from different viewpoints, the resolution of the image samples in the first image sample set is higher than a second resolution threshold, the second resolution threshold is greater than the first resolution threshold, and the hybrid neural radiation field model is used to store at least the target image information of the first image sample set in a voxel grid, the frequency of the target image information being lower than a frequency threshold; controlling the VR device or AR device to render the original image using the hybrid neural radiation field model to obtain a first rendered image monitored from the current viewpoint, wherein the resolution of the first rendered image is lower than the first resolution threshold; controlling the VR device or AR device to enlarge the resolution of the first rendered image using the super-resolution model to obtain a second rendered image monitored from the current viewpoint, wherein the resolution of the second rendered image is higher than the second resolution threshold; and driving the VR device or AR device to render and display the second rendered image.

[0009] According to another aspect of the embodiments of this application, another image rendering method for an object is also provided. The method may include: acquiring an original image monitored from the current viewpoint by calling a first interface, wherein the first interface includes a first parameter, the value of which is the original image, and the resolution of the original image is lower than a first resolution threshold; calling a hybrid neural radiation field model and a super-resolution model generated after joint training, wherein the samples used during joint training are a first image sample set captured from different viewpoints, the resolution of the image samples in the first image sample set is higher than a second resolution threshold, the second resolution threshold is greater than the first resolution threshold, and the hybrid neural radiation field model is used to store at least the target image information of the first image sample set in a voxel grid, the frequency of the target image information being lower than a frequency threshold; rendering the original image using the hybrid neural radiation field model to obtain a first rendered image monitored from the current viewpoint, wherein the resolution of the first rendered image is lower than the first resolution threshold; amplifying the resolution of the first rendered image using the super-resolution model to obtain a second rendered image monitored from the current viewpoint, wherein the resolution of the second rendered image is higher than the second resolution threshold; and outputting the second rendered image by calling a second interface, wherein the second interface includes a second parameter, the value of which is the second rendered image.

[0010] According to one aspect of the embodiments of this application, an image rendering apparatus for an object is provided. The apparatus may include: a first acquisition unit, configured to initiate image monitoring of the object to be monitored and acquire an original image monitored from the current viewpoint, wherein the resolution of the original image is lower than a first resolution threshold; a first invocation unit, configured to invoke a hybrid neural radiation field model and a super-resolution model generated after joint training, wherein the samples used during joint training are a first image sample set captured from different viewpoints, the resolution of the image samples in the first image sample set is higher than a second resolution threshold, the second resolution threshold is greater than the first resolution threshold, and the hybrid neural radiation field model is used to store at least the target image information of the first image sample set in a voxel grid, the frequency of the target image information being lower than a frequency threshold; a first rendering unit, configured to render the original image using the hybrid neural radiation field model to obtain a first rendered image monitored from the current viewpoint, wherein the resolution of the first rendered image is lower than the first resolution threshold; and a first processing unit, configured to use the super-resolution model to amplify the resolution of the first rendered image to obtain a second rendered image monitored from the current viewpoint, wherein the resolution of the second rendered image is higher than the second resolution threshold.

[0011] According to another aspect of the embodiments of this application, an image rendering apparatus for another object is also provided. The apparatus may include: a second acquisition unit, configured to acquire a first image sample set captured from different viewpoints of the object to be monitored, wherein the resolution of the image samples in the first image sample set is higher than a resolution threshold; and a training unit, configured to jointly train an initial hybrid neural radiation field model and an initial super-resolution model using the first image sample set to obtain a hybrid neural radiation field model and a super-resolution model, wherein the hybrid neural radiation field model is used to store at least the target image information of the first image sample set in a voxel grid, the frequency of the target image information is lower than a frequency threshold, and is used to render the original image of the object to be monitored detected at the current viewpoint to obtain a first rendered image detected at the current viewpoint, the resolution of the original image being lower than another resolution threshold, the other resolution threshold being lower than a resolution threshold; and the super-resolution model is used to enlarge the resolution of the first rendered image to obtain a second rendered image detected at the current viewpoint.

[0012] According to another aspect of the embodiments of this application, an image rendering apparatus for another object is also provided. The device may include: a display unit for initiating image monitoring of the object to be monitored and displaying the original image monitored from the current viewpoint on the display screen of a virtual reality (VR) device or an augmented reality (AR) device, wherein the resolution of the original image is lower than a first resolution threshold; a second invocation unit for invoking a hybrid neural radiation field model and a super-resolution model generated after joint training, wherein the samples used during joint training are a first image sample set captured from different viewpoints, the resolution of the image samples in the first image sample set is higher than a second resolution threshold, the second resolution threshold is greater than the first resolution threshold, and the hybrid neural radiation field model is used to store at least the target image information of the first image sample set in a voxel grid, wherein the frequency of the target image information is lower than a frequency threshold; a first control unit for controlling the VR device or AR device to render the original image using the hybrid neural radiation field model to obtain a first rendered image monitored from the current viewpoint, wherein the resolution of the first rendered image is lower than the first resolution threshold; a second control unit for controlling the VR device or AR device to enlarge the resolution of the first rendered image using the super-resolution model to obtain a second rendered image monitored from the current viewpoint, wherein the resolution of the second rendered image is higher than the second resolution threshold; and a driving unit for driving the VR device or AR device to render and display the second rendered image.

[0013] According to another aspect of the embodiments of this application, an image rendering apparatus for another object is also provided. The apparatus may include: a third calling unit, configured to acquire a raw image monitored from the current viewpoint by calling a first interface, wherein the first interface includes a first parameter, the parameter value of which is the raw image, and the resolution of the raw image is lower than a first resolution threshold; and a fourth calling unit, configured to call a hybrid neural radiation field model and a super-resolution model generated after joint training, wherein the samples used during joint training are a first image sample set captured from different viewpoints, the resolution of the image samples in the first image sample set is higher than a second resolution threshold, the second resolution threshold is greater than the first resolution threshold, and the hybrid neural radiation field model is used to at least render the target image information of the first image sample set. The information is stored in a voxel grid, and the frequency of the target image information is less than a frequency threshold; the second rendering unit is used to render the original image using a hybrid neural radiation field model to obtain a first rendered image monitored at the current viewpoint, wherein the resolution of the first rendered image is less than a first resolution threshold; the second processing unit is used to enlarge the resolution of the first rendered image using a super-resolution model to obtain a second rendered image monitored at the current viewpoint, wherein the resolution of the second rendered image is higher than a second resolution threshold; the fifth calling unit is used to output the second rendered image by calling a second interface, wherein the second interface includes a second parameter, and the value of the second parameter is the second rendered image.

[0014] According to another aspect of the embodiments of this application, a computer-readable storage medium is also provided, the computer-readable storage medium including a stored program, wherein, when the program is running, it controls the device where the storage medium is located to execute the image rendering method of any of the above-mentioned objects.

[0015] According to another aspect of the embodiments of this application, a processor is also provided, which is used to run a program, wherein the image rendering method of any of the above-mentioned objects is executed during program execution.

[0016] In this embodiment, image monitoring of the target object is initiated to acquire the original image monitored from the current viewpoint, wherein the resolution of the original image is lower than a first resolution threshold; a hybrid neural radiation field model and a super-resolution model generated after joint training are invoked, wherein the samples used during joint training are a first image sample set captured from different viewpoints, the resolution of the image samples in the first image sample set is higher than a second resolution threshold, the second resolution threshold is greater than the first resolution threshold, and the hybrid neural radiation field model is used to store at least the target image information of the first image sample set in a voxel grid, wherein the frequency of the target image information is lower than a frequency threshold; the hybrid neural radiation field model is used to render the original image to obtain a first rendered image monitored from the current viewpoint, wherein the resolution of the first rendered image is lower than the first resolution threshold; the super-resolution model is used to enlarge the resolution of the first rendered image to obtain a second rendered image monitored from the current viewpoint, wherein the resolution of the second rendered image is higher than the second resolution threshold. In other words, the embodiments of this application use a hybrid neural radiation field model and a super-resolution model for joint training, so that the optimization of the parameters of the hybrid neural radiation field model and the super-resolution model influence each other. The hybrid neural radiation field model and the super-resolution model generated after joint training process the original image, thereby achieving the technical effect of improving the rendering efficiency of the image and solving the technical problem of low image rendering efficiency. Attached Figure Description

[0017] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, illustrate exemplary embodiments and are used to explain this application, but do not constitute an undue limitation of this application. In the drawings:

[0018] Figure 1 This is a schematic diagram of the hardware environment of a virtual reality device for image rendering of an object according to an embodiment of this application;

[0019] Figure 2 This is a structural block diagram of a computing environment for image rendering of an object according to an embodiment of this application;

[0020] Figure 3 This is a flowchart of an image rendering method for an object according to an embodiment of this application;

[0021] Figure 4 This is a flowchart of another image rendering method for an object according to an embodiment of this application;

[0022] Figure 5 This is a flowchart of another image rendering method for an object according to an embodiment of this application;

[0023] Figure 6 This is a schematic diagram of a rendering result according to an embodiment of this application;

[0024] Figure 7 This is a flowchart of another image rendering method for an object according to an embodiment of this application;

[0025] Figure 8 This is a schematic diagram of a super-resolution algorithm architecture for neural radiation fields based on related technologies;

[0026] Figure 9 This is a schematic diagram of a super-resolution direct pixel algorithm architecture according to an embodiment of this application;

[0027] Figure 10 This is a structural block diagram of a service mesh for an image rendering method of an object according to an embodiment of this application;

[0028] Figure 11 This is a schematic diagram of an image rendering apparatus for an object according to an embodiment of this application;

[0029] Figure 12 This is a schematic diagram of an image rendering apparatus for another object according to an embodiment of this application;

[0030] Figure 13 This is a schematic diagram of an image rendering apparatus for another object according to an embodiment of this application;

[0031] Figure 14 This is a schematic diagram of an image rendering apparatus for another object according to an embodiment of this application;

[0032] Figure 15 This is a structural block diagram of a computer terminal according to an embodiment of this application. Detailed Implementation

[0033] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present application, and not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative effort should fall within the scope of protection of the present application.

[0034] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0035] First, some nouns or terms that appear in the description of the embodiments of this application shall be interpreted as follows:

[0036] Neural Radiance Field (NeRF) is a three-dimensional new perspective synthesis algorithm. Given a finite number of photos of any object or scene from different perspectives, the algorithm can obtain photos of the object or scene viewed from any new perspective.

[0037] Implicit Neural Radiance Field (INRF) can be obtained by using purely implicit mapping (neural network) to realize the neural radiation field that associates spatial coordinates with scene attributes.

[0038] Explicit Neural Radiance Field (NERF) can be used to represent scene properties using a spatially location-dependent 3D mesh.

[0039] Hybrid Neural Radiation Field (HNRF): This can be a neural radiation field that combines neural networks and three-dimensional meshes simultaneously;

[0040] Direct Voxel Grid Optimization (DVGO) can be a type of hybrid neural radiation field with an ultrafast convergence method.

[0041] Super Resolution (SR) is an algorithm used to improve image resolution, thereby enlarging the image and enhancing its clarity.

[0042] Generative Adversarial Networks (GANs) are algorithms for generating signals. Given a number of images that follow a certain distribution, the algorithm can generate new images that conform to that distribution.

[0043] A multi-layer perceptron (MLP) can be a basic neural network that can perform non-linear transformations on input information to better complete subsequent tasks.

[0044] Rectified Linear Units (ReLUs) are one of the most commonly used activation functions in neural networks, and can be used to introduce nonlinearity into the network.

[0045] Example 1

[0046] According to an embodiment of this application, an image rendering method for an object is provided. It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. Furthermore, although a logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order than that shown here.

[0047] Figure 1 This is a schematic diagram of the hardware environment of a virtual reality device for image rendering of an object according to an embodiment of this application. For example... Figure 1 As shown, the virtual reality device 104 is connected to the terminal 106, and the terminal 106 is connected to the server 102 via a network. The virtual reality device 104 is not limited to: virtual reality helmets, virtual reality glasses, virtual reality all-in-one machines, etc. The terminal 104 is not limited to PCs, mobile phones, tablets, etc. The server 102 can be a server corresponding to a media file operator. The network includes, but is not limited to: wide area network, metropolitan area network, or local area network.

[0048] Optionally, the virtual reality device 104 in this embodiment includes a memory, a processor, and a transmission device. The memory stores an application program that can be used to execute: initiate image monitoring of the object to be monitored, acquire the original image monitored from the current viewpoint, wherein the resolution of the original image is lower than a first resolution threshold; invoke a hybrid neural radiation field model and a super-resolution model generated after joint training, wherein the samples used during joint training are a first image sample set captured from different viewpoints, the resolution of the image samples in the first image sample set is higher than a second resolution threshold, and the hybrid neural radiation field model is used to store at least the low-frequency information of the first image sample set in a voxel grid; render the original image using the hybrid neural radiation field model to obtain a first rendered image monitored from the current viewpoint, wherein the resolution of the first rendered image is lower than the first resolution threshold; and enlarge the resolution of the first rendered image using the super-resolution model to obtain a second rendered image monitored from the current viewpoint, wherein the resolution of the second rendered image is higher than the second resolution threshold, thereby solving the technical problem of low image rendering efficiency and achieving the goal of improving image rendering efficiency.

[0049] The terminal in this embodiment can be used to: display the original image monitored from the current viewpoint on the presentation screen of a virtual reality device or an augmented reality device; invoke a hybrid neural radiation field model and a super-resolution model generated after joint training, wherein the samples used during joint training are a first image sample set captured from different viewpoints, the resolution of the image samples in the first image sample set is higher than a second resolution threshold, and the hybrid neural radiation field model is used to store at least the low-frequency information of the first image sample set in a voxel grid; control the virtual reality (VR) device or augmented reality (AR) device to render the original image using the hybrid neural radiation field model to obtain a first rendered image monitored from the current viewpoint, wherein the resolution of the first rendered image is lower than the first resolution threshold; control the VR device or AR device to use the super-resolution model to enlarge the resolution of the first rendered image to obtain a second rendered image monitored from the current viewpoint, wherein the resolution of the second rendered image is higher than the second resolution threshold; and send the second rendered image to the virtual reality device 104, which displays it at the target projection location after receiving the second rendered image.

[0050] Optionally, the virtual reality device 104 in this embodiment includes an eye-tracking head-mounted display (HMD) and an eye-tracking module that function the same as in the embodiments described above. That is, the screen in the HMD displays real-time images, and the eye-tracking module in the HMD acquires the real-time movement path of the user's eyes. In this embodiment, the terminal acquires the user's position and movement information in real three-dimensional space through a tracking system, and calculates the three-dimensional coordinates of the user's head in virtual three-dimensional space, as well as the user's field of vision orientation in virtual three-dimensional space.

[0051] Figure 1 The hardware structure block diagram shown can serve not only as an exemplary block diagram of the aforementioned AR / VR device (or mobile device), but also as an exemplary block diagram of the aforementioned server. In one optional embodiment, Figure 2 The use of the above is illustrated in a block diagram. Figure 1 The AR / VR device (or mobile device) shown is an example of a computing node in computing environment 201. Figure 2 This is a structural block diagram of a computing environment for image rendering of an object according to an embodiment of this application, such as... Figure 2 As shown, computing environment 201 includes multiple computing nodes (such as servers) running on a distributed network (shown as 210-1, 210-2, ..., in the diagram). Each computing node contains local processing and memory resources, and end user 202 can remotely run applications or store data within computing environment 201. Applications can be provided as multiple services 220-1, 220-2, 220-3, and 220-4 within computing environment 201, representing services "A", "D", "E", and "H", respectively.

[0052] End user 202 can provide and access services through a web browser or other software application on a client. In some embodiments, the provisioning and / or requests of end user 202 can be provided to ingress gateway 230. Ingress gateway 230 may include a corresponding agent to handle the provisioning and / or requests for services (one or more services provided in computing environment 201).

[0053] Services are provided or deployed based on various virtualization technologies supported by the computing environment 201. In some embodiments, services may be provided based on virtual machine (VM)-based virtualization, container-based virtualization, and / or similar methods. VM-based virtualization can simulate a real computer by initializing a virtual machine, executing programs and applications without directly accessing any actual hardware resources. While the machine is virtualized by a virtual machine, container-based virtualization can launch containers to virtualize an entire operating system (OS), allowing multiple workloads to run on a single OS instance.

[0054] In one embodiment based on container virtualization, several containers of a service can be assembled into a Pod (e.g., a Kubernetes Pod). For example, such as Figure 2 As shown, service 220-2 can be equipped with one or more Pods 240-1, 240-2, ..., 240-N (collectively referred to as Pods). Each Pod can include a proxy 245 and one or more containers 242-1, 242-2, ..., 242-M (collectively referred to as containers). One or more containers in a Pod handle requests related to one or more corresponding functions of the service. The proxy 245 typically controls service-related network functions such as routing and load balancing. Other services can also be Pods similar to Pods.

[0055] During operation, executing a user request from end user 202 may require calling one or more services in computing environment 201, and executing one or more functions of one service may require calling one or more functions of another service. For example... Figure 2 As shown, service "A" 220-1 receives user requests from terminal user 202 from ingress gateway 230. Service "A" 220-1 can call service "D" 220-2, and service "D" 220-2 can request service "E" 220-3 to perform one or more functions.

[0056] The aforementioned computing environment can be a cloud computing environment, where resource allocation is managed by cloud services, allowing functionality development without needing to consider implementation, adjustment, or server scaling. This computing environment allows developers to execute event-responsive code without building or maintaining complex infrastructure. Services can be partitioned into a set of functions that can automatically and independently scale, rather than scaling a single hardware device to handle potential loads.

[0057] Under the aforementioned operating environment, this application provides the following: Figure 3 The image rendering method for the object shown is illustrated. It should be noted that the image rendering method for the object in this embodiment can be provided by... Figure 1The mobile terminal in the illustrated embodiment is executed. Figure 3 This is a flowchart of an image rendering method for an object according to an embodiment of this application. For example... Figure 3 As shown, the method may include the following steps:

[0058] Step S302: Start image monitoring of the object to be monitored and acquire the original image monitored from the current viewpoint, wherein the resolution of the original image is lower than the first resolution threshold.

[0059] In the technical solution provided by step S302 of this application, image monitoring of the object to be monitored can be initiated to obtain the original image monitored from the current viewpoint. The original image can be a low-resolution image, for example, an image with a resolution of 1K. The resolution of the original image can be lower than a first resolution threshold, which can be a predetermined value or a value stipulated by laws and regulations.

[0060] Optionally, since videos with a resolution of 4K or higher are defined as ultra-high-definition videos, the first resolution threshold can be preset to 2K, so the original video can be an image with a resolution of 1K.

[0061] Step S304: Invoke the hybrid neural radiation field model and super-resolution model generated after joint training. The samples used during joint training are the first image sample set captured from different viewpoints. The resolution of the image samples in the first image sample set is higher than the second resolution threshold, the second resolution threshold is greater than the first resolution threshold, and the hybrid neural radiation field model is used to store the target image information of the first image sample set in a voxel grid. The frequency of the target image information is less than the frequency threshold.

[0062] In the technical solution provided in step S304 of this application, the hybrid neural radiation field and super-resolution model generated after joint training can be invoked. The hybrid neural radiation field model can be a Direct Voxel Grid Optimization (DVGO) model, which can train and converge on a 1K resolution image in 10 minutes. It should be noted that the hybrid neural radiation field model can also be other models, or it can be a combination of the DVGO model and other hybrid neural radiation field models (e.g., Instant-NGP, Tensor RF, etc.). This is only an example and no specific limitations are imposed on the hybrid neural radiation field model. The super-resolution model (SR) can be a simplified version of the real-ESRGAN algorithm. The parameters of the super-resolution model can be 1 / 5 of the original real-ESRGAN. It should be noted that this is only an example and no specific limitations are imposed on the super-resolution model.

[0063] Optionally, image samples from different viewpoints can be captured to obtain a first image sample set. This first image sample set can be high-resolution ground truth (GT) images or training images. The resolution of the image samples in the first image sample set can be higher than a second resolution threshold. The second resolution threshold can be a threshold satisfied by the high-resolution images, for example, 4K. It should be noted that in this embodiment, the second resolution threshold is greater than the aforementioned first resolution threshold. A hybrid neural radiation field model can be used to store the target image information of the first image sample set in a voxel grid. The voxel grid can be used to store the target image information in the first image sample set. The target image information can be low-frequency information with a frequency lower than a frequency threshold. The frequency threshold can be a preset value or a value selected based on actual conditions, for example, 1. If the image information has a frequency less than 1, it can be determined that the image information is the target image information, and the target image information can be stored in the voxel grid. Low-frequency information can be used to characterize areas in the image sample where brightness or grayscale values ​​change slowly, for example, it can be used to characterize large flat areas in the image sample.

[0064] Low-frequency and high-frequency information in an image sample are relative concepts. For example, if the intensity of each position in an image sample is equal, then only low-frequency information exists in the image sample. From the spectrum image of the image, there is only one main peak, which is located at the position of frequency 0. Therefore, it can be determined that the image information in the image sample is all target image information, and the target image information can be stored in a voxel grid.

[0065] Optionally, the hybrid neural radiation field model and the super-resolution model can be jointly trained using high-resolution image samples to generate the hybrid neural radiation field model and the super-resolution model, obtain the original image monitored from the current viewpoint, and process the original image using the trained hybrid neural radiation field and super-resolution model.

[0066] Step S306: Render the original image using a hybrid neural radiation field model to obtain a first rendered image monitored from the current viewpoint, wherein the resolution of the first rendered image is less than a first resolution threshold.

[0067] In the technical solution provided in step S306 of this application, a hybrid neural radiation field model can be used to render the original image to obtain a first rendered image monitored from the current viewpoint. The first rendered image can be a low-resolution image (LR True Image), and its resolution can be less than a first resolution threshold. Rendering can be performed on the original image using voxel rendering.

[0068] For example, the acquired raw image can be processed by voxel rendering based on the jointly trained hybrid neural radiation field model, resulting in a low-resolution image with a resolution of 1K.

[0069] Step S308: Use a super-resolution model to enlarge the resolution of the first rendered image to obtain a second rendered image monitored from the current viewpoint, wherein the resolution of the second rendered image is higher than the second resolution threshold.

[0070] In the technical solution provided in step S308 of this application, a first rendered image obtained by the hybrid neural radiation field model is acquired. A super-resolution model can be used to process the resolution of the first rendered image to obtain a second rendered image monitored from the current viewpoint. The second rendered image can be a high-resolution rendered image, or an image with a resolution higher than a second resolution threshold. For example, the second resolution threshold can be 4K, and the second rendered image can be an image with a resolution greater than or equal to 4K.

[0071] For example, a first rendered image can be obtained by a hybrid neural radiation field model. The resolution of the first rendered image can be magnified using a super-resolution model. After magnifying the first rendered image by 4 times, a high-resolution image with a resolution of 4K can be obtained from the current viewpoint, which is the aforementioned second rendered image.

[0072] In this embodiment, by using a hybrid neural radiation field model and a super-resolution model for joint training, the problem of poor image quality when rendering high-resolution images in related technologies is solved, achieving the effect of reducing the computational cost of high-resolution images while improving image quality. Furthermore, since the super-resolution model is based on convolution, it can effectively utilize the neighborhood information of the image. Therefore, this method improves image quality on the one hand, and avoids the high time consumption of directly rendering high-resolution images using the neural radiation field model on the other.

[0073] Through steps S302 to S308 of this application, image monitoring of the object to be monitored is initiated to obtain the original image monitored from the current viewpoint, wherein the resolution of the original image is lower than a first resolution threshold; the hybrid neural radiation field model and the super-resolution model generated after joint training are invoked, wherein the samples used during joint training are a first image sample set captured from different viewpoints, the resolution of the image samples in the first image sample set is higher than a second resolution threshold, the second resolution threshold is greater than the first resolution threshold, and the hybrid neural radiation field model is used to store at least the target image information of the first image sample set in a voxel grid, wherein the frequency of the target image information is lower than a frequency threshold; the hybrid neural radiation field model is used to render the original image to obtain a first rendered image monitored from the current viewpoint, wherein the resolution of the first rendered image is lower than the first resolution threshold; the super-resolution model is used to enlarge the resolution of the first rendered image to obtain a second rendered image monitored from the current viewpoint, wherein the resolution of the second rendered image is higher than the second resolution threshold. In other words, the embodiments of this application use a hybrid neural radiation field model and a super-resolution model for joint training, so that the optimization of model parameters influences each other. The hybrid neural radiation field model and the super-resolution model generated after joint training process the original image, thereby achieving the technical effect of improving the rendering efficiency of the image and solving the technical problem of low image rendering efficiency.

[0074] The method described in this embodiment will be further described below.

[0075] As an optional implementation, the method further includes: using target image information to at least represent the viewpoint position of the object to be monitored when image monitoring is performed on the object to be monitored.

[0076] In this embodiment, low-frequency information can be used to at least represent the viewpoint position of the object to be monitored when image monitoring is performed on the object in the image. The object to be monitored can be an object, person, animal, etc., in the image; this is merely an example and not a specific limitation. The viewpoint position can be the location of the object to be monitored, or it can be a spatial location.

[0077] This application embodiment stores low-frequency information in a voxel grid. On the one hand, it can directly determine which spatial locations in the image contain information about the object to be monitored, avoiding invalid sampling. On the other hand, it can store the low-frequency information of the object to be monitored in the corresponding spatial locations, reducing the amount of computation in the data processing process.

[0078] As an optional implementation, the method further includes: rendering the original image using a hybrid neural radiation field model to obtain a first rendered image monitored from the current viewpoint, including: acquiring first ray sampling data in the original image, wherein the first ray sampling data is used to characterize the sampled ray corresponding to the object to be monitored in the original image; acquiring first voxel features corresponding to the first ray sampling data based on a voxel grid; and using the hybrid neural radiation field model to perform voxel rendering on the original image using the first voxel features to obtain the first rendered image.

[0079] In this embodiment, first ray sampling data can be acquired from the original image. Based on the first ray sampling data, the first voxel feature corresponding to the first ray sampling data can be obtained from the voxel mesh. A hybrid neural radiation field model can be used to perform voxel rendering on the original image using the first voxel feature to obtain a first rendered image. The first ray sampling data can be used to characterize the sampled ray (subsampling ray) corresponding to the object to be monitored in the original image. The first voxel feature can include the pixel, volume, and element information of the object to be monitored, and can be a pixel in three-dimensional space.

[0080] As an optional implementation, the method further includes: using a super-resolution model to magnify the resolution of the first rendered image to obtain a second rendered image monitored from the current viewpoint, including: acquiring second ray sampling data from image segments in the first rendered image, wherein the second ray sampling data is used to characterize the sampled rays corresponding to the monitored object in the image segments; acquiring second voxel features corresponding to the second ray sampling data based on a voxel grid; and using a super-resolution model to magnify the resolution of the rendered image corresponding to the second voxel features by a target factor to obtain the second rendered image.

[0081] In this embodiment, second ray sampling data from image patches in the first rendered image can be acquired. Second voxel features corresponding to the second ray sampling data are obtained from the voxel mesh. A super-resolution model can be used to magnify the resolution of the rendered image corresponding to the second voxel features according to a target factor, resulting in the second rendered image monitored from the current viewpoint. The second ray sampling data can be rays from image patches based on the first rendered image, and can be used to characterize the sampled rays corresponding to the monitored object within the image patch. The target factor can be a pre-set factor, such as 4x. The second voxel features can include information such as pixels, volume, and element information.

[0082] Optionally, patch-based rays generated from image patches of the first rendered image can be input into a hybrid neural radiation field model. This model performs voxel rendering on the original image, producing a first rendered image of image patches with texture information. The hybrid neural radiation field model can then transfer this first rendered image to a super-resolution model and acquire second ray sampling data from the image patches in the first rendered image. Second voxel features corresponding to the second ray sampling data are obtained based on the voxel mesh. The super-resolution model can then be used to scale up the resolution of the rendered image corresponding to the second voxel features by a target factor to obtain the second rendered image.

[0083] In this embodiment of the application, the method for rendering the image is not to enlarge the rendered image by ordinary random pixels, but to enlarge the rendered image based on the second voxel features corresponding to the second ray sampling data of the image block. This image block-based approach is beneficial for the subsequent super-resolution model to use directly, and it does not impair the randomness of training when the image block size is small.

[0084] For example, the second ray sampling data in the image slice of the first rendered image can be obtained, and the second voxel feature corresponding to the second ray sampling data can be obtained from the voxel mesh. A super-resolution model can be used to magnify the resolution of the rendered image corresponding to the second voxel feature by 4 times to obtain a second rendered image magnified by 4 times at the current viewpoint.

[0085] As an optional implementation, the method further includes: using a target multiplier to perform mixed degradation processing on image sample sets from different viewpoints with resolutions higher than a second resolution threshold to obtain a second image sample set with a resolution lower than a first resolution threshold; using the second image sample set to obtain an initial super-resolution model based on convolution training; and using the initial super-resolution model to train and obtain a super-resolution model.

[0086] In this embodiment, image sample sets from different viewpoints with resolutions higher than a second resolution threshold can be subjected to hybrid degradation processing based on a target magnification to obtain a second image sample set with a resolution lower than a first resolution threshold. This second image sample set can be used to train an initial super-resolution model based on convolution, and the initial super-resolution model can be used to train the super-resolution model. Hybrid degradation processing can be a hybrid simulated degradation method, which can involve degrading common artificially simulated images, such as adding high-frequency noise, blurring, or compressing noise. Alternatively, degradation methods can be combined in a random order and quantity to achieve hybrid degradation of the real scene. The second image sample set can be low-resolution images (Low Resolution, LR) with a resolution lower than the first resolution threshold, where the second resolution threshold is greater than the first resolution threshold.

[0087] Optionally, high-resolution images from different viewpoints with resolutions higher than a second resolution threshold can be acquired. Multiple high-resolution (HR) images are used as training images to obtain an image sample set. The target image samples can be downsampled and subjected to hybrid degradation. For example, the image sample set can be reduced by a target factor, and a hybrid degradation method can be obtained by randomly selecting the degradation method. Based on this hybrid degradation method, the image sample set is subjected to hybrid degradation processing, resulting in low-resolution images (below the first resolution threshold) that are reduced by a factor of 4 and have varying degrees of noise and blurring, thus obtaining a second image sample set. The second image samples can be used to train an initial super-resolution model based on convolution, and the initial super-resolution model can be used to train another super-resolution model.

[0088] In the embodiments of this application, the super-resolution model is obtained based on convolution training, which can make good use of the neighborhood information of the image, improve the image quality, and avoid the high time consumption required for the neural radiation field model to directly render high-resolution images, thereby achieving the technical effect of improving the rendering efficiency of the model.

[0089] As an optional implementation, the method further includes: using an initial hybrid neural radiation field model, rendering the first image sample set using light sampling data samples from the first image sample set to obtain a first output image, wherein the initial hybrid neural radiation field model is used to train the hybrid neural radiation field model; and using the first output image as the input image of the initial super-resolution model to train the initial super-resolution model to obtain the super-resolution model.

[0090] In this embodiment, an initial hybrid neural radiation field model can be used to render the first image sample set using ray sampling data samples from the first image sample set, obtaining a first output image. This first output image can be used as the input image for an initial super-resolution model to train the model, thus obtaining the super-resolution model. The ray sampling data samples can be the corresponding ray sampling samples from the image samples in the first image sample set. The first output image can include a high-resolution image rendered by the initial hybrid neural radiation field model, or it can be an image chunk obtained by slicing the high-resolution image.

[0091] Optionally, the first image sample is rendered using the light sampling data samples in the first image sample set using the initial hybrid neural radiation field model to obtain the first output image. The first output image can be used directly as the input of the initial super-resolution model, or the first output image can be input into the initial super-resolution model to train the initial super-resolution model to obtain the super-resolution model.

[0092] In this embodiment, the first output image is typically segmented, and the resulting image segments are then sequentially input into the super-resolution model to complete the training of the initial super-resolution model. Processing the image segments helps save memory consumption without compromising the randomness of the model training process. Unlike related technologies, this embodiment does not simply use random pixels from ordinary training as training samples; instead, it sequentially inputs the resulting image segments into the super-resolution model to complete the training of the initial super-resolution model, thereby improving the efficiency and accuracy of the super-resolution model training process.

[0093] As can be seen from the above, in the embodiments of this application, when the initial super-resolution model is trained alone to obtain the super-resolution model, the data used is independent of the scene of the neural radiation field. That is, the training of the super-resolution model is a one-time process. After training, it can be jointly fine-tuned with any hybrid neural radiation field model. Therefore, the output of the model before the super-resolution model (e.g., DVGO) can be directly used as the input of the super-resolution model, thereby reducing the training cost of the super-resolution model.

[0094] As an optional implementation, the first output image is used as the input image of the initial super-resolution model, and the initial super-resolution model is trained to obtain the super-resolution model, including: obtaining the second output image obtained by the initial super-resolution model through magnification of the input image; establishing a loss function based on the second output image; and adjusting the parameters of the initial super-resolution model based on the loss function to obtain the super-resolution model.

[0095] In this embodiment, the first output image is used as the input image of the initial super-resolution model. The initial super-resolution model magnifies the input image to obtain the second output image. A loss function can be established based on the second output image, and the parameters of the initial super-resolution model can be adjusted based on the loss function to obtain the super-resolution model. The second output image can be the image obtained by the initial super-resolution model after magnifying the input image by a target factor. The loss function can serve as a constraint on the super-resolution model and can be used to accelerate model convergence. The loss function can be a perceptual loss function (PCPLoss), a GAN loss function, etc., and can be one or a mixture of multiple loss functions. This is only an example, and no specific restrictions are placed on the number and type of loss functions.

[0096] Optionally, the output image of the initial mixed neural radiation field can be acquired, and the output image can be segmented to obtain image blocks. The image blocks are used as input to the initial super-resolution model and input sequentially into the initial super-resolution model. The initial super-resolution model magnifies the input image to obtain a second output image. The second output image can be a random image block magnified by a target factor. A loss function can be established based on the second output image, and the parameters of the initial super-resolution model can be adjusted based on the loss function to obtain the super-resolution model.

[0097] As an optional implementation, the loss function includes at least one of the following: a first loss function, used to make the difference between the target type pixels in the second output image and the target type pixels in the corresponding real image less than a pixel threshold; a second loss function, used to make the difference between the pixel distribution information of the second output image and the pixel distribution information of the corresponding real image less than a distribution information threshold; and a third loss function, used to make the difference between the features of the second output image in the feature space and the features of the corresponding real image in the feature space less than a feature threshold.

[0098] In this embodiment, the loss function may include a first loss function, a second loss function, and a third loss function. The first loss function may be an L1 loss function, which can be used to promote the consistency of basic pixel features. The second loss function may be an adversarial network loss function, which can be used to promote the pixel distribution between the rendered image (second output image) and the real image to be similar, thereby further ensuring the consistency of rendering from multiple perspectives and achieving the highest resolution image rendering effect perceived by the human eye. The third loss function may be a perception-based loss function, which can be used to promote the consistency of the image in the feature space.

[0099] Optionally, the three functions mentioned above—perception-based loss function, adversarial network loss function, and L1 loss function—can be used as constraints on the super-resolution model, thereby accelerating the convergence of the super-resolution model during training and improving the model's accuracy.

[0100] In related technologies, peak signal-to-noise ratio (PSNR) oriented loss functions, namely L1 and L2 loss functions, are used. These pixel-matching loss functions easily lead to over-smoothing of image details, a phenomenon particularly noticeable at high resolutions. Therefore, these methods suffer from a lack of visual clarity at high resolutions. In this embodiment, when training the super-resolution model with low-resolution images, an L1 loss function is used to ensure rapid training convergence. Furthermore, during subsequent joint training with high-resolution images, a combination of L1 loss function, a perception-based loss function, and an adversarial network loss function is employed. The addition of a perception-oriented loss function not only improves image clarity but also enhances the consistency of images across different viewpoints.

[0101] As an alternative implementation, the parameters of the initial hybrid neural radiation field model are adjusted based on the gradient of the trained super-resolution model to obtain the hybrid neural radiation field model.

[0102] In this embodiment, the gradient of the trained super-resolution model can be passed to the initial hybrid neural radiation field model, thereby adjusting the parameters of the initial hybrid neural radiation field model based on the gradient of the trained super-resolution model to obtain the hybrid neural radiation field model.

[0103] In this embodiment, the initial hybrid neural radiation field model and the super-resolution model are jointly trained. There is gradient transfer between the two models. When the parameters of the super-resolution model are optimized, it will also affect the optimization of the initial hybrid neural radiation field model. The gradient generated by the loss function of the super-resolution model will also be transferred to the neural radiation field model, thereby enabling the neural radiation field model to obtain domain information that cannot be obtained by ordinary training methods.

[0104] In related technologies, the output of the first model is simply saved as an image, post-processed, and then used as the input to the second model. There is no gradient transfer between the two models (NO Grad), so the parameter optimization of the second model has no impact on the first model. However, in this embodiment, the output of the hybrid neural radiation field can be directly used as the input to the super-resolution model, and gradient transfer exists between the two models. Thus, when the parameters of the super-resolution model are optimized, they can directly affect the optimization of the parameters of the hybrid neural radiation field model. In other words, this embodiment achieves the effect of reducing the computational burden of high-resolution images while improving their quality by jointly fine-tuning the hybrid neural radiation field model and the super-resolution model.

[0105] As an optional implementation, the light sampling data samples from the same training batch input to the initial hybrid neural radiation field model come from the same image block in the first image sample set, while the light sampling data samples from different training batches input to the initial hybrid neural radiation field model come from random image blocks in the first image sample set.

[0106] In this embodiment, light acquisition data samples can be obtained from uniform image blocks in the first image sample set, and light acquisition data samples from different training batches can be obtained from random image blocks in the first image sample set. Light acquisition data samples from the same training batch and light acquisition data samples from different training batches can be used as inputs into the initial hybrid neural radiation field model.

[0107] Optionally, the hybrid neural radiation field improves the low-resolution image, while the final high-resolution image is achieved by the super-resolution model. Since the hybrid neural radiation field generates pixels based on randomly sampled rays, and the super-resolution model obtains new blocks based on image slicing, the random ray sampling method used during the joint training of the neural radiation field and the super-resolution model is random slicing. That is, the rays in each batch during training come from one slice, while the rays between different batches are completely random. Because 3D network optimization requires rendering equations, which result in only being able to calculate the value of a single pixel at a time, it's impossible to consider the values ​​of neighboring pixels. 2D networks, on the other hand, do not require rendering equations and can directly calculate the value of an image region containing multiple pixels. Therefore, they can simultaneously calculate the value of a pixel and its neighboring pixels, thus better utilizing the texture information in the image. Based on the above reasons, this application proposes that during training, the light rays in each batch come from a single slice, while the light rays between different batches are completely random. This solves the problem of inconsistent input between the 3D and 2D networks, utilizes neighborhood information in the 2D network that the 3D network does not utilize, and optimizes the 3D network through gradient backpropagation. Thus, it ensures the consistency of high-resolution rendered images from different perspectives and improves the rendering quality.

[0108] As an alternative implementation, the initial hybrid neural radiation field model is trained on a third set of image samples with a resolution lower than a first resolution threshold.

[0109] In this embodiment, the initial hybrid neural radiation field model can be trained based on a third image sample set with a resolution lower than a first resolution threshold. The third image sample set may contain low-resolution images with a resolution lower than the first resolution threshold.

[0110] In this embodiment, to address the issues of high time consumption and discontinuous viewpoints caused by separately optimizing the neural radiation field and the super-resolution model, patch-based dray sampling is used to combine image segments with rays, achieving joint optimization of the two modules. Furthermore, a perception-guided loss function and an adversarial network loss function are incorporated to ensure consistency in rendering multiple viewpoints, achieving a rendering effect that conforms to human eye perception standards, and realizing a high-resolution image that conforms to current standards as perceived by the human eye.

[0111] In this embodiment, by jointly training a hybrid neural radiation field model and a super-resolution model, the optimization of model parameters influences each other. The hybrid neural radiation field model and super-resolution model generated after joint training process the original image, thereby achieving the technical effect of improving the rendering efficiency of the image and solving the technical problem of low image rendering efficiency.

[0112] The image rendering method of the object in the embodiments of this application will be further described below from the perspective of model training.

[0113] Figure 4 This is a flowchart of another image rendering method for an object according to an embodiment of this application. For example... Figure 4 As shown, the method may include the following steps:

[0114] Step S402: Obtain a first image sample set of the object to be monitored captured from different perspectives, wherein the resolution of the image samples in the first image sample set is higher than the resolution threshold.

[0115] In the technical solution provided by step S402 of this application, a first image sample set of the object to be monitored can be captured from different viewpoints to obtain the first image sample set. The first image sample set may include high-resolution images with a resolution higher than a resolution threshold. The resolution threshold may be a predetermined value or a value stipulated by laws and regulations, such as 4K.

[0116] Step S404: The initial hybrid neural radiation field model and the initial super-resolution model are jointly trained using the first image sample set to obtain the hybrid neural radiation field model and the super-resolution model. The hybrid neural radiation field model is used to store the target image information of the first image sample set in a voxel grid. The frequency of the target image information is less than the frequency threshold. It is also used to render the original image of the object to be monitored under the current viewpoint to obtain the first rendered image under the current viewpoint. The resolution of the original image is lower than another resolution threshold, and the other resolution threshold is lower than the resolution threshold. The super-resolution model is used to enlarge the resolution of the first rendered image to obtain the second rendered image under the current viewpoint.

[0117] In the technical solution provided by step S404 of this application, the initial hybrid neural radiation field and the initial super-resolution model can be jointly trained using a first image sample set to obtain a hybrid neural radiation field model and a super-resolution model. The hybrid neural radiation field model can be used to render the original image of the object to be monitored from the current viewpoint, obtaining a first rendered image monitored from the current viewpoint. The original image can be a low-resolution image with a resolution lower than another resolution threshold, for example, a low-resolution image with a resolution of 2K. The other resolution threshold is less than a resolution threshold, and the other resolution threshold can be a predetermined value or a value stipulated by laws and regulations, for example, 2K. The super-resolution model can be used to enlarge the resolution of the first rendered image to obtain a second rendered image monitored from the current viewpoint.

[0118] Optionally, the hybrid neural radiation field improves the low-resolution image, while the final high-resolution image is achieved by the super-resolution model. Therefore, in this embodiment, the hybrid neural radiation field model processes the first image sample set to obtain a low-resolution image below a first resolution threshold. The super-resolution model processes the low-resolution first rendered image to obtain a high-resolution second rendered image.

[0119] In this embodiment, the initial hybrid neural radiation field and the initial super-resolution model can be jointly trained using a first image sample set. This involves comparing the images in the high-resolution first image sample set with the output image (the second rendered image) of the initial super-resolution model, identifying differences, substituting them into the loss function to calculate the function value, and adjusting the parameters of the initial super-resolution model based on this function value to obtain the super-resolution model. Backpropagation can be used to calculate and propagate gradients, thereby optimizing the parameters of both the initial neural radiation field model and the super-resolution model, causing the model parameters to change towards values ​​more suitable for the output close to the first image sample set. Due to the joint training, the gradients simultaneously affect the parameters of both models during propagation, ensuring that the parameters of both models are optimized concurrently, achieving better results than separate optimization.

[0120] Through steps S402 to S404 of this application, a first image sample set of the object to be monitored captured from different viewpoints is obtained, wherein the resolution of the image samples in the first image sample set is higher than a resolution threshold; the first image sample set is used to jointly train the initial hybrid neural radiation field model and the initial super-resolution model to obtain the hybrid neural radiation field model and the super-resolution model, wherein the hybrid neural radiation field model is used to store at least the target image information of the first image sample set in a voxel grid, the frequency of the target image information is less than a frequency threshold, and is used to render the original image of the object to be monitored at the current viewpoint to obtain the first rendered image monitored at the current viewpoint, the resolution of the original image is lower than another resolution threshold, the other resolution threshold is lower than the resolution threshold, and the super-resolution model is used to enlarge the resolution of the first rendered image to obtain the second rendered image monitored at the current viewpoint, thereby achieving the technical effect of improving the rendering efficiency of the image, and thus solving the technical problem of low rendering efficiency of the image.

[0121] According to an embodiment of this application, an image rendering method for objects in virtual reality scenarios such as virtual reality (VR) devices and augmented reality (AR) devices is also provided. It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. Furthermore, although a logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order than that shown here.

[0122] Figure 5 This is a flowchart of another image rendering method for an object according to an embodiment of this application. For example... Figure 5 As shown, the method may include the following steps:

[0123] Step S502: Start image monitoring of the object to be monitored, and display the original image monitored from the current perspective on the display screen of the virtual reality (VR) device or augmented reality (AR) device, wherein the resolution of the original image is lower than the first resolution threshold.

[0124] Step S504: Invoke the hybrid neural radiation field model and super-resolution model generated after joint training. The samples used during joint training are the first image sample set captured from different viewpoints. The resolution of the image samples in the first image sample set is higher than the second resolution threshold, the second resolution threshold is higher than the first resolution threshold, and the hybrid neural radiation field model is used to store the target image information of the first image sample set in a voxel grid. The frequency of the target image information is lower than the frequency threshold.

[0125] Step S506: Control the VR device or AR device to render the original image using a hybrid neural radiation field model to obtain a first rendered image monitored from the current viewpoint, wherein the resolution of the first rendered image is less than a first resolution threshold.

[0126] Step S508: Control the VR device or AR device to use a super-resolution model to magnify the resolution of the first rendered image to obtain a second rendered image monitored from the current viewpoint, wherein the resolution of the second rendered image is higher than the second resolution threshold.

[0127] Step S510: Drive the VR device or AR device to render and display the second rendered image.

[0128] Optionally, in this embodiment, the image rendering method described above can be applied to a hardware environment consisting of a server and a virtual reality device. The original image monitored from the current viewpoint is displayed on the screen of the virtual reality (VR) device or augmented reality (AR) device. The server can be a server corresponding to a media file operator. The aforementioned network includes, but is not limited to, a wide area network (WAN), a metropolitan area network (MAN), or a local area network (LAN). The aforementioned virtual reality device is not limited to, for example, a virtual reality headset, virtual reality glasses, or a standalone virtual reality device.

[0129] Optionally, the virtual reality device includes: a memory, a processor, and a transmission device. The memory stores an application that can be used to execute: initiate image monitoring of the object to be monitored; display the original image monitored from the current viewpoint on the display screen of a virtual reality (VR) device or augmented reality (AR) device, wherein the resolution of the original image is lower than a first resolution threshold; invoke a hybrid neural radiation field model and a super-resolution model generated after joint training, wherein the samples used during joint training are a first image sample set captured from different viewpoints, the resolution of the image samples in the first image sample set is higher than a second resolution threshold, the second resolution threshold is greater than the first resolution threshold, and the hybrid neural radiation field model is used to store at least the target image information of the first image sample set in a voxel grid, the frequency of the target image information is lower than a frequency threshold; control the VR device or AR device to render the original image using the hybrid neural radiation field model to obtain a first rendered image monitored from the current viewpoint, wherein the resolution of the first rendered image is lower than the first resolution threshold; control the VR device or AR device to use the super-resolution model to enlarge the resolution of the first rendered image to obtain a second rendered image monitored from the current viewpoint, wherein the resolution of the second rendered image is higher than the second resolution threshold; drive the VR device or AR device to render and display the second rendered image.

[0130] It should be noted that the image rendering method for objects applied in VR or AR devices described above in this embodiment may include... Figure 3The method of the illustrated embodiment is used to drive a VR device or AR device to display a second rendered image.

[0131] Optionally, the processor in this embodiment can invoke the application stored in the memory via the transmission device to perform the above steps. The transmission device can receive media files sent by the server via a network, and can also be used for data transmission between the processor and the memory.

[0132] Optionally, in a virtual reality device, there is a head-mounted display with eye tracking. The screen in the HMD is used to display the video footage. The eye tracking module in the HMD is used to acquire the real-time movement path of the user's eyes. The tracking system is used to track the user's position and movement information in real three-dimensional space. The computing and processing unit is used to acquire the user's real-time position and movement information from the tracking system and calculate the three-dimensional coordinates of the user's head in the virtual three-dimensional space, as well as the user's field of vision orientation in the virtual three-dimensional space.

[0133] In this embodiment, the virtual reality device can be connected to a terminal, and the terminal can be connected to a server via a network. The virtual reality device is not limited to virtual reality headsets, virtual reality glasses, virtual reality all-in-one machines, etc., and the terminal is not limited to PCs, mobile phones, tablets, etc. The server can be a server corresponding to a media file operator, and the network includes, but is not limited to, wide area networks, metropolitan area networks, or local area networks.

[0134] Figure 6 This is a schematic diagram of a rendering result according to an embodiment of this application, such as... Figure 6 As shown, a second rendered image is displayed on the screen of a virtual reality (VR) device or an augmented reality (AR) device, and the original image detected by the VR or AR device at the current viewpoint is retrieved (e.g., ...). Figure 6 The original image is used to call the hybrid neural radiation field model and super-resolution model generated after joint training; the VR or AR device is controlled to render the original image using the hybrid neural radiation field model to obtain the first rendered image monitored from the current viewpoint; the VR or AR device is controlled to use the super-resolution model to enlarge the resolution of the first rendered image to obtain the second rendered image monitored from the current viewpoint, such as... Figure 6 The method shows how to acquire an original image containing a rabbit as the object to be monitored, process the original image to obtain a high-resolution second rendered image, and drive VR or AR devices to render and display the second rendered image.

[0135] Under the aforementioned operating environment, one embodiment of this application also provides another such... Figure 7 The image rendering method for the object shown is illustrated. It should be noted that the image rendering method for the object in this embodiment can be provided by... Figure 1 The mobile terminal in the illustrated embodiment is executed. Figure 7 This is a flowchart of another image rendering method for an object according to an embodiment of this application.

[0136] like Figure 7 As shown, the method may include the following steps:

[0137] Step S702: Obtain the original image monitored from the current viewpoint by calling the first interface, wherein the first interface includes a first parameter, the parameter value of the first parameter is the original image, and the resolution of the original image is lower than a first resolution threshold.

[0138] In the technical solution provided by step S702 of this application, the first interface can be an interface for data interaction between the server and the client. The client can use the original image monitored under the current viewpoint to be acquired as a first parameter of the first interface to achieve the purpose of acquiring the original image monitored under the current viewpoint using the first interface.

[0139] Step S704: Invoke the hybrid neural radiation field model and super-resolution model generated after joint training. The samples used during joint training are the first image sample set captured from different viewpoints. The resolution of the image samples in the first image sample set is higher than the second resolution threshold, the second resolution threshold is greater than the first resolution threshold, and the hybrid neural radiation field model is used to store the target image information of the first image sample set in a voxel grid. The frequency of the target image information is less than the frequency threshold.

[0140] Step S706: Render the original image using a hybrid neural radiation field model to obtain a first rendered image monitored from the current viewpoint, wherein the resolution of the first rendered image is less than a first resolution threshold.

[0141] Step S708: Use a super-resolution model to enlarge the resolution of the first rendered image to obtain a second rendered image monitored from the current viewpoint, wherein the resolution of the second rendered image is higher than the second resolution threshold.

[0142] Step S710: Output the second rendered image by calling the second interface, wherein the second interface includes a second parameter, and the value of the second parameter is the second rendered image.

[0143] In the technical solution provided by step S710 of this application, the second interface can be an interface for data interaction between the server and the client. The server can pass the second rendered image into the second interface as a parameter of the second interface to achieve the purpose of outputting the second rendered image.

[0144] Through the above steps, the original image monitored from the current viewpoint is obtained by calling the first interface, where the first interface includes a first parameter whose value is the original image, and the resolution of the original image is lower than a first resolution threshold. The hybrid neural radiation field model and super-resolution model generated after joint training are then called. During joint training, the samples used are a set of first image samples captured from different viewpoints. The resolution of the image samples in the first image sample set is higher than a second resolution threshold, and the second resolution threshold is greater than the first resolution threshold. The hybrid neural radiation field model is used to store at least the target image information of the first image sample set in a voxel grid, and the frequency of the target image information is lower than a frequency threshold. The hybrid neural radiation field model is used to render the original image to obtain a first rendered image monitored from the current viewpoint, where the resolution of the first rendered image is lower than the first resolution threshold. The super-resolution model is used to amplify the resolution of the first rendered image to obtain a second rendered image monitored from the current viewpoint, where the resolution of the second rendered image is higher than the second resolution threshold. The second rendered image is output by calling the second interface, where the second interface includes a second parameter whose value is the second rendered image. This achieves the technical effect of improving image rendering efficiency and solves the technical problem of low image rendering efficiency.

[0145] Example 2

[0146] As algorithms and hardware related to 3D scenes mature, concepts such as VR, AR, and metaverse have gained popularity. Among the 3D scene-related algorithms developed in recent years, neural radiation field algorithms, which only require multiple photos taken from different angles with a regular mobile phone for the same object, are relatively convenient to use and produce high-quality rendering effects. Therefore, neural radiation field algorithms have become one of the fastest-growing 3D algorithms.

[0147] However, the original neural radiation field has many problems that prevent its practical application. For example, the algorithm's training time is measured in days, resulting in a long training period; after training, the method still has a long rendering time for images; and the rendered images are prone to problems such as blurred details and foggy noise. These problems are particularly noticeable when rendering high-resolution images.

[0148] Currently, several improved neural radiation field algorithms have emerged to address the time-consuming training and rendering issues associated with speed and cost, achieving significant progress. Examples include Plenoxels (3D scene reconstruction) and InstantNGP (3D reconstruction algorithm). Other algorithms improve image rendering quality, such as Mip-Ne RF (Multi-Scale Neural Radiation Field) and NeRF-SR (New View Synthesis), but all these algorithms are primarily designed for low-resolution (less than 1K) scenes. Once extended to high-resolution scenes, the inefficient rendering of objects remains a significant challenge.

[0149] As mentioned above, algorithms addressing speed-related costs often incur high storage costs, which are particularly pronounced in high-resolution scene representations. For example, 3D scene reconstruction algorithms, while offering fast training speeds, require large storage spaces, limiting their practicality. Similarly, algorithms for improving image rendering quality also suffer from high speed costs in high-resolution scenarios. For instance, the original neural radiation field algorithm, while requiring less storage, has long training and inference times, hindering large-scale deployment. Furthermore, current algorithms are all pixel-oriented, focusing on restoring the peak signal-to-noise ratio (PSNR). This pixel-oriented approach easily leads to image blurring in high-resolution scenes, and no effective solutions have yet been proposed to address these issues.

[0150] In related technologies, a super-resolution algorithm for neural radiation fields is proposed. Figure 8 This is a schematic diagram of a super-resolution algorithm architecture for neural radiation fields based on related technologies. For example... Figure 8 As shown, the super-resolution algorithm for neural radiation fields employs the most basic pure network structure of neural radiation fields and uses a phased training approach. In the first phase, a rendered image is obtained through sampling rays, multilayer perceptrons, and voxel rendering. The neural radiation field is then trained using low-resolution true images (LR True Image) with an L2 loss function. In the second phase, the trained neural radiation field is used to render high-resolution true images (HR GT image patches) and simultaneously render depth maps (Rendered image patches). Similar patches for the viewpoint to be rendered are then found in the high-resolution true images. This warps the depth maps, aligning the high-resolution true image patches with the rendered image patches. These two sets of aligned patches are then used to train an encoder-decoder network to fine-tune the high-resolution rendering results of the neural radiation field.

[0151] While the neural radiation field super-resolution algorithm employs subpixel sampling to improve the final renderable resolution, this method significantly increases training time and still cannot solve the rendering quality problem for high-resolution images. Furthermore, this method uses an encoder-decoder structure for secondary reconstruction. By using regions similar to the rendering viewpoint on the high-resolution image as common input to train the encoder-decoder network from the subpixel sampled neural radiation field-rendered image, this method, while utilizing information from the known high-resolution image, further increases training and rendering time. Moreover, since it is performed separately from the neural radiation field algorithm training, it cannot guarantee consistent detail rendering across different viewpoints. This results in the algorithm not producing high-resolution results, and even when rendering images at high resolution (e.g., 4K), significant noise remains in the rendered results, and the texture of the same object is inconsistent across different viewpoints. Because this method uses the original neural radiation field structure and an additional encoder-decoder, it still suffers from very long training and testing times.

[0152] To address the aforementioned issues, this application proposes a Direct Voxel Super Resolution (DVSR) algorithm to resolve the problems of neural radiation fields in high-resolution scenes. It can achieve visual effects at current standard resolutions (e.g., 4K) with relatively low time and storage costs.

[0153] It should be noted that videos with resolutions of 4K and above are defined as ultra-high-definition videos, representing the main development direction of the video industry in the future. Most currently popular videos have a resolution of 2K, and neural radiation field algorithms in related technologies struggle to render videos above 2K. The algorithm proposed in this application can produce ultra-high-definition visual effects of 4K and above, which is currently unattainable by related technologies.

[0154] The super-resolution direct pixel algorithm for high-resolution neural radiation fields proposed in the embodiments of this application will be further explained below.

[0155] Figure 9 This is a schematic diagram of a super-resolution direct pixel algorithm architecture according to an embodiment of this application. Figure 9 As shown, this embodiment trains a hybrid neural radiation field model (Hybrid-nerf) and a super-resolution model using low-resolution images. Then, these two trained models are combined and fine-tuned using high-resolution image slices. The hybrid neural radiation field model can be optimized using a direct voxel grid.

[0156] In related technologies, the output of the first model is simply saved as an image, post-processed, and then used as the input to the second model; there is no gradient transfer between the two models (NO Grad). Therefore, the parameter optimization of the second model has no impact on the first model. However, in the embodiment of this application, the output of the previous model (DVGO) can be directly used as the input to the next model (SR), and gradient transfer exists between the two models. Thus, when the parameters of the second model are optimized, they can directly affect the optimization of the parameters of the first model. That is, the embodiment of this application achieves the effect of reducing the computational load of high-resolution images while improving their quality by jointly fine-tuning the neural radiation field model and the super-resolution model.

[0157] In this embodiment, training with high-resolution image slices allows for comparison between the high-resolution slice image and the output image of the super-resolution model. Differences are identified, and these differences are substituted into the loss function to calculate the function value. Then, backpropagation is used to calculate and propagate the gradient, thereby optimizing the model parameters to better align with the output slice values. Based on this joint fine-tuning, the gradient propagation simultaneously affects the parameters of both models, ensuring that the parameters of both models are optimized concurrently, achieving better results than separate optimization.

[0158] As an alternative embodiment, a hybrid neural radiation field model and a super-resolution model can be trained using low-resolution (1K) image samples, wherein the hybrid neural radiation field model and the super-resolution model can employ the same loss function (e.g., L2 loss).

[0159] like Figure 9 As shown, the embodiments of this application use a hybrid neural radiation field structure, which can store coarse low-frequency information of the scene in a small voxel grid, and represent higher-frequency detailed information using a multilayer perceptron.

[0160] Alternatively, by using a voxel grid, on the one hand, it is possible to directly determine which locations in space contain object information, thus avoiding invalid sampling; on the other hand, low-frequency information of objects can be stored in the corresponding spatial locations, reducing computation in the data processing process.

[0161] Meanwhile, since low-frequency information is already stored in the voxel grid, the multilayer perceptron only needs to supplement high-frequency information. For natural images, the amount of low-frequency information is far greater than the amount of high-frequency information. Therefore, the multilayer perceptron only needs a small network to learn how to supplement high-frequency information. Thus, the multilayer perceptron after voxel gridding only needs a small network to further supplement high-frequency information based on the information in the voxel grid, thereby improving computational speed and reducing rendering costs. High-frequency information can be used to characterize rapidly changing parts of the image, such as edges, noise, or details.

[0162] It should be noted that low-frequency and high-frequency information in an image are relative concepts. For example, if the intensity changes drastically at various locations in an image (e.g., the intensity difference between locations is greater than 10), then the image contains not only low-frequency information but also high-frequency information. From the image's spectrum, there is not only one main peak but also multiple side peaks. Therefore, it can be determined that the image information corresponding to the multiple side peaks can be low-frequency information, and the information corresponding to the main peak can be high-frequency information.

[0163] For example, a very small network can be a network consisting of a 3-layer multilayer perceptron and two linear rectified activation functions without any parameters. The number of multilayer perceptrons in the first two layers can be 64*64, and the number of multilayer perceptrons in the last layer can be 64*3. It should be noted that the above parameters are only illustrative and are not specific limitations.

[0164] In related technologies, a pure network structure of the original neural radiation field is used. The advantage of this implicit neural radiation field is that it occupies a small storage space, generally only about 5M. However, this method has a large computational load. For each input spatial sampling point, it is necessary to perform calculations through a multi-layer neural network, and this multi-layer network is generally more than 8 layers, with more than 128 neurons in each layer. As a result, for high-resolution target images with many pixels, the training time of this method is too long, and the rendering time of an image is also long, resulting in a high overall time cost. In contrast, the embodiment of this application uses a hybrid neural radiation field structure to represent higher frequency detail information using a multilayer perceptron. It utilizes the explicit neural radiation field to directly store the coarse low-frequency information of the scene in a small voxel grid; and utilizes the multilayer perceptron set in the implicit neural radiation field to store higher frequency detail information with a very small storage space. In other words, the embodiments of this application combine the advantages of explicit neural radiation fields and implicit neural radiation fields, avoiding both the computational time problem of implicit neural radiation fields alone and the high storage occupation problem of explicit neural radiation fields, thereby improving training speed and inference speed, and having extremely high flexibility and easy expansion.

[0165] For example, a hybrid neural radiation field model can be the DVGO model, which can be trained and converged on a 1K resolution image in 10 minutes. It should be noted that the hybrid neural radiation field model can also be other models, or it can be a combination of the DVGO model and other hybrid neural radiation field models. This is just an example and no specific restrictions are placed on the hybrid neural radiation field model.

[0166] For another example, the super-resolution model can be a simplified version of the super-resolution algorithm, with parameters that are about 1 / 5 of the original real-ESRGAN. During training, the super-resolution model can be trained using a hybrid simulated degradation method based on real-scene super-resolution, thus achieving a 4x upscaling. This hybrid simulated degradation method can involve common artificial image degradation processes, such as adding high-frequency noise, blurring, adding compressed noise, downsampling, etc. These degradation methods can be combined in a random order and number to achieve a hybrid simulated degradation of the real-scene image.

[0167] like Figure 9 As shown, the training process of the super-resolution model can include: using some high-resolution images as training images, and then performing the above-mentioned hybrid downsampling on the training images. The downsampling can be 4x, and other downsampling methods can be randomly combined. This results in low-resolution images that have been reduced by 4x and have had varying degrees of noise and blur added. These high-resolution and low-resolution images can be combined to train the super-resolution model.

[0168] It should be noted that the data used when training this super-resolution model is independent of the neural radiation field scene. Therefore, the training of the super-resolution model is a one-time process. After training, it can be jointly fine-tuned with any hybrid neural radiation field model. In other words, the output of the previous model (e.g., DVGO) can be directly used as the input of the super-resolution model, thereby reducing the training cost of the super-resolution model.

[0169] As an alternative embodiment, high-resolution images can be used to jointly fine-tune the trained hybrid neural radiation field and super-resolution model.

[0170] like Figure 9 As shown, during joint fine-tuning, the light rays generated by image chunks based on random images can be input into the direct voxel mesh optimization model. The direct voxel mesh optimization model performs voxel rendering and other processing on the data to render the image chunks with texture information.

[0171] In this embodiment of the application, the method for rendering the image is not to enlarge the rendered image by ordinary random pixels, but to enlarge the rendered image based on the second voxel features corresponding to the second ray sampling data of the image block. This image block-based approach is beneficial for the subsequent super-resolution model to use directly, and it does not impair the randomness of training when the image block size is small.

[0172] Optionally, the direct voxel mesh optimization model can input the rendered image chunks with texture information into the super-resolution model, which then enlarges the chunks to obtain image chunks of a random image magnified by 4 times (4x) (the second rendered image).

[0173] Optionally, the super-resolution model can be constrained by three functions: a perceptual loss function, an adversarial network loss function, and an L1 loss function. The L1 loss function can promote basic pixel consistency. The perceptual loss function can promote consistency in the image's feature space. The adversarial network loss function can promote pixel distribution similarity between the rendered image and the real image.

[0174] In related technologies, a peak signal-to-noise ratio-oriented loss function is used, that is, Figure 8 The L1 and L2 loss functions mentioned above, which are pixel-matching loss functions, can easily lead to over-smoothing of image details, especially at high resolutions. Therefore, these methods often fail to achieve the desired clarity at high resolutions. In this embodiment, when training the super-resolution model with low-resolution images, an L1 loss function is used to ensure rapid convergence. Furthermore, during subsequent joint training with high-resolution images, a combination of the L1 loss function, a perception-guided loss function, and an adversarial network loss function is employed. The addition of a perception-guided loss function not only improves image clarity but also enhances the consistency of images across different viewpoints.

[0175] In this embodiment, since the direct voxel grid optimization model and the super-resolution model are jointly trained, the gradient generated by the loss function of the super-resolution model is also passed to the direct voxel grid optimization model, thereby enabling the neural radiation field model to acquire domain information that cannot be obtained by ordinary training methods.

[0176] Optionally, the hybrid neural radiation field improves the low-resolution image, while the final high-resolution image is achieved by the super-resolution model. Since the hybrid neural radiation field generates pixels based on randomly sampled rays, and the super-resolution model obtains new image patches based on image slicing, the random ray sampling method used during the joint training of the neural radiation field and the super-resolution model is random slicing. That is, the rays in each batch during training come from one patch, while the rays between different batches are completely random. 3D network optimization requires rendering equations, which result in only being able to calculate the value of a single pixel at a time, without considering its neighbors. 2D networks, on the other hand, do not require rendering equations and can directly calculate the image of a region containing multiple pixels. Therefore, they can simultaneously calculate the value of a pixel and its neighbors, thus better utilizing the texture information in the image. This approach solves the problem of inconsistent inputs between 3D and 2D networks, utilizes neighborhood information not used by 3D networks in 2D networks, and optimizes the 3D network through gradient backpropagation. This ensures consistency of the high-resolution rendered image across different viewpoints while improving rendering quality.

[0177] As an alternative implementation, high-resolution (e.g., 4K resolution) images can be directly generated based on a jointly trained direct voxel grid optimization model and a super-resolution model.

[0178] For example, such as Figure 9 As shown, random rays from a 1K resolution image can be sampled into the direct voxel mesh optimization model. After voxel rendering, a low-resolution image with a resolution of 1K is obtained. Then, after being magnified by 4 times by the super-resolution model, a high-resolution image with a resolution of 4K (HR Image) is obtained.

[0179] In this embodiment, by using a hybrid neural radiation field model and a super-resolution model for joint training, the problem of poor image quality when rendering high-resolution images in related technologies is solved, achieving the effect of reducing the computational cost of high-resolution images while improving image quality. Furthermore, since the super-resolution model is based on convolution, it can effectively utilize the neighborhood information of the image. Therefore, this method improves image quality on the one hand, and avoids the high time consumption of directly rendering high-resolution images using the neural radiation field model on the other.

[0180] In this embodiment of the application, in order to solve the problems of high time consumption and discontinuous viewpoint caused by optimizing the neural radiation field and the super-resolution model separately, the image blocks and light rays are combined by using image patch-based ray sampling to achieve the purpose of jointly optimizing the two modules. Furthermore, a perception-guided loss function and an adversarial network loss function are added to ensure the consistency of rendering multiple viewpoints, achieving a rendering effect that meets the standards for human eye perception, and realizing the current high-resolution image that meets the standards for human eye perception.

[0181] In another alternative embodiment, Figure 10 The use of the above is illustrated in a block diagram. Figure 1 The computer terminal (or mobile device) shown is an example of a service mesh. Figure 10 This is a service mesh structure diagram of an image rendering method for an object according to an embodiment of this application, such as... Figure 10 As shown, the Service Mesh 1000 is mainly used to facilitate secure and reliable communication between multiple microservices. Microservices refer to the decomposition of an application into multiple smaller services or instances, which are distributed across different clusters / machines to run.

[0182] like Figure 10 As shown, a microservice may include application service instance A and application service instance B, which together form the functional application layer of service mesh 1000. In one implementation, application service instance A runs as a container / process 1008 on machine / workload container group 1014 (POD), and application service instance B runs as a container / process 1010 on machine / workload container group 1015 (POD).

[0183] In one implementation, application service instance A can be a product query service, and application service instance B can be a product order placement service.

[0184] like Figure 10 As shown, application service instance A and grid agent (sidecar) 1003 coexist in machine workload container group 1014, and application service instance B and grid agent 1005 coexist in machine workload container 1014. Grid agents 1003 and 1005 form the data plane layer of service mesh 1000. Grid agents 1003 and 1005 run as container / process 1004, which can receive requests 1012 for product query services, and as grid agent 1006. Grid agent 1003 and application service instance A can communicate bidirectionally, and grid agent 1005 and application service instance B can also communicate bidirectionally. Furthermore, grid agents 1003 and 1005 can also communicate bidirectionally with each other.

[0185] In one implementation, all traffic from application service instance A is routed to the appropriate destination via mesh proxy 1003, and all network traffic from application service instance B is routed to the appropriate destination via mesh proxy 1005. It should be noted that the network traffic mentioned here includes, but is not limited to, Hypertext Transfer Protocol (HTTP), Representational State Transfer (REST), high-performance, general-purpose open-source frameworks (gRPC), and open-source in-memory data structure storage systems (Redis).

[0186] In one implementation, the functionality of the extended data plane layer can be achieved by writing custom filters for the agents (Envoy) in service mesh 1000. The service mesh agent configuration can enable the service mesh to correctly proxy service traffic, achieving service interoperability and service governance. Mesh agents 1003 and 1005 can be configured to perform at least one of the following functions: service discovery, health checking, routing, load balancing, authentication and authorization, and observability.

[0187] like Figure 10 As shown, the service mesh 1000 also includes a control plane layer. This control plane layer can consist of a set of services running in a dedicated namespace, hosted by a managed control plane component 1001 within machine / workload container groups (machine / Pod) 10002. For example... Figure 10As shown, the managed control plane component 1001 communicates bidirectionally with grid agents 1003 and 1005. The managed control plane component 1001 is configured to perform several control and management functions. For example, the managed control plane component 1001 receives telemetry data transmitted by grid agents 1003 and 1005 and can further aggregate this telemetry data. In addition to these services, the managed control plane component 1001 can also provide a user-facing application programming interface (API) to facilitate manipulation of network behavior and provision of configuration data to grid agents 1003 and 1005. It should be noted that, for the foregoing method embodiments, for the sake of simplicity, they are all described as a series of actions. However, those skilled in the art should understand that this application is not limited to the described order of actions, because according to this application, some steps can be performed in other orders or simultaneously. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are preferred embodiments, and the actions and modules involved are not necessarily essential to this application.

[0188] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that this application is not limited to the described order of actions, as some steps may be performed in other orders or simultaneously according to this application. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are preferred embodiments, and the actions and modules involved are not necessarily essential to this application.

[0189] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods according to the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes several instructions to cause a terminal device (which may be a mobile phone, computer, server, or network device, etc.) to execute the methods of the various embodiments of this application.

[0190] Example 3

[0191] According to embodiments of this application, a method for implementing the above is also provided. Figure 3 The image rendering method of the object shown is the image rendering device of the object.

[0192] Figure 11 This is a schematic diagram of an image rendering apparatus for an object according to an embodiment of this application, such as... Figure 11As shown, the image rendering device 1100 of the object may include: a first acquisition unit 1102, a first calling unit 1104, a first rendering unit 1106, and a first processing unit 1108.

[0193] The first acquisition unit 1102 is used to initiate image monitoring of the object to be monitored and acquire the original image monitored from the current viewpoint, wherein the resolution of the original image is lower than a first resolution threshold.

[0194] The first calling unit 1104 is used to call the hybrid neural radiation field model and the super-resolution model generated after joint training. The samples used during joint training are a first image sample set captured from different viewpoints. The resolution of the image samples in the first image sample set is higher than a second resolution threshold, the second resolution threshold is greater than the first resolution threshold, and the hybrid neural radiation field model is used to store at least the target image information of the first image sample set in a voxel grid. The frequency of the target image information is less than a frequency threshold.

[0195] The first rendering unit 1106 is used to render the original image using a hybrid neural radiation field model to obtain a first rendered image monitored at the current viewpoint, wherein the resolution of the first rendered image is less than a first resolution threshold.

[0196] The first processing unit 1108 is used to enlarge the resolution of the first rendered image using a super-resolution model to obtain a second rendered image monitored from the current viewpoint, wherein the resolution of the second rendered image is higher than a second resolution threshold.

[0197] It should be noted that the first acquisition unit 1102, the first invocation unit 1104, the first rendering unit 1106, and the first processing unit 1108 mentioned above correspond to steps S302 to S308 in Embodiment 1. The four units and the corresponding steps implement the same instances and application scenarios, but are not limited to the content disclosed in Embodiment 1. It should be noted that the above units can be hardware components or software components stored in memory and processed by one or more processors. The above units can also be part of the device and run in the AR / VR device provided in Embodiment 1.

[0198] According to embodiments of this application, a method for implementing the above is also provided. Figure 4 The image rendering method of the object shown is the image rendering device of the object.

[0199] Figure 12 This is a schematic diagram of an image rendering apparatus for another object according to an embodiment of this application, such as... Figure 12 As shown, the image rendering device 1200 for the object may include a second acquisition unit 1202 and a training unit 1204.

[0200] The second acquisition unit 1202 is used to acquire a first image sample set of the object to be monitored captured from different perspectives, wherein the resolution of the image samples in the first image sample set is higher than the resolution threshold.

[0201] Training unit 1204 is used to jointly train the initial hybrid neural radiation field model and the initial super-resolution model using the first image sample set to obtain the hybrid neural radiation field model and the super-resolution model. The hybrid neural radiation field model is used to store the target image information of the first image sample set in a voxel grid. The frequency of the target image information is less than the frequency threshold. It is also used to render the original image of the object to be monitored under the current viewpoint to obtain the first rendered image under the current viewpoint. The resolution of the original image is lower than another resolution threshold. The other resolution threshold is lower than the resolution threshold. The super-resolution model is used to enlarge the resolution of the first rendered image to obtain the second rendered image under the current viewpoint.

[0202] It should be noted that the second acquisition unit 1202 and training unit 1204 mentioned above correspond to steps S402 to S404 in Embodiment 1. The two units and the corresponding steps implement the same instances and application scenarios, but are not limited to the content disclosed in Embodiment 1. It should be noted that the above units can be hardware components or software components stored in memory and processed by one or more processors. The above units can also be part of the device and run in the AR / VR device provided in Embodiment 1.

[0203] According to embodiments of this application, a method for implementing the above is also provided. Figure 5 The image rendering method of the object shown is the image rendering device of the object.

[0204] Figure 13 This is a schematic diagram of an image rendering apparatus for another object according to an embodiment of this application, such as... Figure 13 As shown, the voice generation device 1300 may include: a display unit 1302, a second calling unit 1304, a first control unit 1306, a second control unit 1308, and a driving unit 1310.

[0205] The display unit 1302 is used to initiate image monitoring of the object to be monitored and display the original image monitored from the current perspective on the display screen of the virtual reality (VR) device or augmented reality (AR) device, wherein the resolution of the original image is lower than a first resolution threshold.

[0206] The second calling unit 1304 is used to call the hybrid neural radiation field model and the super-resolution model generated after joint training. The samples used during joint training are a first image sample set captured from different viewpoints. The resolution of the image samples in the first image sample set is higher than a second resolution threshold, the second resolution threshold is greater than the first resolution threshold, and the hybrid neural radiation field model is used to store at least the target image information of the first image sample set in a voxel grid. The frequency of the target image information is less than a frequency threshold.

[0207] The first control unit 1306 is used to control the VR device or AR device to render the original image using a hybrid neural radiation field model to obtain a first rendered image monitored at the current viewpoint, wherein the resolution of the first rendered image is less than a first resolution threshold.

[0208] The second control unit 1308 is used to control the VR device or AR device to use a super-resolution model to amplify the resolution of the first rendered image to obtain a second rendered image monitored from the current viewpoint, wherein the resolution of the second rendered image is higher than a second resolution threshold.

[0209] The driving unit 1310 is used to drive a VR device or AR device to render and display a second rendered image.

[0210] It should be noted that the aforementioned display unit 1302, second calling unit 1304, first control unit 1306, second control unit 1308, and driving unit 1310 correspond to steps S502 to S510 in Embodiment 1. The five units and their corresponding steps implement the same instances and application scenarios, but are not limited to the content disclosed in Embodiment 1. It should also be noted that the aforementioned units can be hardware or software components stored in memory and processed by one or more processors. These units can also function as part of a device within the AR / VR device provided in Embodiment 1.

[0211] According to an embodiment of this application, another method for implementing the above is also provided. Figure 7 The image rendering method of the object shown is the image rendering device of the object.

[0212] Figure 14 This is a schematic diagram of an image rendering apparatus for another object according to an embodiment of this application, such as... Figure 14 As shown, the image rendering device 1400 of the object may include: a third calling unit 1402, a fourth calling unit 1404, a second rendering unit 1406, a second processing unit 1408 and a fifth calling unit 1410.

[0213] The third calling unit 1402 is used to obtain the original image monitored under the current view by calling the first interface, wherein the first interface includes a first parameter, the parameter value of the first parameter is the original image, and the resolution of the original image is lower than a first resolution threshold.

[0214] The fourth calling unit 1404 is used to call the hybrid neural radiation field model and super-resolution model generated after joint training. The samples used during joint training are the first image sample set captured from different viewpoints. The resolution of the image samples in the first image sample set is higher than the second resolution threshold, the second resolution threshold is higher than the first resolution threshold, and the hybrid neural radiation field model is used to store the target image information of the first image sample set in a voxel grid. The frequency of the target image information is lower than the frequency threshold.

[0215] The second rendering unit 1406 is used to render the original image using a hybrid neural radiation field model to obtain a first rendered image monitored at the current viewpoint, wherein the resolution of the first rendered image is less than a first resolution threshold.

[0216] The second processing unit 1408 is used to enlarge the resolution of the first rendered image using a super-resolution model to obtain a second rendered image monitored from the current viewpoint, wherein the resolution of the second rendered image is higher than a second resolution threshold.

[0217] The fifth calling unit 1410 is used to output a second rendered image by calling a second interface, wherein the second interface includes a second parameter, and the value of the second parameter is the second rendered image.

[0218] It should be noted that the third calling unit 1402, the fourth calling unit 1404, the second rendering unit 1406, the second processing unit 1408, and the fifth calling unit 1410 mentioned above correspond to steps S702 to S710 in Embodiment 1. The five units and the corresponding steps implement the same instances and application scenarios, but are not limited to the content disclosed in Embodiment 1. It should be noted that the above units can be hardware components or software components stored in memory and processed by one or more processors. The above units can also be part of the device and run in the AR / VR device provided in Embodiment 1.

[0219] In the image rendering apparatus of this embodiment, the optimization of model parameters is mutually influenced by the joint training of the hybrid neural radiation field model and the super-resolution model. The original image is processed by the hybrid neural radiation field model and the super-resolution model generated after joint training, thereby achieving the technical effect of improving the rendering efficiency of the image and solving the technical problem of low image rendering efficiency.

[0220] Example 4

[0221] Embodiments of this application may provide a computer terminal, which may be any computer terminal device in a group of computer terminals. Optionally, in this embodiment, the aforementioned computer terminal may also be replaced by a mobile terminal or other terminal device.

[0222] Optionally, in this embodiment, the computer terminal may be located in at least one of a plurality of network devices in a computer network.

[0223] In this embodiment, the computer terminal can execute the following steps of the image rendering method for the object of the application: initiate image monitoring of the object to be monitored, acquire the original image monitored from the current viewpoint, wherein the resolution of the original image is lower than a first resolution threshold; call the hybrid neural radiation field model and super-resolution model generated after joint training, wherein the samples used during joint training are a first image sample set captured from different viewpoints, the resolution of the image samples in the first image sample set is higher than a second resolution threshold, the second resolution threshold is greater than the first resolution threshold, and the hybrid neural radiation field model is used to store at least the target image information of the first image sample set in a voxel grid, the frequency of the target image information is lower than a frequency threshold; render the original image using the hybrid neural radiation field model to obtain a first rendered image monitored from the current viewpoint, wherein the resolution of the first rendered image is lower than the first resolution threshold; use the super-resolution model to enlarge the resolution of the first rendered image to obtain a second rendered image monitored from the current viewpoint, wherein the resolution of the second rendered image is higher than the second resolution threshold.

[0224] Optionally, Figure 15 This is a structural block diagram of a computer terminal according to an embodiment of this application. Figure 15 As shown, the computer terminal A may include one or more (only one is shown in the figure) processors 1502, memory 1504, and transmission devices 1506.

[0225] The memory can be used to store software programs and modules, such as the program instructions / modules corresponding to the image rendering method and apparatus for the object in this embodiment. The processor executes various functional applications and predictions by running the software programs and modules stored in the memory, thereby realizing the aforementioned image rendering method for the object. The memory may include high-speed random access memory, and may also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory may further include memory remotely located relative to the processor, and these remote memories can be connected to terminal A via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.

[0226] The processor can invoke information and applications stored in memory via a transmission device to perform the following steps: initiate image monitoring of the object to be monitored, acquire the original image monitored from the current viewpoint, wherein the resolution of the original image is lower than a first resolution threshold; invoke the hybrid neural radiation field model and super-resolution model generated after joint training, wherein the samples used during joint training are a first image sample set captured from different viewpoints, the resolution of the image samples in the first image sample set is higher than a second resolution threshold, the second resolution threshold is greater than the first resolution threshold, and the hybrid neural radiation field model is used to store at least the target image information of the first image sample set in a voxel grid, the frequency of the target image information is lower than a frequency threshold; render the original image using the hybrid neural radiation field model to obtain a first rendered image monitored from the current viewpoint, wherein the resolution of the first rendered image is lower than the first resolution threshold; and enlarge the resolution of the first rendered image using the super-resolution model to obtain a second rendered image monitored from the current viewpoint, wherein the resolution of the second rendered image is higher than the second resolution threshold.

[0227] Optionally, the processor may also execute program code for the following steps: acquiring first ray sampling data in the original image, wherein the first ray sampling data is used to characterize the sampling ray corresponding to the object to be monitored in the original image; acquiring first voxel features corresponding to the first ray sampling data based on a voxel grid; and using a hybrid neural radiation field model to perform voxel rendering on the original image using the first voxel features to obtain a first rendered image.

[0228] Optionally, the processor may also execute program code for the following steps: acquiring second ray sampling data from an image slice in the first rendered image, wherein the second ray sampling data is used to characterize the sampling ray corresponding to the object to be monitored in the image slice; acquiring second voxel features corresponding to the second ray sampling data based on a voxel grid; and using a super-resolution model to enlarge the resolution of the rendered image corresponding to the second voxel features by a target factor to obtain the second rendered image.

[0229] Optionally, the processor may also execute program code for the following steps: using an initial hybrid neural radiation field model, rendering the first image sample set using light sampling data samples from the first image sample set to obtain a first output image, wherein the initial hybrid neural radiation field model is used to train the hybrid neural radiation field model; using the first output image as the input image of the initial super-resolution model, training the initial super-resolution model to obtain the super-resolution model.

[0230] Optionally, the processor may also execute program code that performs the following steps: obtaining a second output image obtained by magnifying the input image using the initial super-resolution model; establishing a loss function based on the second output image; and adjusting the parameters of the initial super-resolution model based on the loss function to obtain the super-resolution model.

[0231] Optionally, the processor may also execute program code with the following steps: a first loss function, used to make the difference between the target type pixels in the second output image and the target type pixels in the corresponding real image less than a pixel threshold; a second loss function, used to make the difference between the pixel distribution information of the second output image and the pixel distribution information of the corresponding real image less than a distribution information threshold; and a third loss function, used to make the difference between the features of the second output image in the feature space and the features of the corresponding real image in the feature space less than a feature threshold.

[0232] Optionally, the processor may also execute program code that adjusts the parameters of the initial hybrid neural radiation field model based on the gradient of the trained super-resolution model to obtain the hybrid neural radiation field model.

[0233] Optionally, the processor may also execute program code for the following steps: light sampling data samples from the same training batch input to the initial hybrid neural radiation field model, which are from the same image slice in the first image sample set; and light sampling data samples from different training batches input to the initial hybrid neural radiation field model, which are from random image slices in the first image sample set.

[0234] As an optional example, the processor can invoke information and applications stored in memory via a transmission device to perform the following steps: acquiring a first image sample set of the object to be monitored captured from different viewpoints, wherein the resolution of the image samples in the first image sample set is higher than a resolution threshold; jointly training an initial hybrid neural radiation field model and an initial super-resolution model using the first image sample set to obtain a hybrid neural radiation field model and a super-resolution model, wherein the hybrid neural radiation field model is used to store at least the target image information of the first image sample set in a voxel grid, the frequency of the target image information is less than a frequency threshold, and is used to render the original image of the object to be monitored at the current viewpoint to obtain a first rendered image monitored at the current viewpoint, wherein the resolution of the original image is lower than another resolution threshold, the other resolution threshold is lower than the resolution threshold, and the super-resolution model is used to enlarge the resolution of the first rendered image to obtain a second rendered image monitored at the current viewpoint.

[0235] As an optional example, the processor can invoke information and applications stored in memory via a transmission device to perform the following steps: initiate image monitoring of the object to be monitored; display the original image monitored from the current viewpoint on the presentation screen of a virtual reality (VR) device or augmented reality (AR) device, wherein the resolution of the original image is lower than a first resolution threshold; invoke a hybrid neural radiation field model and a super-resolution model generated after joint training, wherein the samples used during joint training are a first image sample set captured from different viewpoints, the resolution of the image samples in the first image sample set is higher than a second resolution threshold, the second resolution threshold is greater than the first resolution threshold, and the hybrid neural radiation field model is used to store at least the target image information of the first image sample set in a voxel grid, the frequency of the target image information is lower than a frequency threshold; control the VR device or AR device to render the original image using the hybrid neural radiation field model to obtain a first rendered image monitored from the current viewpoint, wherein the resolution of the first rendered image is lower than the first resolution threshold; control the VR device or AR device to use the super-resolution model to enlarge the resolution of the first rendered image to obtain a second rendered image monitored from the current viewpoint, wherein the resolution of the second rendered image is higher than the second resolution threshold; drive the VR device or AR device to render and display the second rendered image.

[0236] As an optional example, the processor can invoke information and applications stored in memory via a transmission device to perform the following steps: acquiring the original image monitored from the current viewpoint by invoking a first interface, wherein the first interface includes a first parameter whose value is the original image, and the resolution of the original image is lower than a first resolution threshold; invoking a hybrid neural radiation field model and a super-resolution model generated after joint training, wherein the samples used during joint training are a first image sample set captured from different viewpoints, the resolution of the image samples in the first image sample set is higher than a second resolution threshold, the second resolution threshold is greater than the first resolution threshold, and the hybrid neural radiation field model is used to store at least the target image information of the first image sample set in a voxel grid, the frequency of the target image information being lower than a frequency threshold; rendering the original image using the hybrid neural radiation field model to obtain a first rendered image monitored from the current viewpoint, wherein the resolution of the first rendered image is lower than the first resolution threshold; amplifying the resolution of the first rendered image using the super-resolution model to obtain a second rendered image monitored from the current viewpoint, wherein the resolution of the second rendered image is higher than the second resolution threshold; and outputting the second rendered image by invoking a second interface, wherein the second interface includes a second parameter whose value is the second rendered image.

[0237] This application embodiment uses a hybrid neural radiation field model and a super-resolution model for joint training, so that the optimization of the parameters of the hybrid neural radiation field model and the super-resolution model influence each other. The hybrid neural radiation field model and the super-resolution model generated after joint training process the original image, thereby achieving the technical effect of improving the rendering efficiency of the image and solving the technical problem of low image rendering efficiency.

[0238] Those skilled in the art will understand that Figure 15 The structure shown is for illustrative purposes only. Computer terminal A can also be a smartphone (such as a tablet, a mobile computer, a mobile internet device (MID), a PAD, or other terminal device). Figure 15 This does not limit the structure of the aforementioned computer terminal A. For example, computer terminal A may also include components that are more complex than those described above. Figure 15 Showing more or fewer components (such as network interfaces, display devices, etc.), or having the same Figure 15 The different configurations shown.

[0239] Those skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be implemented by a program instructing the hardware related to the terminal device. The program can be stored in a computer-readable storage medium, which may include: flash drive, read-only memory (ROM), random access memory (RAM), disk or optical disk, etc.

[0240] Example 6

[0241] Embodiments of this application also provide a computer-readable storage medium. Optionally, in this embodiment, the computer-readable storage medium can be used to store the program code executed by the image rendering method of the object provided in Embodiment 1.

[0242] Optionally, in this embodiment, the computer-readable storage medium may be located in any computer terminal in a group of computer terminals in a computer network, or in any mobile terminal in a group of mobile terminals.

[0243] Optionally, in this embodiment, the computer-readable storage medium is configured to store program code for performing the following steps: initiating image monitoring of the object to be monitored, acquiring the original image monitored from the current viewpoint, wherein the resolution of the original image is lower than a first resolution threshold; calling the hybrid neural radiation field model and the super-resolution model generated after joint training, wherein the samples used during joint training are a first image sample set captured from different viewpoints, the resolution of the image samples in the first image sample set is higher than a second resolution threshold, the second resolution threshold is greater than the first resolution threshold, and the hybrid neural radiation field model is used to store at least the target image information of the first image sample set in a voxel grid, the frequency of the target image information is lower than a frequency threshold; rendering the original image using the hybrid neural radiation field model to obtain a first rendered image monitored from the current viewpoint, wherein the resolution of the first rendered image is lower than the first resolution threshold; and amplifying the resolution of the first rendered image using the super-resolution model to obtain a second rendered image monitored from the current viewpoint, wherein the resolution of the second rendered image is higher than the second resolution threshold.

[0244] Optionally, the computer-readable storage medium may also execute program code that performs the following steps: acquiring first ray sampling data in the original image, wherein the first ray sampling data is used to characterize the sampling ray corresponding to the object to be monitored in the original image; acquiring first voxel features corresponding to the first ray sampling data based on a voxel grid; and using a hybrid neural radiation field model to perform voxel rendering on the original image using the first voxel features to obtain a first rendered image.

[0245] Optionally, the aforementioned computer-readable storage medium may also execute program code that performs the following steps: acquiring second ray sampling data from an image slice in a first rendered image, wherein the second ray sampling data is used to characterize the sampling ray corresponding to the object to be monitored in the image slice; acquiring second voxel features corresponding to the second ray sampling data based on a voxel grid; and using a super-resolution model to enlarge the resolution of the rendered image corresponding to the second voxel features by a target factor to obtain a second rendered image.

[0246] Optionally, the computer-readable storage medium may also execute program code that performs the following steps: using an initial hybrid neural radiation field model, rendering the first image sample set using light sampling data samples from the first image sample set to obtain a first output image, wherein the initial hybrid neural radiation field model is used to train the hybrid neural radiation field model; using the first output image as the input image of the initial super-resolution model, training the initial super-resolution model to obtain the super-resolution model.

[0247] Optionally, the aforementioned computer-readable storage medium may also execute program code that performs the following steps: obtaining a second output image obtained by magnifying the input image using the initial super-resolution model; establishing a loss function based on the second output image; and adjusting the parameters of the initial super-resolution model based on the loss function to obtain the super-resolution model.

[0248] Optionally, a first loss function is used to ensure that the difference between the target type pixels in the second output image and the target type pixels in the corresponding real image is less than a pixel threshold; a second loss function is used to ensure that the difference between the pixel distribution information of the second output image and the pixel distribution information of the corresponding real image is less than a distribution information threshold; and a third loss function is used to ensure that the difference between the features of the second output image in the feature space and the features of the corresponding real image in the feature space is less than a feature threshold.

[0249] Optionally, the aforementioned computer-readable storage medium may also execute program code that performs the following steps: adjusting the parameters of the initial hybrid neural radiation field model based on the gradient of the trained super-resolution model to obtain the hybrid neural radiation field model.

[0250] Optionally, the computer-readable storage medium may also execute program code that performs the following steps: light sampling data samples from the same training batch input to the initial hybrid neural radiation field model, which are from the same image slice in the first image sample set; and light sampling data samples from different training batches input to the initial hybrid neural radiation field model, which are from random image slices in the first image sample set.

[0251] As an optional example, a computer-readable storage medium is configured to store program code for performing the following steps: acquiring a first image sample set of the object to be monitored captured at different viewpoints, wherein the resolution of the image samples in the first image sample set is higher than a resolution threshold; jointly training an initial hybrid neural radiation field model and an initial super-resolution model using the first image sample set to obtain a hybrid neural radiation field model and a super-resolution model, wherein the hybrid neural radiation field model is used to store at least the target image information of the first image sample set in a voxel grid, the frequency of the target image information is less than a frequency threshold, and is used to render the original image of the object to be monitored at the current viewpoint to obtain a first rendered image monitored at the current viewpoint, wherein the resolution of the original image is lower than another resolution threshold, the other resolution threshold is less than a resolution threshold, and the super-resolution model is used to enlarge the resolution of the first rendered image to obtain a second rendered image monitored at the current viewpoint.

[0252] As an optional example, a computer-readable storage medium is configured to store program code for performing the following steps: initiating image monitoring of the object to be monitored; displaying the original image monitored from the current viewpoint on the presentation screen of a virtual reality (VR) device or augmented reality (AR) device, wherein the resolution of the original image is lower than a first resolution threshold; invoking a hybrid neural radiation field model and a super-resolution model generated after joint training, wherein the samples used during joint training are a first image sample set captured from different viewpoints, the resolution of the image samples in the first image sample set is higher than a second resolution threshold, the second resolution threshold is greater than the first resolution threshold, and the hybrid neural radiation field model is used to store at least the target image information of the first image sample set in a voxel grid, the frequency of the target image information being lower than a frequency threshold; controlling the VR device or AR device to render the original image using the hybrid neural radiation field model to obtain a first rendered image monitored from the current viewpoint, wherein the resolution of the first rendered image is lower than the first resolution threshold; controlling the VR device or AR device to enlarge the resolution of the first rendered image using the super-resolution model to obtain a second rendered image monitored from the current viewpoint, wherein the resolution of the second rendered image is higher than the second resolution threshold; and driving the VR device or AR device to render and display the second rendered image.

[0253] As an optional example, a computer-readable storage medium is configured to store program code for performing the following steps: acquiring a raw image monitored from the current viewpoint by calling a first interface, wherein the first interface includes a first parameter whose value is the raw image, and the resolution of the raw image is lower than a first resolution threshold; calling a hybrid neural radiation field model and a super-resolution model generated after joint training, wherein the samples used during joint training are a first image sample set captured from different viewpoints, the resolution of the image samples in the first image sample set is higher than a second resolution threshold, the second resolution threshold is greater than the first resolution threshold, and the hybrid neural radiation field model is used to store at least the target image information of the first image sample set in a voxel grid, the frequency of the target image information being lower than a frequency threshold; rendering the raw image using the hybrid neural radiation field model to obtain a first rendered image monitored from the current viewpoint, wherein the resolution of the first rendered image is lower than the first resolution threshold; using the super-resolution model to enlarge the resolution of the first rendered image to obtain a second rendered image monitored from the current viewpoint, wherein the resolution of the second rendered image is higher than the second resolution threshold; and outputting the second rendered image by calling a second interface, wherein the second interface includes a second parameter whose value is the second rendered image.

[0254] The sequence numbers of the embodiments in this application are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.

[0255] In the above embodiments of this application, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments.

[0256] In the several embodiments provided in this application, it should be understood that the disclosed technical content can be implemented in other ways. The device embodiments described above are merely illustrative; for example, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection of units or modules may be electrical or other forms.

[0257] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0258] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0259] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as a USB flash drive, read-only memory (ROM), random access memory (RAM), portable hard drive, magnetic disk, or optical disk.

[0260] The above are merely preferred embodiments of this application. It should be noted that those skilled in the art can make various improvements and modifications without departing from the principles of this application, and these improvements and modifications should also be considered within the scope of protection of this application.

Claims

1. A method for rendering an image of an object, characterized in that, include: Initiate image monitoring of the object to be monitored and acquire the original image monitored from the current viewpoint, wherein the resolution of the original image is lower than a first resolution threshold; The hybrid neural radiation field model and super-resolution model generated after joint training are invoked. The samples used during joint training are a first image sample set captured from different viewpoints. The resolution of the image samples in the first image sample set is higher than a second resolution threshold. The second resolution threshold is greater than the first resolution threshold. The hybrid neural radiation field model is used to store at least the target image information of the first image sample set in a voxel grid. The frequency of the target image information is less than a frequency threshold. The original image is rendered using the hybrid neural radiation field model to obtain a first rendered image monitored under the current viewpoint, wherein the resolution of the first rendered image is less than the first resolution threshold. The resolution of the first rendered image is magnified using the super-resolution model to obtain a second rendered image monitored under the current viewpoint, wherein the resolution of the second rendered image is higher than the second resolution threshold.

2. The method according to claim 1, characterized in that, The target image information is used to at least indicate the viewpoint position of the object under monitoring when image monitoring is performed on the object under monitoring.

3. The method according to claim 1, characterized in that, The original image is rendered using the hybrid neural radiation field model to obtain a first rendered image monitored under the current viewpoint, including: Acquire first ray sampling data from the original image, wherein the first ray sampling data is used to characterize the sampling ray corresponding to the object to be monitored in the original image; The first voxel feature corresponding to the first ray sampling data is obtained based on the voxel grid; Using the hybrid neural radiation field model, the original image is voxel-rendered using the first voxel features to obtain the first rendered image.

4. The method according to claim 1, characterized in that, The resolution of the first rendered image is increased using the super-resolution model to obtain a second rendered image monitored at the current viewpoint, including: Acquire second ray sampling data from an image slice in the first rendered image, wherein the second ray sampling data is used to characterize the sampling ray corresponding to the object to be monitored in the image slice; The second voxel feature corresponding to the second ray sampling data is obtained based on the voxel grid; Using the super-resolution model, the resolution of the rendered image corresponding to the second voxel feature is magnified by the target factor to obtain the second rendered image.

5. The method according to claim 4, characterized in that, The target multiplier is used to perform mixed degradation processing on image sample sets from different viewpoints with resolutions higher than the second resolution threshold to obtain a second image sample set with a resolution lower than the first resolution threshold. The second image sample set is used to obtain an initial super-resolution model based on convolution training. The initial super-resolution model is used to train the super-resolution model.

6. The method according to claim 5, characterized in that, The method further includes: Using an initial hybrid neural radiation field model, the first image sample set is rendered using light sampling data samples from the first image sample set to obtain a first output image, wherein the initial hybrid neural radiation field model is used to train the hybrid neural radiation field model. The first output image is used as the input image of the initial super-resolution model to train the initial super-resolution model, thereby obtaining the super-resolution model.

7. The method according to claim 6, characterized in that, The first output image is used as the input image of the initial super-resolution model, and the initial super-resolution model is trained to obtain the super-resolution model, including: Obtain the second output image obtained by magnifying the input image using the initial super-resolution model; A loss function is established based on the second output image; The parameters of the initial super-resolution model are adjusted based on the loss function to obtain the super-resolution model.

8. The method according to claim 7, characterized in that, The loss function includes at least one of the following: A first loss function is used to ensure that the difference between the target type pixels in the second output image and the target type pixels in the corresponding real image is less than a pixel threshold. The second loss function is used to ensure that the difference between the pixel distribution information of the second output image and the pixel distribution information of the corresponding real image is less than the distribution information threshold. The third loss function is used to ensure that the difference between the features of the second output image in the feature space and the features of the corresponding real image in the feature space is less than the feature threshold.

9. The method according to claim 7, characterized in that, The method further includes: The parameters of the initial hybrid neural radiation field model are adjusted based on the gradient of the trained super-resolution model to obtain the hybrid neural radiation field model.

10. The method according to claim 6, characterized in that, The light sampling data samples from the same training batch input to the initial hybrid neural radiation field model are from the same image slice in the first image sample set, while the light sampling data samples from different training batches input to the initial hybrid neural radiation field model are from random image slices in the first image sample set.

11. The method according to claim 6, characterized in that, The initial hybrid neural radiation field model was trained on a third image sample set with a resolution lower than the first resolution threshold.

12. A method for rendering an image of an object, characterized in that, include: Obtain a first image sample set of the object to be monitored captured from different perspectives, wherein the resolution of the image samples in the first image sample set is higher than a resolution threshold; The initial hybrid neural radiation field model and the initial super-resolution model are jointly trained using the first image sample set to obtain the hybrid neural radiation field model and the super-resolution model. The hybrid neural radiation field model is used to store the target image information of the first image sample set in a voxel grid. The frequency of the target image information is less than a frequency threshold. It is also used to render the original image of the object to be monitored under the current viewpoint to obtain a first rendered image under the current viewpoint. The resolution of the original image is lower than another resolution threshold, and the other resolution threshold is less than the resolution threshold. The super-resolution model is used to enlarge the resolution of the first rendered image to obtain a second rendered image under the current viewpoint.

13. A method for rendering an image of an object, characterized in that, include: Initiate image monitoring of the object to be monitored, and display the original image monitored from the current perspective on the display screen of the virtual reality (VR) device or augmented reality (AR) device, wherein the resolution of the original image is lower than a first resolution threshold; The hybrid neural radiation field model and super-resolution model generated after joint training are invoked. The samples used during joint training are a first image sample set captured from different viewpoints. The resolution of the image samples in the first image sample set is higher than a second resolution threshold. The second resolution threshold is greater than the first resolution threshold. The hybrid neural radiation field model is used to store at least the target image information of the first image sample set in a voxel grid. The frequency of the target image information is less than a frequency threshold. The VR device or the AR device is controlled to render the original image using the hybrid neural radiation field model to obtain a first rendered image monitored under the current viewpoint, wherein the resolution of the first rendered image is less than the first resolution threshold. The VR device or the AR device is controlled to use the super-resolution model to magnify the resolution of the first rendered image to obtain a second rendered image monitored under the current viewpoint, wherein the resolution of the second rendered image is higher than the second resolution threshold. The VR device or AR device is driven to render and display the second rendered image.

14. A method for rendering an image of an object, characterized in that, include: The original image monitored from the current viewpoint is obtained by calling the first interface, wherein the first interface includes a first parameter, the parameter value of the first parameter is the original image, and the resolution of the original image is lower than a first resolution threshold. The hybrid neural radiation field model and super-resolution model generated after joint training are invoked. The samples used during joint training are a first image sample set captured from different viewpoints. The resolution of the image samples in the first image sample set is higher than a second resolution threshold. The second resolution threshold is greater than the first resolution threshold. The hybrid neural radiation field model is used to store at least the target image information of the first image sample set in a voxel grid. The frequency of the target image information is less than a frequency threshold. The original image is rendered using the hybrid neural radiation field model to obtain a first rendered image monitored under the current viewpoint, wherein the resolution of the first rendered image is less than the first resolution threshold. The resolution of the first rendered image is magnified using the super-resolution model to obtain a second rendered image monitored under the current viewpoint, wherein the resolution of the second rendered image is higher than the second resolution threshold. The second rendered image is output by calling the second interface, wherein the second interface includes a second parameter, and the value of the second parameter is the second rendered image.